gitkaz/mlx_gguf_server
This is a FastAPI based LLM server. Load multiple LLM models (MLX or llama.cpp) simultaneously using multiprocessing.
43
/ 100
Emerging
No Package
No Dependents
Maintenance
13 / 25
Adoption
6 / 25
Maturity
9 / 25
Community
15 / 25
Stars
17
Forks
4
Language
Python
License
MIT
Category
Last pushed
Mar 27, 2026
Commits (30d)
0
Get this data via API
curl "https://pt-edge.onrender.com/api/v1/quality/transformers/gitkaz/mlx_gguf_server"
Open to everyone — 100 requests/day, no key needed. Get a free key for 1,000/day.
Higher-rated alternatives
beehive-lab/GPULlama3.java
GPU-accelerated Llama3.java inference in pure Java using TornadoVM.
48
srgtuszy/llama-cpp-swift
Swift bindings for llama-cpp library
37
RhinoDevel/mt_llm
Pure C wrapper library to use llama.cpp with Linux and Windows as simple as possible.
33
JackZeng0208/llama.cpp-android-tutorial
llama.cpp tutorial on Android phone
33
dougeeai/llama-cpp-python-wheels
Pre-built wheels for llama-cpp-python across platforms and CUDA versions
30