Up to 4× faster local models.Every token exactly the same.
mlx-dspark speeds up language models on Apple Silicon with speculative decoding: a small drafter proposes the next few tokens, and your model checks them all in one pass.
pip install mlx-dspark
Watch a model race. Pick one:
- if drafted
- if accepted
- in rejected
- if the model's own token
One round, slowed down.Draft, verify, accept.
-
Draft
A drafter a fraction of your model's size reads its hidden state and proposes the next block of tokens, all at once.
-
Verify
Your model scores every drafted position in a single forward pass. On a Mac that costs close to one normal step, because decoding is limited by reading the weights from memory, and one pass reads them once.
-
Accept
It keeps the drafts up to the first one it disagrees with, then adds its own next token. Worst case is one token per pass, like normal decoding. Best case is the whole block plus one.
Because your model checks every token, the output is the one it would have produced on its own: speculative decoding here is lossless, with greedy and sampled decoding alike. How it works in detail
Find the right modelfor the Mac you have.
Every model below is measured and vouched for: its drafter downloads automatically, with settings tuned to your machine on first run.
Memory is each model's measured peak at chat length, judged against about 70% of your Mac's RAM (macOS and your other apps need the rest). Long conversations add KV cache on top: a 27B adds about 11 GB at 128k tokens.
Measured on a real Mac,not promised.
Speedup over plain decoding of the same model on three kinds of prompt. Each pair runs at the draft cap it derives by itself, with no flags. Full results and method
Point at a model to see its numbers. Click it to open its page.
Use it with the tools you have.One server, three API dialects.
One engine, four ways in. The server speaks the OpenAI, Anthropic and Responses APIs on one port, so the tools you already use connect by changing one URL.
- Claude Code, fully local, in one command
- Codex through the Responses API
- pi, OpenCode, Open WebUI, Continue and anything OpenAI-compatible
- Python, if you'd rather script it
mlx-dspark generate --model mlx-community/Qwen3-8B-8bit \
--prompt "Explain how rainbows form."
mlx-dspark serve --model mlx-community/Qwen3-8B-8bit
curl http://127.0.0.1:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "Qwen3-8B-8bit",
"messages": [{"role": "user", "content": "Hi!"}]}'
mlx-dspark serve --model mlx-community/Qwen3-8B-8bit --no-thinking
# in a second terminal
mlx-dspark claude
from mlx_dspark import load_pair, speculative_generate
target, tok, drafter, cfg = load_pair("mlx-community/Qwen3-8B-8bit")
res = speculative_generate(target, tok, drafter, "Explain rainbows.")
print(res.text, res.mean_accept_len)
Or skip the terminal.There is a Mac app.
The native Mac app wraps the same engine: chat with saved sessions, a model manager that tells you whether a model fits before you download it, one-click setup for coding agents, and a Race that checks the lossless claim token by token on your own prompt.
brew tap ARahim3/mlx-dspark https://github.com/ARahim3/mlx-dspark
brew install --cask mlx-dspark
Install the Mac app or grab the DMG from Releases.
One command to start.Your Mac does the tuning.
pip install mlx-dspark
mlx-dspark serve --model mlx-community/Qwen3-8B-8bit
The first run downloads the model and its drafter, then spends about five seconds measuring your Mac to pick its settings. Quickstart