mlx-dspark

Models

Qwen3-4B

Small and quick: ~89–126 tok/s at 1.98× in ~8 GB, with DeepSeek's official drafter.

Made by
Alibaba Qwen
Size
4B
Architecture
Dense transformer
Features
Tool callingThinking

Run it

mlx-dspark serve --model mlx-community/Qwen3-4B-8bit

Serves an OpenAI + Anthropic API on http://127.0.0.1:8080. For a one-off answer, use mlx-dspark generate --model mlx-community/Qwen3-4B-8bit --prompt "…".

1.98× faster than plain decoding, mean of chat, code and math
89–127 tokens per second (plain: 51.7)
3.23 tokens accepted per round
~8 GB peak memory at chat length
Chat 1.78×
Code 1.72×
Math 2.46×
Dashed: plain decoding (1×). Solid: mlx-dspark on the same prompt, draft cap 7.

Drafter

The drafter downloads with the model the first time you run it, and loads 4-bit quantized (acceptance doesn't depend on the drafter's precision).

Reading the prompt

Prefill runs at 1141 tok/s with CPU co-prefill (886 without), so a 20,000-token prompt takes about 18 s the first time. After that the prefix cache skips it.

Measured 2026-09-27 on an M4 Pro, 48 GB with mlx 0.32.2.

Qwen3-4B with DeepSeek's official DSpark drafter, and z-lab's DFlash drafter available via --mode dflash. The plain baseline here matches mlx_lm.generate (51–52 tok/s either way), so the speedup is against a fair reference.

Batching also works well on this dense model: 4 concurrent requests reach ~2.5× the aggregate throughput of serving them one at a time (--max-batch 4).

Numbers not matching your Mac? They shouldn't be identical: the draft cap is derived from your machine's measured curves. Run mlx-dspark benchmark --model mlx-community/Qwen3-4B-8bit --trials 3 to get your own. See how these are measured.