mlx-dspark

Models

Results and method

Every measured number in one place, how it was measured, and why some models gain far more than others.

Point at a model to see its numbers. Click it to open its page.

The dashed line is plain decoding of the same model. Everything to its right is time saved, with identical output.

The full table#

Each model at the draft cap it derives on its own (no flags), against plain greedy decoding of the same model. M4 Pro, 48 GB, warm, 4-bit drafter, 200 new tokens, end-to-end tokens per second, median of 3 runs over the three prompts.

ModelDrafterCapAcceptedPlainmlx-dsparkSpeedupChat / code / mathMeasured
Qwen3.8-27B 8-bit DFlash 2 7 5.42 8.2 30.8 3.74× 3.07 / 3.93 / 4.23 2026-09-28
LFM2.5-1.2B bf16 DSpark 7 5.28 99.6 333.5 3.35× 2.40 / 3.83 / 3.82 2026-09-27
Gemma-4 12B 8-bit DSpark 7 4.98 18.4 59.8 3.25× 2.62 / 2.90 / 4.24 2026-09-27
MiniCPM5-2B bf16 DSpark 7 4.87 50.6 159.7 3.16× 2.12 / 3.15 / 4.21 2026-09-27
Nanbeige4.2-3B bf16 DSpark 7 3.95 17.3 50.6 2.92× 1.98 / 2.50 / 4.31 2026-09-27
LFM2.5-2.6B bf16 DSpark 7 4.00 43.6 121.7 2.79× 2.24 / 2.65 / 3.49 2026-09-27
Qwen3.8-27B 4-bit DFlash 2 7 5.09 14.7 40.0 2.72× 2.05 / 3.12 / 3.02 2026-09-28
Muse-Glimmer-30B 8-bit DSpark (community) 4 3.31 8.2 20.2 2.47× 1.97 / 2.45 / 2.99 2026-08-12*
Ornith-1.0-9B 8-bit DSpark (community) 4 3.64 26.7 64.2 2.40× 2.21 / 2.53 / 2.48 2026-07-18*
Qwen3.6-27B 8-bit DSpark (community) 4 3.15 8.4 19.2 2.29× 2.26 / 1.96 / 2.67 2026-07-25*
Qwen3-8B 8-bit DSpark 7 3.37 29.5 64.5 2.19× 1.71 / 2.03 / 2.83 2026-09-27
Qwen3-14B 8-bit DSpark 4 2.87 15.3 31.0 2.03× 1.62 / 2.11 / 2.36 2026-07-22*
Qwen3-4B 8-bit DSpark 7 3.23 51.7 102.5 1.98× 1.78 / 1.72 / 2.46 2026-09-27
Muse-Glimmer-30B 4-bit DSpark (community) 2 2.45 14.0 24.4 1.74× 1.57 / 1.70 / 1.94 2026-08-12*
Qwen3.6-35B-A3B 4-bit DSpark (community) adaptive 4.72 86.9 114.5 1.32× 1.05 / 1.24 / 1.67 2026-07-26*
LFM2.5-8B-A1B bf16 DSpark 4 — 65.0 82.0 1.26× 1.04 / 1.30 / 1.44 2026-08-22*
Nemotron-3.5-Lightning-30B-A3B 4-bit DSpark (NVIDIA) 3 3.28 91.4 100.9 1.10× 0.95 / 1.23 / 1.13 2026-08-12*
Ternary-Bonsai-27B 2-bit (ternary) DSpark 2 2.60 25.4 27.2 1.07× 1.01 / 1.13 / 1.07 2026-07-16*

Rows marked * were measured before the September 2026 verify kernels. Those kernels moved every 8-bit model re-measured since from cap 4 to cap 7 (Gemma-4 12B went from 2.82× to 3.25× in the same run), so the older 8-bit rows are probably conservative. Qwen3.6-35B-A3B's row uses --confidence-threshold 0.3; its zero-flag default measures 1.27×. The per-model pages have each row's flags and caveats.

Reproduce any row on your own Mac:

mlx-dspark benchmark --model mlx-community/Qwen3-8B-8bit --trials 3

Two drafters on Qwen3.8-27B#

Qwen3.8-27B has two strong drafters: Inco AI's DFlash 2 (the default) and Red Hat's DSpark head (--mode dspark). Both verify the same width (cap 7). Paired in one process, arms interleaved per prompt, High Power mode:

QuantDrafterMeanChatCodeMathAcceptedtok/s
4-bit (plain 14.7 tok/s)DFlash 2 default 2.72×2.05×3.12×3.02×5.0940.0
4-bitDSpark (Red Hat) 2.63×2.25×2.81×2.83×4.8338.6
8-bit (plain 8.2 tok/s)DFlash 2 default 3.74×3.07×3.93×4.23×5.4230.8
8-bitDSpark (Red Hat) 3.79×3.03×4.01×4.32×5.5531.1

DFlash 2 leads at 4-bit; the two tie at 8-bit, which is the best DSpark-mode ratio in the project. DFlash 2 adds a candidate-path selector (the model's top-16 candidates per position, walked as one coherent chain) and dynamic convolutions to the DFlash backbone. Acceptance rises by about 1.1–1.4 tokens without widening the verify, which is exactly the kind of gain that converts on Apple Silicon.

At long context#

Agent clients send 10–40k-token prompts. Decode speed at depth, against plain decoding at the same depth, accepted tokens per round in parentheses (agent-style content, thinking off, best of 2):

Model and drafterContextPlainCap 3Cap 7
Qwen3.8-27B 4-bit, DFlash 216k14.1 tok/s 1.69× (2.74)1.79× (3.16)
32k12.9 tok/s 1.66× (2.58)1.85× (3.24)
Qwen3.8-27B 4-bit, DSpark (Red Hat)16k14.2 tok/s 1.89× (3.0)1.89× (3.38)
32k13.1 tok/s 1.76× (2.74)1.83× (3.32)
MiniCPM5-2B bf16, DSpark16k42.2 tok/s 1.62× (2.21)1.54× (2.5)
32k36.0 tok/s 2.06× (2.78)2.09× (3.51)
Nanbeige4.2-3B bf16, DSpark16k13.9 tok/s 1.97× (2.36)1.89× (2.78)
32k10.6 tok/s 2.21× (2.74)1.83× (3.2)

Speedups at depth are lower than at chat length for a structural reason: attention over a long KV cache is extra verify work that a single decode step pays only once. You don't need to pick the cap: with no --max-draft, it shrinks automatically when the measured depth cost says a narrower verify pays. See Long context.

Reading the prompt#

Prefill speed decides how long you wait before the first token. With CPU co-prefill on (the default for the CLI and server), 2048-token prompt, median of 3:

ModelPrefillWithout CPU co-prefillA 20k-token prompt takes
LFM2.5-1.2B bf163240 tok/s—about 6 s
LFM2.5-8B-A1B bf162170 tok/s—about 9 s
MiniCPM5-2B bf161640 tok/s—about 12 s
LFM2.5-2.6B bf161420 tok/s—about 14 s
Qwen3.6-35B-A3B 4-bit960 tok/s—about 21 s
Qwen3-4B 8-bit1141 tok/s886 tok/s (1.29×)about 18 s
Nanbeige4.2-3B bf16526 tok/s—about 38 s
Qwen3-8B 8-bit636 tok/s488 tok/s (1.30×)about 31 s
Gemma-4 12B 8-bit381 tok/s293 tok/s (1.30×)about 52 s
Qwen3.8-27B 8-bit182 tok/s138 tok/s (1.32×)about 110 s
Qwen3.8-27B 4-bit184 tok/s130 tok/s (1.42×)about 109 s
Qwen3.6-27B 8-bit126 tok/s—about 159 s

Model size barely predicts this. Prefill is compute-bound, so small dense models lead, and mixtures of experts punch far above their total size: an A3B model does only about 3.8B parameters' worth of arithmetic per token. A dash means CPU co-prefill doesn't apply to that model yet (bf16 layers and expert layers aren't split). Either way this is a first-request cost: the prefix cache skips it on every later turn.

Copy-heavy editing goes further#

The table above is fresh generation. When a model re-emits or refactors code already in its context, the everyday agent workload, lookup drafts reach well past it with output still bit-identical:

Task Model Before lookup drafts With them
Re-emit a file Gemma-4 12B 8-bit 3.03× 4.51× (75 tok/s)
Rename refactor Gemma-4 12B 8-bit 4.33×
Rename refactor Ornith-1.0-9B 8-bit 2.79× 3.57× (93 tok/s)
Re-emit a file Ornith-1.0-9B 8-bit 2.45×

Chat and fresh code are unchanged. See lookup drafts.

Why some models gain more than others#

The drafter sets the ceiling. Acceptance, how many drafted tokens survive per round, is set by how well the drafter matches the model. Official drafters and carefully qualified community ones accept 4–5.5 tokens per round on these prompts.

What a step costs sets the payoff. Speculation saves whole passes through the model. If a pass is cheap, there's little to save. That's why:

Precision changes the verify curve. 2-bit Bonsai is compute-bound from a verify width of 2, which caps it near 1.1×. 8-bit stays flat to width 8 with the current kernels, which is why most 8-bit models draft a full block of 7. See Draft caps.

How it's measured#

Your Mac will differ, which is the point of deriving the cap on it. A faster chip can land on a lower cap, because cheaper verification shifts the optimum. Run the benchmark and compare ratios, not absolute tokens per second.