mlx-dspark

How it works

Reading the prompt

Decoding speed is half your wait. The other half is prefill, reading the prompt, and for agents it's most of it.

Everything else on this site measures decoding, how fast tokens come out once generation starts. The other half of your wall clock is prefill: reading the prompt. For a chat message that's nothing. For a pasted file, a long conversation or an agent (Claude Code sends 18–26k tokens with every request), it's most of the wait.

ModelPrefillWithout CPU co-prefillA 20k-token prompt takes
LFM2.5-1.2B bf163240 tok/s—about 6 s
LFM2.5-8B-A1B bf162170 tok/s—about 9 s
MiniCPM5-2B bf161640 tok/s—about 12 s
LFM2.5-2.6B bf161420 tok/s—about 14 s
Qwen3.6-35B-A3B 4-bit960 tok/s—about 21 s
Qwen3-4B 8-bit1141 tok/s886 tok/sabout 18 s
Nanbeige4.2-3B bf16526 tok/s—about 38 s
Qwen3-8B 8-bit636 tok/s488 tok/sabout 31 s
Gemma-4 12B 8-bit381 tok/s293 tok/sabout 52 s
Qwen3.8-27B 8-bit182 tok/s138 tok/sabout 110 s
Qwen3.8-27B 4-bit184 tok/s130 tok/sabout 109 s
Qwen3.6-27B 8-bit126 tok/s—about 159 s

M4 Pro, 2048-token prompt, median of 3, default settings.

Model size barely predicts this. Prefill is compute-bound, so small dense models lead, and mixtures of experts punch far above their total size: an A3B model does only about 3.8B parameters' worth of arithmetic per token, however many experts it stores. The reverse shows up on Qwen3.8-27B: the 4-bit build prefills at the same rate as the 8-bit, because weight bits change decode speed (bandwidth-bound), not prefill. Looped Nanbeige runs every token through its layers twice, so the "3B" pays a 6B's arithmetic.

What mlx-dspark does#

Skips work. Prefill logits that every caller throws away aren't computed, and wide quantized weights are dequantized once instead of once per output tile. That's worth 1.07–1.15×, bit-identical, with no extra memory.

Adds a second engine: CPU co-prefill. With that, the GPU runs prefill at about 85% of its measured matrix-multiply peak, so going faster needs more hardware. Above a measured row count, each wide quantized matmul hands a calibrated fraction of its rows (about 0.3 on an M4 Pro) to the CPU's matrix units, which run alongside the GPU on the same arrays, with no copy and no second thread. That's the 1.3–1.4× column above, for about 0.4 GB of extra peak memory.

--cpu-split 0 turns it off; /health reports the calibrated setting (cpu_split), and the serve banner prints it.

The Apple Neural Engine was measured for the same job and doesn't fit: its fast weight formats can't hold MLX's group-quantized weights exactly, and the exact fp16 form of a 27B model's MLP would be a 34 GB copy.

The real lever: don't read it twice#

A 20k-token first request is slow on any engine. What makes a long conversation or an agent session usable is that every turn after the first skips the part it has already read: 62 s cold, then about 1 s on the next turn, on a 27B with an 8k-token system prompt. That's the prefix cache.

Watching a long cold prompt? GET /events streams prefill progress ({"type": "prefill", "processed": …, "total": …}), so a client can show a progress bar instead of what looks like a hung server.