How it works
Long context
Agents send 10–40k-token prompts. What changes at that depth, what mlx-dspark does about it, and how to keep memory in check.
Agent clients like Claude Code, Codex and pi send 10–40k tokens with every request. At that depth, speculative decoding used to fade on a Mac: at 32k tokens, DSpark on Qwen3.8-27B could fall below plain decoding. Three engine changes (v0.20.0) hold the speedup there:
- A multi-row attention kernel. MLX's decode attention re-reads the whole KV cache once per verified row. The new kernel reads each KV tile once for all rows, and also covers the DSpark drafter's own attention over the context.
- A preallocated drafter context. The DSpark drafter no longer copies its whole context twice per layer per round.
- A 4096-token drafter window. DeepSpec-style heads lose acceptance when they attend a context far longer than they were trained on. Attending only the most recent 4096 tokens restores their short-context acceptance at any depth. Heads trained with their own sliding window (DFlash 2, Red Hat's Qwen3.8 head, Nemotron, Muse) keep theirs.
Measured at 16k and 32k#
Decode speed, prompt reading excluded, with agent-style content (source code plus a coding task, thinking off), against plain decoding at the same depth. Accepted tokens per round in parentheses:
| Model and drafter | Context | Plain | Cap 3 | Cap 7 |
|---|---|---|---|---|
| Qwen3.8-27B 4-bit, DFlash 2 | 16k | 14.1 tok/s | 1.69× (2.74) | 1.79× (3.16) |
| 32k | 12.9 tok/s | 1.66× (2.58) | 1.85× (3.24) | |
| Qwen3.8-27B 4-bit, DSpark (Red Hat) | 16k | 14.2 tok/s | 1.89× (3.0) | 1.89× (3.38) |
| 32k | 13.1 tok/s | 1.76× (2.74) | 1.83× (3.32) | |
| MiniCPM5-2B bf16, DSpark | 16k | 42.2 tok/s | 1.62× (2.21) | 1.54× (2.5) |
| 32k | 36.0 tok/s | 2.06× (2.78) | 2.09× (3.51) | |
| Nanbeige4.2-3B bf16, DSpark | 16k | 13.9 tok/s | 1.97× (2.36) | 1.89× (2.78) |
| 32k | 10.6 tok/s | 2.21× (2.74) | 1.83× (3.2) |