mlx-dspark

Models

Qwen3-8B

The agent workhorse: 2.19× mean, and the fastest wall clock under Claude Code because its prefix cache trims cleanly.

Made by
Alibaba Qwen
Size
8B
Architecture
Dense transformer
Features
Tool callingThinking

Run it

mlx-dspark serve --model mlx-community/Qwen3-8B-8bit

Serves an OpenAI + Anthropic API on http://127.0.0.1:8080. For a one-off answer, use mlx-dspark generate --model mlx-community/Qwen3-8B-8bit --prompt "…".

2.19× faster than plain decoding, mean of chat, code and math
50–83 tokens per second (plain: 29.5)
3.37 tokens accepted per round
~11 GB peak memory at chat length
Chat 1.71×
Code 2.03×
Math 2.83×
Dashed: plain decoding (1×). Solid: mlx-dspark on the same prompt, draft cap 7.

Drafter

The drafter downloads with the model the first time you run it, and loads 4-bit quantized (acceptance doesn't depend on the drafter's precision).

Reading the prompt

Prefill runs at 636 tok/s with CPU co-prefill (488 without), so a 20,000-token prompt takes about 31 s the first time. After that the prefix cache skips it.

Measured 2026-09-27 on an M4 Pro, 48 GB with mlx 0.32.2.

Qwen3-8B with DeepSeek's official DSpark drafter. It's the model this project measured agents on.

Why it wins for agents#

In a real Claude Code session (read a buggy file, fix it with the Edit tool), Qwen3-8B finished in about 2:20, nearly 2× faster than two models with higher acceptance. Claude Code sends 18–26k tokens of system prompt and tool schemas with every request, so prefill dominates the clock. A dense model's prefix cache trims cleanly, so each turn reuses almost all of it.

Use it with --no-thinking for agent work: with thinking on, the model reasons before every tool call. The same task took 3:17 and 2762 output tokens, against ~2:20 and 169 without.

Also#

z-lab's DFlash drafter (z-lab/Qwen3-8B-DFlash-b16) loads with --mode dflash. At this size it is a wash against DSpark (its full block is a net loss even in z-lab's own runner), so DSpark stays the default.

Numbers not matching your Mac? They shouldn't be identical: the draft cap is derived from your machine's measured curves. Run mlx-dspark benchmark --model mlx-community/Qwen3-8B-8bit --trials 3 to get your own. See how these are measured.