mlx-dspark

Models

Qwen3.6-35B-A3B

Raw speed above all: ~91–145 tok/s from a 35B MoE. The ratio is modest (1.3×) because the baseline is already fast.

Made by
Alibaba Qwen
Size
35B (3.8B active)
Architecture
Hybrid mixture of experts (~3.8B active)
Features
Tool callingThinking

Run it

mlx-dspark serve --model mlx-community/Qwen3.6-35B-A3B-4bit --confidence-threshold 0.3

Serves an OpenAI + Anthropic API on http://127.0.0.1:8080. For a one-off answer, use mlx-dspark generate --model mlx-community/Qwen3.6-35B-A3B-4bit --prompt "…".

1.32× faster than plain decoding, mean of chat, code and math
91–145 tokens per second (plain: 86.9)
4.72 tokens accepted per round
~23 GB peak memory at chat length
Chat 1.05×
Code 1.24×
Math 1.67×
Dashed: plain decoding (1×). Solid: mlx-dspark on the same prompt, adaptive cap, --confidence-threshold 0.3.

Drafter

The drafter downloads with the model the first time you run it, and loads 4-bit quantized (acceptance doesn't depend on the drafter's precision). Lookup drafts are off by default for this pair (measured as a net loss); --lookup-drafts turns them on.

Reading the prompt

Prefill runs at 960 tok/s, so a 20,000-token prompt takes about 21 s the first time. After that the prefix cache skips it.

Measured 2026-07-26 on an M4 Pro, 48 GB with mlx 0.32.0. This predates the September 2026 verify kernels, which raised every 8-bit model re-measured since, so it is probably conservative.

Only ~3.8B of this model's 35B parameters are active per token, so plain decoding already runs at 86.9 tok/s, faster than every other target here, including the 4B. The drafter is good: acceptance reaches 7.0 tokens per round on math. It still converts to only 1.32×, because speculation's value scales with what a target step costs, and here a step is ~11.5 ms while the dense 1.53B drafter costs ~5.7 ms of every round.

Flags that matter#

  • --confidence-threshold 0.3 lifts it from 1.27× (the zero-flag default) to 1.32×, and math from 1.50× to 1.67×. The verify curve rises from the first extra row and acceptance swings from 2.8 (chat) to 7.0 (math), so deciding per round how far to draft pays off. 0.7 over-throttles.
  • Lookup drafts are off by default: on a mixture of experts every extra verify row pulls in a fresh set of experts, so a free draft is not free.

For raw tokens per second nothing here decodes faster at this size.

Numbers not matching your Mac? They shouldn't be identical: the draft cap is derived from your machine's measured curves. Run mlx-dspark benchmark --model mlx-community/Qwen3.6-35B-A3B-4bit --trials 3 to get your own. See how these are measured.