mlx-dspark

Models

Qwen3.6-27B

2.29× mean on the 8-bit target with a strong block-15 community drafter.

Made by
Alibaba Qwen
Size
27B
Architecture
Hybrid: linear + full attention
Features
Tool callingThinking

Run it

mlx-dspark serve --model mlx-community/Qwen3.6-27B-8bit

Serves an OpenAI + Anthropic API on http://127.0.0.1:8080. For a one-off answer, use mlx-dspark generate --model mlx-community/Qwen3.6-27B-8bit --prompt "…".

2.29× faster than plain decoding, mean of chat, code and math
16–22 tokens per second (plain: 8.4)
3.15 tokens accepted per round
~32 GB peak memory at chat length
Chat 2.26×
Code 1.96×
Math 2.67×
Dashed: plain decoding (1×). Solid: mlx-dspark on the same prompt, draft cap 4.

Drafter

The drafter downloads with the model the first time you run it, and loads 4-bit quantized (acceptance doesn't depend on the drafter's precision).

Reading the prompt

Prefill runs at 126 tok/s, so a 20,000-token prompt takes about 159 s the first time. After that the prefix cache skips it.

Measured 2026-07-25 on an M4 Pro, 48 GB with mlx 0.32.0. This predates the September 2026 verify kernels, which raised every 8-bit model re-measured since, so it is probably conservative.

Qwen3.6-27B with satgeze's community DSpark drafter: a block-15 head (7 is typical), trained against the bf16 target and warm-started from z-lab's DFlash head for the same model.

The rule of thumb it follows: match the target's precision to what the drafter was trained against. A bf16-trained drafter wants the 8-bit target. The 4-bit target resolves the same drafter and should trade ratio for absolute speed, but it isn't measured.

These numbers predate the September 2026 verify kernel, which moved other 8-bit targets from cap 4 to 7.

Numbers not matching your Mac? They shouldn't be identical: the draft cap is derived from your machine's measured curves. Run mlx-dspark benchmark --model mlx-community/Qwen3.6-27B-8bit --trials 3 to get your own. See how these are measured.