mlx-dspark

Models

Qwen3-14B

2.03× mean with DeepSeek's official drafter. Measured before the September kernels, so likely higher today.

Made by
Alibaba Qwen
Size
14B
Architecture
Dense transformer
Features
Tool callingThinking

Run it

mlx-dspark serve --model mlx-community/Qwen3-14B-8bit

Serves an OpenAI + Anthropic API on http://127.0.0.1:8080. For a one-off answer, use mlx-dspark generate --model mlx-community/Qwen3-14B-8bit --prompt "…".

2.03× faster than plain decoding, mean of chat, code and math
25–36 tokens per second (plain: 15.3)
2.87 tokens accepted per round
~19 GB peak memory at chat length
Chat 1.62×
Code 2.11×
Math 2.36×
Dashed: plain decoding (1×). Solid: mlx-dspark on the same prompt, draft cap 4.

Drafter

The drafter downloads with the model the first time you run it, and loads 4-bit quantized (acceptance doesn't depend on the drafter's precision).

Measured 2026-07-22 on an M4 Pro, 48 GB with mlx 0.32.0. This predates the September 2026 verify kernels, which raised every 8-bit model re-measured since, so it is probably conservative.

Qwen3-14B with DeepSeek's official DSpark drafter (there's no z-lab DFlash adapter at this size).

These numbers predate the September 2026 verify kernel. That kernel moved every re-measured 8-bit model's derived cap from 4 to 7 (Qwen3-8B went 2.13× to 2.19×, Gemma-4 12B 2.82× to 3.25×), so this row is probably conservative. Run mlx-dspark benchmark --model mlx-community/Qwen3-14B-8bit --trials 3 for today's number on your Mac.

Numbers not matching your Mac? They shouldn't be identical: the draft cap is derived from your machine's measured curves. Run mlx-dspark benchmark --model mlx-community/Qwen3-14B-8bit --trials 3 to get your own. See how these are measured.