mlx-dspark

Models

Gemma-4 12B

A big ratio and real speed: 3.25× mean and ~48–78 tok/s, with DeepSeek's official DSpark drafter.

Made by
Google
Size
12B
Architecture
Dense, sliding-window + global attention
Features
Tool callingNo thinking by default

Run it

mlx-dspark serve --model mlx-community/gemma-4-12B-it-8bit

Serves an OpenAI + Anthropic API on http://127.0.0.1:8080. For a one-off answer, use mlx-dspark generate --model mlx-community/gemma-4-12B-it-8bit --prompt "…".

3.25× faster than plain decoding, mean of chat, code and math
48–78 tokens per second (plain: 18.4)
4.98 tokens accepted per round
~15 GB peak memory at chat length
Chat 2.62×
Code 2.90×
Math 4.24×
Dashed: plain decoding (1×). Solid: mlx-dspark on the same prompt, draft cap 7.

Drafter

The drafter downloads with the model the first time you run it, and loads 4-bit quantized (acceptance doesn't depend on the drafter's precision).

Reading the prompt

Prefill runs at 381 tok/s with CPU co-prefill (293 without), so a 20,000-token prompt takes about 52 s the first time. After that the prefix cache skips it.

Measured 2026-09-27 on an M4 Pro, 48 GB with mlx 0.32.2.

Google's Gemma-4 12B instruct model, loaded through mlx-vlm, with DeepSeek's official DSpark drafter. The 2026-09 verify kernel moved its derived cap from 4 to 7, which took the mean from 2.82× to 3.25× (math 3.12× to 4.24×).

Copy-heavy editing goes further#

When the model re-emits or refactors code that is already in its context, lookup drafts reach 4.51× on file re-emission (75 tok/s) and 4.33× on a rename refactor, with output still bit-identical.

Good to know#

  • Gemma-4 does not think by default, so --no-thinking is a no-op here.
  • It uses a rotating (sliding-window) KV cache. Prefix caching is exact until the window first wraps, then switches to checkpoints automatically.
  • z-lab's DFlash drafter also loads (--mode dflash). On current MLX it measures below DSpark on this model, so treat it as the head-to-head option.
  • A 4-bit build of the target (…-it-4bit) resolves the same drafter. It gives a smaller ratio (~1.45× on older kernels) and more raw tokens per second.
  • Gemma-4's tool-call syntax and its <|tool_response> turn marker are handled natively. Claude Code works well with it, but it doesn't converge on pi's tool protocol.

Numbers not matching your Mac? They shouldn't be identical: the draft cap is derived from your machine's measured curves. Run mlx-dspark benchmark --model mlx-community/gemma-4-12B-it-8bit --trials 3 to get your own. See how these are measured.