mlx-dspark

Models

Qwen3.8-27B

The project's best ratio: a 27B reasoning model at 3.7× on 8-bit, and the fastest 27B-class decode (~40 tok/s) on 4-bit, with Inco AI's DFlash 2 drafter.

Made by
Alibaba Qwen
Size
27B
Architecture
Hybrid: 48 linear-attention + 16 full-attention layers
Context
256k tokens
Features
Tool callingThinking

Run it (4-bit)

mlx-dspark serve --model mlx-community/Qwen3.8-27B-4bit

Serves an OpenAI + Anthropic API on http://127.0.0.1:8080. For a one-off answer, use mlx-dspark generate --model mlx-community/Qwen3.8-27B-4bit --prompt "…".

2.72× faster than plain decoding, mean of chat, code and math
30–46 tokens per second (plain: 14.7)
5.09 tokens accepted per round
~18 GB peak memory at chat length
Chat 2.05×
Code 3.12×
Math 3.02×
Dashed: plain decoding (1×). Solid: mlx-dspark on the same prompt, draft cap 7.

Two drafters, measured side by side

DrafterModeMeanChatCodeMathAccepttok/s
DFlash 2 default --mode dflash 2.72× 2.05× 3.12× 3.02× 5.09 40.0
DSpark (Red Hat) --mode dspark 2.63× 2.25× 2.81× 2.83× 4.83 38.6

Drafter

The drafter downloads with the model the first time you run it, and loads 4-bit quantized (acceptance doesn't depend on the drafter's precision). Lookup drafts are off by default for this pair (measured as a net loss); --lookup-drafts turns them on.

At long context

Decode speed with an agent-sized prompt, against plain decoding at the same depth (accepted tokens per round in parentheses). With no --max-draft, the cap shrinks automatically when a narrower verify pays.

DrafterContextPlainCap 3Cap 7
DFlash 216k14.1 tok/s 1.69× (2.74) 1.79× (3.16)
DFlash 232k12.9 tok/s 1.66× (2.58) 1.85× (3.24)
DSpark (Red Hat)16k14.2 tok/s 1.89× (3.0) 1.89× (3.38)
DSpark (Red Hat)32k13.1 tok/s 1.76× (2.74) 1.83× (3.32)

Reading the prompt

Prefill runs at 184 tok/s with CPU co-prefill (130 without), so a 20,000-token prompt takes about 109 s the first time. After that the prefix cache skips it.

Measured 2026-09-28 on an M4 Pro, 48 GB with mlx 0.32.2, paired in one process (arms interleaved per prompt), High Power mode.

Qwen3.8-27B is a 27B reasoning model with a hybrid backbone: 48 of its 64 layers are linear attention, which keeps a fixed-size recurrent state instead of a growing KV cache. That is why a "256k-context" 27B is usable at long context on a Mac at all.

Two drafters, one default#

--mode auto (the default, and what the Mac app uses) picks DFlash 2 from Inco AI for both quants. DFlash 2 adds a candidate-path selector and dynamic convolutions to the DFlash backbone. Acceptance goes up by about 1.1 to 1.4 tokens per round without widening the verify, and that is the kind of gain that turns into real speed on Apple Silicon.

Red Hat's DSpark head (--mode dspark) is a first-class alternative, not a fallback. It ties DFlash 2 at 8-bit (3.79× vs 3.74×) and holds acceptance at long context with its own trained 2048-token window.

Which quant#

  • 8-bit (~29 GB) has the best ratio in the project: 3.74× mean, 4.23× on math.
  • 4-bit (~18 GB) has the highest absolute speed: 40 tok/s, the fastest decode of any 27B-class target here.

Long prompts#

The KV cache costs 0.086 GB per 1k tokens (identical for both quants, because the cache is bf16 regardless of weight bits). That comes to about 11 GB on top of the weights at 128k, and about 23 GB at the full 256k. Cap it with --context-window if RAM is the limit; requests past the cap get the "prompt is too long" error that agent clients compact on.

The template supports reasoning_effort (low, medium, xhigh). Set a server default with --reasoning-effort, or send it per request.

Lookup drafts are off by default for this pair (measured as a net loss on the 4-bit verify curve). Every divergence from plain greedy decoding is a floating-point tie, on both quants.

Numbers not matching your Mac? They shouldn't be identical: the draft cap is derived from your machine's measured curves. Run mlx-dspark benchmark --model mlx-community/Qwen3.8-27B-4bit --trials 3 to get your own. See how these are measured.