Models
Qwen3.8-27B
The project's best ratio: a 27B reasoning model at 3.7× on 8-bit, and the fastest 27B-class decode (~40 tok/s) on 4-bit, with Inco AI's DFlash 2 drafter.
- Made by
- Alibaba Qwen
- Size
- 27B
- Architecture
- Hybrid: 48 linear-attention + 16 full-attention layers
- Context
- 256k tokens
- Features
- Tool callingThinking
Run it (4-bit)
mlx-dspark serve --model mlx-community/Qwen3.8-27B-4bit
Serves an OpenAI + Anthropic API on http://127.0.0.1:8080. For a one-off answer, use mlx-dspark generate --model mlx-community/Qwen3.8-27B-4bit --prompt "…".
Two drafters, measured side by side
| Drafter | Mode | Mean | Chat | Code | Math | Accept | tok/s |
|---|---|---|---|---|---|---|---|
| DFlash 2 default | --mode dflash |
2.72× | 2.05× | 3.12× | 3.02× | 5.09 | 40.0 |
| DSpark (Red Hat) | --mode dspark |
2.63× | 2.25× | 2.81× | 2.83× | 4.83 | 38.6 |
Drafter
incoai/Qwen3.8-27B-DFlash2resolves automatically with--mode auto(the default).RedHatAI/Qwen3.8-27B-speculator.dsparkavailable with--mode dspark.
The drafter downloads with the model the first time you run it, and loads 4-bit quantized (acceptance doesn't depend on the drafter's precision).
Lookup drafts are off by default for this pair (measured as a net loss); --lookup-drafts turns them on.
At long context
Decode speed with an agent-sized prompt, against plain decoding at the same depth (accepted tokens per round in parentheses). With no --max-draft, the cap shrinks automatically when a narrower verify pays.
| Drafter | Context | Plain | Cap 3 | Cap 7 |
|---|---|---|---|---|
| DFlash 2 | 16k | 14.1 tok/s | 1.69× (2.74) | 1.79× (3.16) |
| DFlash 2 | 32k | 12.9 tok/s | 1.66× (2.58) | 1.85× (3.24) |
| DSpark (Red Hat) | 16k | 14.2 tok/s | 1.89× (3.0) | 1.89× (3.38) |
| DSpark (Red Hat) | 32k | 13.1 tok/s | 1.76× (2.74) | 1.83× (3.32) |
Reading the prompt
Prefill runs at 184 tok/s with CPU co-prefill (130 without), so a 20,000-token prompt takes about 109 s the first time. After that the prefix cache skips it.
Measured 2026-09-28 on an M4 Pro, 48 GB with mlx 0.32.2, paired in one process (arms interleaved per prompt), High Power mode.
Run it (8-bit)
mlx-dspark serve --model mlx-community/Qwen3.8-27B-8bit
Serves an OpenAI + Anthropic API on http://127.0.0.1:8080. For a one-off answer, use mlx-dspark generate --model mlx-community/Qwen3.8-27B-8bit --prompt "…".
Two drafters, measured side by side
| Drafter | Mode | Mean | Chat | Code | Math | Accept | tok/s |
|---|---|---|---|---|---|---|---|
| DFlash 2 default | --mode dflash |
3.74× | 3.07× | 3.93× | 4.23× | 5.42 | 30.8 |
| DSpark (Red Hat) | --mode dspark |
3.79× | 3.03× | 4.01× | 4.32× | 5.55 | 31.1 |
Drafter
incoai/Qwen3.8-27B-DFlash2resolves automatically with--mode auto(the default).RedHatAI/Qwen3.8-27B-speculator.dsparkavailable with--mode dspark.
The drafter downloads with the model the first time you run it, and loads 4-bit quantized (acceptance doesn't depend on the drafter's precision).
Lookup drafts are off by default for this pair (measured as a net loss); --lookup-drafts turns them on.
Reading the prompt
Prefill runs at 182 tok/s with CPU co-prefill (138 without), so a 20,000-token prompt takes about 110 s the first time. After that the prefix cache skips it.
Measured 2026-09-28 on an M4 Pro, 48 GB with mlx 0.32.2, paired in one process (arms interleaved per prompt), High Power mode.
Qwen3.8-27B is a 27B reasoning model with a hybrid backbone: 48 of its 64 layers are linear attention, which keeps a fixed-size recurrent state instead of a growing KV cache. That is why a "256k-context" 27B is usable at long context on a Mac at all.
Two drafters, one default#
--mode auto (the default, and what the Mac app uses) picks DFlash 2 from Inco AI for both quants. DFlash 2 adds a candidate-path selector and dynamic convolutions to the DFlash backbone. Acceptance goes up by about 1.1 to 1.4 tokens per round without widening the verify, and that is the kind of gain that turns into real speed on Apple Silicon.
Red Hat's DSpark head (--mode dspark) is a first-class alternative, not a fallback. It ties DFlash 2 at 8-bit (3.79× vs 3.74×) and holds acceptance at long context with its own trained 2048-token window.
Which quant#
- 8-bit (~29 GB) has the best ratio in the project: 3.74× mean, 4.23× on math.
- 4-bit (~18 GB) has the highest absolute speed: 40 tok/s, the fastest decode of any 27B-class target here.
Long prompts#
The KV cache costs 0.086 GB per 1k tokens (identical for both quants, because the cache is bf16 regardless of weight bits). That comes to about 11 GB on top of the weights at 128k, and about 23 GB at the full 256k. Cap it with --context-window if RAM is the limit; requests past the cap get the "prompt is too long" error that agent clients compact on.
The template supports reasoning_effort (low, medium, xhigh). Set a server default with --reasoning-effort, or send it per request.
Lookup drafts are off by default for this pair (measured as a net loss on the 4-bit verify curve). Every divergence from plain greedy decoding is a floating-point tie, on both quants.
Numbers not matching your Mac? They shouldn't be identical: the draft cap is derived from your machine's measured curves. Run mlx-dspark benchmark --model mlx-community/Qwen3.8-27B-4bit --trials 3 to get your own. See how these are measured.