mlx-dspark

Models

Ternary-Bonsai-27B

A 27B reasoning model in ~8.5 GB. Speculation adds ~1.15× on code; use --max-draft auto.

Made by
PrismML
Size
27B
Architecture
Ternary (2-bit pack) hybrid
Features
Thinking

Run it

mlx-dspark serve --model prism-ml/Ternary-Bonsai-27B-mlx-2bit --max-draft auto

Serves an OpenAI + Anthropic API on http://127.0.0.1:8080. For a one-off answer, use mlx-dspark generate --model prism-ml/Ternary-Bonsai-27B-mlx-2bit --prompt "…".

1.07× faster than plain decoding, mean of chat, code and math
26–29 tokens per second (plain: 25.4)
2.60 tokens accepted per round
~12 GB peak memory at chat length
Chat 1.01×
Code 1.13×
Math 1.07×
Dashed: plain decoding (1×). Solid: mlx-dspark on the same prompt, draft cap 2.

Drafter

The drafter downloads with the model the first time you run it, and loads 4-bit quantized (acceptance doesn't depend on the drafter's precision).

Measured 2026-07-16 on an M4 Pro, 48 GB with mlx 0.32.0. This predates the September 2026 verify kernels, which raised every 8-bit model re-measured since, so it is probably conservative.

PrismML's Bonsai 27B is a 1.7-bit ternary rebuild of Qwen3.6-27B: a full 27B-class reasoning model in about 8.5 GB. PrismML publishes its DSpark drafter as GGUF only, and it doesn't accelerate on Macs through their own tooling. mlx-dspark runs it from a 1:1 repack (Rahim/Ternary-Bonsai-27B-dspark, quantized to 4-bit at load).

The honest ceiling#

Extra verify rows on a 2-bit model are compute-bound (they cost the same as on a 4-bit model), while its plain step is very fast. So chat hovers at break-even, and code gains about 1.15×. --max-draft auto is the recommended setting: it picks the cap from your Mac's measured curves plus live acceptance, and can park speculation on content where it would lose.

mlx-dspark generate --model prism-ml/Ternary-Bonsai-27B-mlx-2bit \
  --max-draft auto --prompt "Implement binary search in Python."

The 1-bit pack runs through mlx-vlm ≥ 0.6.5, but speculation is a net loss there (0.71–0.77×), so mlx-dspark refuses it with a pointer to plain generation.

Numbers not matching your Mac? They shouldn't be identical: the draft cap is derived from your machine's measured curves. Run mlx-dspark benchmark --model prism-ml/Ternary-Bonsai-27B-mlx-2bit --trials 3 to get your own. See how these are measured.