Models
Nanbeige4.2-3B
4.31× on math: a looped 3B whose expensive step is exactly where a drafter pays. Holds 2.2× at 32k context.
- Made by
- Nanbeige
- Size
- 3B
- Architecture
- Looped depth: 22 layers run twice
- Features
- Thinking
Before you download
A fresh download needs --trust-remote-code (the repos carry auto_map pointers; nothing is imported on this route).
Run it
mlx-dspark serve --model MercuriusDream/Nanbeige4.2-3B-mlx-bf16 --trust-remote-code
Serves an OpenAI + Anthropic API on http://127.0.0.1:8080. For a one-off answer, use mlx-dspark generate --model MercuriusDream/Nanbeige4.2-3B-mlx-bf16 --prompt "…".
Drafter
Nanbeige/Nanbeige4.2-3B-DSparkresolves automatically with--mode auto(the default).
The drafter downloads with the model the first time you run it, and loads 4-bit quantized (acceptance doesn't depend on the drafter's precision).
At long context
Decode speed with an agent-sized prompt, against plain decoding at the same depth (accepted tokens per round in parentheses). With no --max-draft, the cap shrinks automatically when a narrower verify pays.
| Context | Plain | Cap 3 | Cap 7 |
|---|---|---|---|
| 16k | 13.9 tok/s | 1.97× (2.36) | 1.89× (2.78) |
| 32k | 10.6 tok/s | 2.21× (2.74) | 1.83× (3.2) |
Reading the prompt
Prefill runs at 526 tok/s, so a 20,000-token prompt takes about 38 s the first time. After that the prefix cache skips it.
Measured 2026-09-27 on an M4 Pro, 48 GB with mlx 0.32.2.
Nanbeige4.2-3B runs its 22 layers twice with shared weights. Each decode step therefore reads a ~6B model's worth of weights, so plain decoding is only ~17 tok/s. That expensive step is exactly where a drafter pays: math reaches 74 tok/s.
The drafter is Nanbeige's official DSpark head.
First run#
Both repos carry auto_map pointers in config.json, which the remote-code check flags on a fresh download. Start this pair with --trust-remote-code. Nothing is actually imported: mlx-dspark uses its own model code and keeps the tokenizer's remote code off.
mlx-dspark serve --model MercuriusDream/Nanbeige4.2-3B-mlx-bf16 --trust-remote-code
mlx-lm doesn't ship the Nanbeige module in a release yet, so mlx-dspark vendors mlx-lm's implementation until it does.
Numbers not matching your Mac? They shouldn't be identical: the draft cap is derived from your machine's measured curves. Run mlx-dspark benchmark --model MercuriusDream/Nanbeige4.2-3B-mlx-bf16 --trials 3 to get your own. See how these are measured.