mlx-dspark

Models

Nanbeige4.2-3B

4.31× on math: a looped 3B whose expensive step is exactly where a drafter pays. Holds 2.2× at 32k context.

Made by
Nanbeige
Size
3B
Architecture
Looped depth: 22 layers run twice
Features
Thinking

Before you download

A fresh download needs --trust-remote-code (the repos carry auto_map pointers; nothing is imported on this route).

Run it

mlx-dspark serve --model MercuriusDream/Nanbeige4.2-3B-mlx-bf16 --trust-remote-code

Serves an OpenAI + Anthropic API on http://127.0.0.1:8080. For a one-off answer, use mlx-dspark generate --model MercuriusDream/Nanbeige4.2-3B-mlx-bf16 --prompt "…".

2.92× faster than plain decoding, mean of chat, code and math
34–75 tokens per second (plain: 17.3)
3.95 tokens accepted per round
~10 GB peak memory at chat length
Chat 1.98×
Code 2.50×
Math 4.31×
Dashed: plain decoding (1×). Solid: mlx-dspark on the same prompt, draft cap 7.

Drafter

The drafter downloads with the model the first time you run it, and loads 4-bit quantized (acceptance doesn't depend on the drafter's precision).

At long context

Decode speed with an agent-sized prompt, against plain decoding at the same depth (accepted tokens per round in parentheses). With no --max-draft, the cap shrinks automatically when a narrower verify pays.

ContextPlainCap 3Cap 7
16k13.9 tok/s 1.97× (2.36) 1.89× (2.78)
32k10.6 tok/s 2.21× (2.74) 1.83× (3.2)

Reading the prompt

Prefill runs at 526 tok/s, so a 20,000-token prompt takes about 38 s the first time. After that the prefix cache skips it.

Measured 2026-09-27 on an M4 Pro, 48 GB with mlx 0.32.2.

Nanbeige4.2-3B runs its 22 layers twice with shared weights. Each decode step therefore reads a ~6B model's worth of weights, so plain decoding is only ~17 tok/s. That expensive step is exactly where a drafter pays: math reaches 74 tok/s.

The drafter is Nanbeige's official DSpark head.

First run#

Both repos carry auto_map pointers in config.json, which the remote-code check flags on a fresh download. Start this pair with --trust-remote-code. Nothing is actually imported: mlx-dspark uses its own model code and keeps the tokenizer's remote code off.

mlx-dspark serve --model MercuriusDream/Nanbeige4.2-3B-mlx-bf16 --trust-remote-code

mlx-lm doesn't ship the Nanbeige module in a release yet, so mlx-dspark vendors mlx-lm's implementation until it does.

Numbers not matching your Mac? They shouldn't be identical: the draft cap is derived from your machine's measured curves. Run mlx-dspark benchmark --model MercuriusDream/Nanbeige4.2-3B-mlx-bf16 --trials 3 to get your own. See how these are measured.