mlx-dspark

Models

MiniCPM5-2B

The best pick for a small Mac: 3.16× mean at ~107–213 tok/s in ~6 GB, and a tool-calling model.

Made by
OpenBMB
Size
2B
Architecture
Dense transformer (llama)
Context
128k tokens
Features
Tool callingThinking

Run it

mlx-dspark serve --model mlx-community/MiniCPM5-2B-bf16

Serves an OpenAI + Anthropic API on http://127.0.0.1:8080. For a one-off answer, use mlx-dspark generate --model mlx-community/MiniCPM5-2B-bf16 --prompt "…".

3.16× faster than plain decoding, mean of chat, code and math
107–213 tokens per second (plain: 50.6)
4.87 tokens accepted per round
~6 GB peak memory at chat length
Chat 2.12×
Code 3.15×
Math 4.21×
Dashed: plain decoding (1×). Solid: mlx-dspark on the same prompt, draft cap 7.

Drafter

The drafter downloads with the model the first time you run it, and loads 4-bit quantized (acceptance doesn't depend on the drafter's precision).

At long context

Decode speed with an agent-sized prompt, against plain decoding at the same depth (accepted tokens per round in parentheses). With no --max-draft, the cap shrinks automatically when a narrower verify pays.

ContextPlainCap 3Cap 7
16k42.2 tok/s 1.62× (2.21) 1.54× (2.5)
32k36.0 tok/s 2.06× (2.78) 2.09× (3.51)

Reading the prompt

Prefill runs at 1640 tok/s, so a 20,000-token prompt takes about 12 s the first time. After that the prefix cache skips it.

Measured 2026-09-27 on an M4 Pro, 48 GB with mlx 0.32.2.

OpenBMB's 2B (a plain dense llama-architecture model with a 128k context, thinking and XML tool calls), paired with OpenBMB's official DSpark head.

The registry points at the bf16 conversion because that is where the ratio is. On MLX 0.32 a bf16 verify is flat out to width 8, so the derived cap is the full block (7). Math runs at 213 tok/s.

Good to know#

  • Thinking is on by default. The reasoning is split into reasoning_content or Anthropic thinking blocks, and enable_thinking: false answers directly.
  • Tool calls come out in MiniCPM5's own <function name="…"><param name="…"> form. The server translates them to OpenAI / Anthropic tool calls like every other syntax.
  • The confidence head doesn't pay here, and lookup drafts are a wash (left on).
  • Other quants resolve the same drafter but are unmeasured. A 2B model's bf16 baseline is already ~50 tok/s.

Numbers not matching your Mac? They shouldn't be identical: the draft cap is derived from your machine's measured curves. Run mlx-dspark benchmark --model mlx-community/MiniCPM5-2B-bf16 --trials 3 to get your own. See how these are measured.