Models
MiniCPM5-2B
The best pick for a small Mac: 3.16× mean at ~107–213 tok/s in ~6 GB, and a tool-calling model.
- Made by
- OpenBMB
- Size
- 2B
- Architecture
- Dense transformer (llama)
- Context
- 128k tokens
- Features
- Tool callingThinking
Run it
mlx-dspark serve --model mlx-community/MiniCPM5-2B-bf16
Serves an OpenAI + Anthropic API on http://127.0.0.1:8080. For a one-off answer, use mlx-dspark generate --model mlx-community/MiniCPM5-2B-bf16 --prompt "…".
Drafter
openbmb/MiniCPM5-2B-DSparkresolves automatically with--mode auto(the default).
The drafter downloads with the model the first time you run it, and loads 4-bit quantized (acceptance doesn't depend on the drafter's precision).
At long context
Decode speed with an agent-sized prompt, against plain decoding at the same depth (accepted tokens per round in parentheses). With no --max-draft, the cap shrinks automatically when a narrower verify pays.
| Context | Plain | Cap 3 | Cap 7 |
|---|---|---|---|
| 16k | 42.2 tok/s | 1.62× (2.21) | 1.54× (2.5) |
| 32k | 36.0 tok/s | 2.06× (2.78) | 2.09× (3.51) |
Reading the prompt
Prefill runs at 1640 tok/s, so a 20,000-token prompt takes about 12 s the first time. After that the prefix cache skips it.
Measured 2026-09-27 on an M4 Pro, 48 GB with mlx 0.32.2.
OpenBMB's 2B (a plain dense llama-architecture model with a 128k context, thinking and XML tool calls), paired with OpenBMB's official DSpark head.
The registry points at the bf16 conversion because that is where the ratio is. On MLX 0.32 a bf16 verify is flat out to width 8, so the derived cap is the full block (7). Math runs at 213 tok/s.
Good to know#
- Thinking is on by default. The reasoning is split into
reasoning_contentor Anthropicthinkingblocks, andenable_thinking: falseanswers directly. - Tool calls come out in MiniCPM5's own
<function name="…"><param name="…">form. The server translates them to OpenAI / Anthropic tool calls like every other syntax. - The confidence head doesn't pay here, and lookup drafts are a wash (left on).
- Other quants resolve the same drafter but are unmeasured. A 2B model's bf16 baseline is already ~50 tok/s.
Numbers not matching your Mac? They shouldn't be identical: the draft cap is derived from your machine's measured curves. Run mlx-dspark benchmark --model mlx-community/MiniCPM5-2B-bf16 --trials 3 to get your own. See how these are measured.