Models
LFM2.5-1.2B
The raw-speed record: 333 tok/s mean (up to ~385) at 3.35×, in ~4 GB.
- Made by
- Liquid AI
- Size
- 1.2B
- Architecture
- Short-convolution + attention hybrid
- Features
- Tool callingNo thinking by default
Run it
mlx-dspark serve --model LiquidAI/LFM2.5-1.2B-Instruct-MLX-bf16
Serves an OpenAI + Anthropic API on http://127.0.0.1:8080. For a one-off answer, use mlx-dspark generate --model LiquidAI/LFM2.5-1.2B-Instruct-MLX-bf16 --prompt "…".
Drafter
LiquidAI/LFM2.5-1.2B-Instruct-DSparkresolves automatically with--mode auto(the default).
The drafter downloads with the model the first time you run it, and loads 4-bit quantized (acceptance doesn't depend on the drafter's precision).
Reading the prompt
Prefill runs at 3240 tok/s, so a 20,000-token prompt takes about 6 s the first time. After that the prefix cache skips it.
Measured 2026-09-27 on an M4 Pro, 48 GB with mlx 0.32.2.
Liquid AI's LFM2.5 is a short-convolution + attention hybrid: most layers are a tiny causal convolution with a 2-row state instead of attention. The small model is both easy to draft and cheap to verify, which is why it reaches 5.28 accepted tokens per round.
The drafter is Liquid AI's official DSpark head. It reuses the target's tied embedding and output head, so it ships only the backbone.
Good to know#
- bf16 is the sweet spot. MLX 0.32.1's wide-matrix kernels make wide verifies cheap for unquantized weights.
- Tool calls come out in LFM2's
<|tool_call_start|>[func(arg="v")]<|tool_call_end|>form and are parsed into standard tool calls. - Prefill runs at ~3240 tok/s, so a 20k-token prompt takes about 6 s.
Numbers not matching your Mac? They shouldn't be identical: the draft cap is derived from your machine's measured curves. Run mlx-dspark benchmark --model LiquidAI/LFM2.5-1.2B-Instruct-MLX-bf16 --trials 3 to get your own. See how these are measured.