mlx-dspark

Models

LFM2.5-2.6B

A small reasoning model at 2.79× mean and ~97–152 tok/s in ~7 GB.

Made by
Liquid AI
Size
2.6B
Architecture
Short-convolution + attention hybrid
Features
Tool callingThinking

Run it

mlx-dspark serve --model LiquidAI/LFM2.5-2.6B-MLX-bf16

Serves an OpenAI + Anthropic API on http://127.0.0.1:8080. For a one-off answer, use mlx-dspark generate --model LiquidAI/LFM2.5-2.6B-MLX-bf16 --prompt "…".

2.79× faster than plain decoding, mean of chat, code and math
98–152 tokens per second (plain: 43.6)
4.00 tokens accepted per round
~7 GB peak memory at chat length
Chat 2.24×
Code 2.65×
Math 3.49×
Dashed: plain decoding (1×). Solid: mlx-dspark on the same prompt, draft cap 7.

Drafter

The drafter downloads with the model the first time you run it, and loads 4-bit quantized (acceptance doesn't depend on the drafter's precision).

Reading the prompt

Prefill runs at 1420 tok/s, so a 20,000-token prompt takes about 14 s the first time. After that the prefix cache skips it.

Measured 2026-09-27 on an M4 Pro, 48 GB with mlx 0.32.2.

The 2.6B is a pure reasoning model: its chat template always opens a <think> block. Turning thinking off (--no-thinking, enable_thinking: false, or the app toggle) closes the block for it, so it answers directly.

Same family and drafter packaging as LFM2.5-1.2B: a short-convolution + attention hybrid, with Liquid AI's official DSpark head reusing the target's embedding and output head.

Numbers not matching your Mac? They shouldn't be identical: the draft cap is derived from your machine's measured curves. Run mlx-dspark benchmark --model LiquidAI/LFM2.5-2.6B-MLX-bf16 --trials 3 to get your own. See how these are measured.