mlx-dspark

Models

LFM2.5-8B-A1B

Supported and lossless, but a modest 1.26× that only pays at bf16. Plain 8-bit decoding is faster in absolute terms.

Made by
Liquid AI
Size
8B (1B active)
Architecture
Conv hybrid, mixture of experts (~1B active)
Features
Tool callingNo thinking by default

Before you download

Plain 8-bit decoding (~114 tok/s) is faster than bf16 with speculation (~82 tok/s). The drafter is for bf16-quality users.

Run it

mlx-dspark serve --model LiquidAI/LFM2.5-8B-A1B-MLX-bf16

Serves an OpenAI + Anthropic API on http://127.0.0.1:8080. For a one-off answer, use mlx-dspark generate --model LiquidAI/LFM2.5-8B-A1B-MLX-bf16 --prompt "…".

1.26× faster than plain decoding, mean of chat, code and math
68–94 tokens per second (plain: 65.0)
— tokens accepted per round
~19 GB peak memory at chat length
Chat 1.04×
Code 1.30×
Math 1.44×
Dashed: plain decoding (1×). Solid: mlx-dspark on the same prompt, draft cap 4.

Drafter

The drafter downloads with the model the first time you run it, and loads 4-bit quantized (acceptance doesn't depend on the drafter's precision). Lookup drafts are off by default for this pair (measured as a net loss); --lookup-drafts turns them on.

Reading the prompt

Prefill runs at 2170 tok/s, so a 20,000-token prompt takes about 9 s the first time. After that the prefix cache skips it.

Measured 2026-08-22 on an M4 Pro, 48 GB with mlx 0.32.1. This predates the September 2026 verify kernels, which raised every 8-bit model re-measured since, so it is probably conservative.

A mixture-of-experts LFM2.5 with ~1B active parameters per token. It runs with zero extra model code and is lossless, but speculation has little to win: a ~1B-active step is already very cheap, and every extra verify row pulls in a fresh set of experts.

  • At bf16 it's a modest win: 1.26× (1.44× on math, peaking at 1.67× on a single math prompt).
  • At 8-bit it's a net loss (0.90–0.97× at every cap).
  • Absolute speed: plain 8-bit decoding (~114 tok/s) beats bf16 with speculation (~82 tok/s). If you want the fastest way to run this model, run the 8-bit plainly with --mode baseline.

This matches Liquid AI's own card (M4 Max, bf16 mean 1.18×). The per-token acceptance agrees too (~69%), so the ceiling is the MoE's verify cost, not the drafter.

Numbers not matching your Mac? They shouldn't be identical: the draft cap is derived from your machine's measured curves. Run mlx-dspark benchmark --model LiquidAI/LFM2.5-8B-A1B-MLX-bf16 --trials 3 to get your own. See how these are measured.