mlx-dspark

Models

Nemotron-3.5-Lightning-30B-A3B

NVIDIA's Mamba-2 + MoE hybrid with its official DSpark head: ~87–112 tok/s, 1.1–1.34× depending on content.

Made by
NVIDIA
Size
30B (3B active)
Architecture
Mamba-2 + attention, mixture of experts (~3B active)
Features
Thinking

Run it

mlx-dspark serve --model mlx-community/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-4bit

Serves an OpenAI + Anthropic API on http://127.0.0.1:8080. For a one-off answer, use mlx-dspark generate --model mlx-community/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-4bit --prompt "…".

1.10× faster than plain decoding, mean of chat, code and math
87–112 tokens per second (plain: 91.4)
3.28 tokens accepted per round
~20 GB peak memory at chat length
Chat 0.95×
Code 1.23×
Math 1.13×
Dashed: plain decoding (1×). Solid: mlx-dspark on the same prompt, draft cap 3.

Drafter

The drafter downloads with the model the first time you run it, and loads 4-bit quantized (acceptance doesn't depend on the drafter's precision). Lookup drafts are off by default for this pair (measured as a net loss); --lookup-drafts turns them on.

Measured 2026-08-12 on an M4 Pro, 48 GB with mlx 0.32.0. This predates the September 2026 verify kernels, which raised every 8-bit model re-measured since, so it is probably conservative.

A hybrid of Mamba-2 state-space layers, attention and a mixture of experts (~3B active), with NVIDIA's official DSpark head. Mamba's recurrent state can't be trimmed like a KV cache, so mlx-dspark records each verify round's inputs and rebuilds the state at the exact accept point (bit-exact).

The speedup is unusually content-sensitive. The best per-prompt measurements at cap 4 reach 1.34× on math and 1.27× on code. Suite chat is around break-even, which is why the derived cap is 3. As with every MoE here, the ceiling is the verify cost of routed experts, not the drafter.

Lookup drafts are off by default (a net loss on every MoE).

Numbers not matching your Mac? They shouldn't be identical: the draft cap is derived from your machine's measured curves. Run mlx-dspark benchmark --model mlx-community/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-4bit --trials 3 to get your own. See how these are measured.