Models
Nemotron-3.5-Lightning-30B-A3B
NVIDIA's Mamba-2 + MoE hybrid with its official DSpark head: ~87–112 tok/s, 1.1–1.34× depending on content.
- Made by
- NVIDIA
- Size
- 30B (3B active)
- Architecture
- Mamba-2 + attention, mixture of experts (~3B active)
- Features
- Thinking
Run it
mlx-dspark serve --model mlx-community/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-4bit
Serves an OpenAI + Anthropic API on http://127.0.0.1:8080. For a one-off answer, use mlx-dspark generate --model mlx-community/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-4bit --prompt "…".
Drafter
mlx-community/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-DSpark-bf16resolves automatically with--mode auto(the default).
The drafter downloads with the model the first time you run it, and loads 4-bit quantized (acceptance doesn't depend on the drafter's precision).
Lookup drafts are off by default for this pair (measured as a net loss); --lookup-drafts turns them on.
Measured 2026-08-12 on an M4 Pro, 48 GB with mlx 0.32.0. This predates the September 2026 verify kernels, which raised every 8-bit model re-measured since, so it is probably conservative.
A hybrid of Mamba-2 state-space layers, attention and a mixture of experts (~3B active), with NVIDIA's official DSpark head. Mamba's recurrent state can't be trimmed like a KV cache, so mlx-dspark records each verify round's inputs and rebuilds the state at the exact accept point (bit-exact).
The speedup is unusually content-sensitive. The best per-prompt measurements at cap 4 reach 1.34× on math and 1.27× on code. Suite chat is around break-even, which is why the derived cap is 3. As with every MoE here, the ceiling is the verify cost of routed experts, not the drafter.
Lookup drafts are off by default (a net loss on every MoE).
Numbers not matching your Mac? They shouldn't be identical: the draft cap is derived from your machine's measured curves. Run mlx-dspark benchmark --model mlx-community/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-4bit --trials 3 to get your own. See how these are measured.