mlx-dspark

Models

Supported models

Every model here is measured and vouched for. Pass it to --model and its drafter, and that pair's best settings, resolve automatically, for any quantization of the model.

These are the pairs we have measured. They're also exactly what --model auto-resolves, so this list and the engine can't disagree: the page is generated from the engine's registry. Numbers are an M4 Pro's. Your Mac derives its own settings on first run.

Show
ModelMean speedupChat / code / mathTokens / sMemoryDefault drafter
Qwen3.8-27B 8-bit 3.74× 3.07 / 3.93 / 4.23 25–35 ~29 GB DFlash 2
LFM2.5-1.2B bf16 3.35× 2.40 / 3.83 / 3.82 239–381 ~4 GB DSpark
Gemma-4 12B 8-bit 3.25× 2.62 / 2.90 / 4.24 48–78 ~15 GB DSpark
MiniCPM5-2B bf16 3.16× 2.12 / 3.15 / 4.21 107–213 ~6 GB DSpark
Nanbeige4.2-3B bf16 2.92× 1.98 / 2.50 / 4.31 34–75 ~10 GB DSpark
LFM2.5-2.6B bf16 2.79× 2.24 / 2.65 / 3.49 98–152 ~7 GB DSpark
Qwen3.8-27B 4-bit 2.72× 2.05 / 3.12 / 3.02 30–46 ~18 GB DFlash 2
Muse-Glimmer-30B 8-bit 2.47× 1.97 / 2.45 / 2.99 16–25 ~40 GB DSpark (community)
Ornith-1.0-9B 8-bit 2.40× 2.21 / 2.53 / 2.48 59–68 ~13 GB DSpark (community)
Qwen3.6-27B 8-bit 2.29× 2.26 / 1.96 / 2.67 16–22 ~32 GB DSpark (community)
Qwen3-8B 8-bit 2.19× 1.71 / 2.03 / 2.83 50–83 ~11 GB DSpark
Qwen3-14B 8-bit 2.03× 1.62 / 2.11 / 2.36 25–36 ~19 GB DSpark
Qwen3-4B 8-bit 1.98× 1.78 / 1.72 / 2.46 89–127 ~8 GB DSpark
Muse-Glimmer-30B 4-bit 1.74× 1.57 / 1.70 / 1.94 22–27 ~26 GB DSpark (community)
Qwen3.6-35B-A3B 4-bit 1.32× 1.05 / 1.24 / 1.67 91–145 ~23 GB DSpark (community)
LFM2.5-8B-A1B bf16 1.26× 1.04 / 1.30 / 1.44 68–94 ~19 GB DSpark
Nemotron-3.5-Lightning-30B-A3B 4-bit 1.10× 0.95 / 1.23 / 1.13 87–112 ~20 GB DSpark (NVIDIA)
Ternary-Bonsai-27B 2-bit (ternary) 1.07× 1.01 / 1.13 / 1.07 26–29 ~12 GB DSpark

Mean speedup is over plain decoding of the same model, averaged across a chat, a code and a math prompt. Tokens / s is the range across those prompts: chat is usually the low end, because open-ended text accepts fewer drafted tokens than code or math. Memory is the measured peak at chat length for the model plus its 4-bit drafter; long contexts add KV cache on top. How these are measured.

Run one#

mlx-dspark serve --model mlx-community/Qwen3.8-27B-4bit

Matching ignores quantization: a -4bit, -8bit or -bf16 build of the same model resolves the same drafter. mlx-dspark models prints this list in your terminal.

Anything else runs too#

This list is what we vouch for, not what works:

Where models come from#

--model takes a Hugging Face repo id or a local path, like mlx-lm. Before downloading anything, the engine looks, in order, at:

  1. the path itself, if it's a folder (--model ~/models/Qwen3.8-27B-4bit always works);
  2. ~/.cache/mlx_dspark/models;
  3. LM Studio's folder (~/.lmstudio/models): its MLX downloads load in place, with no second download;
  4. any folders you list in MLX_DSPARK_MODEL_DIRS (:-separated), such as an external drive, ~/models or a NAS, as publisher/model trees, publisher_model folders or bare model folders;
  5. your Hugging Face cache (wherever HF_HOME puts it), and only then a download.

Only MLX checkpoints load from these folders (a config.json plus .safetensors). mlx-dspark doctor --json lists every folder searched, in order.

Downloads are big

A 27B model at 8-bit is about 29 GB. Each model page lists what you're signing up for, and the Mac app shows download size and fit before it starts.