Reference
CLI reference
Every subcommand and flag, generated from the command-line parser itself, so it can't drift from the code.
For a guided tour, see Command line. Most flags have measured defaults: when a flag says unset, the engine uses the value calibrated for this Mac and model, or the value measured best for this pair.
mlx-dspark serve#
The API server: OpenAI Chat Completions and Completions, Anthropic Messages, and OpenAI Responses on one port. See Run a server.
--mode MODE- 'auto' (default) picks the target's measured-best speculation (the registry row's stamped mode, e.g. DFlash 2 on Qwen3.8-27B; else dspark -> dflash -> drafter-free lookup), so any repo serves
--model MODEL- target model: an HF repo or local path (e.g. mlx-community/Qwen3-8B-8bit). Matched drafter auto-resolves for known targets; else pass --drafter. See
mlx-dspark models. --no-model- start without loading any model: the server comes up instantly, generation routes answer 503 until a client loads one via POST /admin/load (model pickers still work — /doctor and /admin/models answer model-free)
--drafter DRAFTER- drafter repo/path (overrides auto-resolve)
--max-draft MAX_DRAFT- cap on tokens verified per round; <=0 = full block; 'auto' = adapt the cap per round. Omit to let dspark/dflash DERIVE the cap from this machine+model+quant's measured curves (static_cap, ~5 s once, cached); lookup=6.
--default-max-tokens DEFAULT_MAX_TOKENS- max_tokens used when a request doesn't send one (default 2048)
--max-tokens-cap MAX_TOKENS_CAP- hard ceiling on per-request max_tokens (default 32768 — thinking models routinely need >8k)
--default-temperature DEFAULT_TEMPERATURE- temperature for requests that don't send one (overrides the model's generation_config; note many mlx-community repos ship none, in which case omitted temperature means greedy unless this is set)
--default-top-p DEFAULT_TOP_P- top_p for requests that don't send one (see --default-temperature)
--default-top-k DEFAULT_TOP_K- top_k for requests that don't send one (see --default-temperature)
--confidence-threshold CONFIDENCE_THRESHOLD- DSpark: stop drafting early in a round when the confidence head's cumulative survival estimate drops below this (0 = off). Pays only where the verify curve still rises inside the cap; see Draft caps.
--drafter-bits DRAFTER_BITS- Quantize the drafter to this many bits at load (default 4). Doesn't change acceptance, only drafter speed.
--kv-bits KV_BITS- quantize the target KV cache (4/8 bits; 0 = off). Cuts the KV bandwidth bill on long contexts; mlx-lm text targets only (disables --max-batch)
--max-batch MAX_BATCH- micro-batch up to N concurrently-queued requests through one batched target forward (dense mlx-lm target + dspark/baseline; ~1.5-2.5x aggregate throughput at 2-4 concurrent). 1 = serialized (default)
--context-window CONTEXT_WINDOW- cap the prompt length (default: the target's own trained limit). An over-long request is refused with the message Claude Code recognises as a context limit, so it compacts and retries instead of failing; lower it to keep the KV cache inside your RAM budget.
--host HOST- Interface to bind (default 127.0.0.1). Use 0.0.0.0 with
--api-keyto serve other devices. --port PORT- Port to listen on (default 8080).
--api-key API_KEY- if set, requests must send 'Authorization: Bearer <key>'
--no-thinking- default responses to non-thinking mode (Qwen3 enable_thinking=False); clients can still override per-request
--reasoning-effort REASONING_EFFORT- default reasoning depth for models whose chat template supports it (Qwen3.8-class reasoning_effort; templates that don't know the kwarg ignore it). Clients can override per-request; /health reports whether the loaded model supports it
--no-prefix-cache- disable multi-turn prefix caching (reuse the shared conversation prefix's KV; on by default for dspark/lookup/baseline on dense or under-window sliding-window targets)
--wired-limit- wire MLX's recommended working set (~75%% of RAM) so weights can't be paged out. OFF by default and rarely worth it: wired pages cannot be reclaimed, so on a machine whose working set is already large this can HANG macOS hard enough to need a power cycle (observed on an M4 Pro). Short of that it has corrupted the verify logits on the gemma-4/mlx-vlm route, which can commit wrong tokens rather than crash. It bought no measurable speed where tested (<1%%). Only consider it if the model nearly fills your RAM and you actually see paging stalls - and validate a long run before trusting it.
--prefix-cache-slots PREFIX_CACHE_SLOTS- number of conversations kept in the prefix cache LRU (default 2, so an agent and a chat don't evict each other every turn)
--prefix-cache-rungs N- checkpoint-mode prefix caching (hybrid/recurrent targets, wrapped gemma-4): also snapshot the recurrent state every N prompt tokens, so a request that diverges mid-prompt (new session on the same system prompt, compacted history) partially reuses the cache instead of missing outright (default 8192; 0 disables)
--lookup-drafts, --no-lookup-drafts- hybrid n-gram drafting inside dspark mode. Unset = each loaded pair's measured default (re-resolved on every /admin/load swap); an explicit value pins it for the whole server
--lookup-long-draft LOOKUP_LONG_DRAFT- match-scaled long-draft ceiling for lookup drafts (default 32; set to 6 to disable — see
mlx-dspark generate -h) --wide-gemm-min N- prefill wide-GEMM crossover row count (default: calibrated once and cached; 0 disables — see
mlx-dspark generate -h) --trust-remote-code- allow a checkpoint to ship Python the loader imports (config.json model_file / auto_map). Refused by default: a crafted model repo would run code as you. Also MLX_DSPARK_TRUST_REMOTE_CODE=1
--cpu-split FRAC- prefill CPU co-prefill row fraction (default: calibrated once and cached; 0 disables — see
mlx-dspark generate -h). /health reports the live state; /admin/load takes a per-swapcpu_splitoverride (0 = off) --small-m, --no-small-m- small-M MMA verify kernel (see
mlx-dspark generate --help). Unset = on where the cached probe proves it faster on this machine; --no-small-m forces the stock kernel — the serve-side A/B that previously required downgrading. /health reports the live state; /admin/load takes a per-swapsmall_mboolean override. --sdpa-split, --no-sdpa-split- wide-verify SDPA split (see
mlx-dspark generate --help). Unset = on where a per-chip probe finds mlx's multi-row cliff; --no-sdpa-split forces the single call. /health reports the state. --multirow-attn, --no-multirow-attn- multi-row decode-attention kernel (see
mlx-dspark generate --help). Unset = on for attention shapes a per-chip probe proves faster; --no-multirow-attn uses mlx's kernels. /health reports the state; /admin/load takes a per-swapmultirow_attnboolean. --drafter-window N- DSpark drafter context window (see
mlx-dspark generate --help); 0 = whole context, unset = the pair default. /health reports it; /admin/load takes a per-swapdrafter_windowinteger. --warmup, --no-warmup- on load, run a tiny throwaway generation to compile the Metal kernels and ramp the GPU clock so the FIRST real request is warm instead of eating the ~2 s cold-start. On by default; --no-warmup skips it. /health reports 'warming_up' during the pass and a 'warmup' flag once ready; /admin/load takes a per-swap boolean.
--memory-guard, --no-memory-guard- when macOS reports memory pressure, free the prefix cache's snapshots (WARN: rungs + older slots; CRITICAL: everything) and return MLX's retained buffers to the OS, at the next round boundary — trading a re-prefill for not swapping the model. On by default; /health reports 'memory_guard'; /admin/load takes a per-swap boolean.
--prefix-cache-dir PREFIX_CACHE_DIR- directory for the L2 SSD spill tier (enables spilling the cache to disk)
--prefix-cache-max-ram-mb PREFIX_CACHE_MAX_RAM_MB- spill the prefix cache to --prefix-cache-dir once it exceeds this many MB of RAM (0 = never spill; requires --prefix-cache-dir)
mlx-dspark claude#
Launch Claude Code against a running mlx-dspark serve. Configures the launched process only — other Claude Code sessions are unaffected.
--url URL- base URL of the running server (default http://127.0.0.1:8080)
--api-key API_KEY- the server's --api-key, if it was started with one
--print-env- print the environment as shell exports instead of launching
--print-settings- print a .claude/settings.local.json 'env' block instead of launching (project-scoped: applies to that project only)
Anything after -- is passed straight to claude, e.g. mlx-dspark claude -- --continue.
mlx-dspark generate#
One prompt, one answer, streamed to the terminal. Also what a bare mlx-dspark --prompt … runs.
--mode MODE- auto (default) = this target's measured-best speculation (the registry row's stamped mode, e.g. DFlash 2 on Qwen3.8-27B; else dspark -> dflash -> drafter-free lookup, so ANY repo runs); dspark = DSpark spec decoding; dflash = z-lab DFlash / DFlash 2; lookup = drafter-free prompt-lookup spec decoding (any target); baseline = plain greedy target
--model MODEL- target model: an HF repo or local path (e.g. mlx-community/Qwen3-8B-8bit). The matched drafter auto-resolves for known targets; else pass --drafter.
--drafter DRAFTER- drafter repo/path (overrides auto-resolve)
--prompt PROMPT- The prompt. It goes through the model's chat template unless
--no-chat-templateis set. --wired-limit- wire MLX's recommended working set (~75%% of RAM) so weights can't be paged out. OFF by default and rarely worth it: wired pages cannot be reclaimed, so on a machine whose working set is already large this can HANG macOS hard enough to need a power cycle (observed on an M4 Pro). Short of that it has corrupted the verify logits on the gemma-4/mlx-vlm route, which can commit wrong tokens rather than crash. It bought no measurable speed where tested (<1%%). Only consider it if the model nearly fills your RAM and you actually see paging stalls - and validate a long run before trusting it.
--max-new-tokens MAX_NEW_TOKENS- Maximum tokens to generate.
--max-draft MAX_DRAFT- tokens verified per round (cap). Omit to let dspark/dflash DERIVE the cap from this machine+model+quant's measured curves (static_cap, ~5 s once, cached); 'auto' adapts it per round instead. lookup=6; dflash <=0 = full block.
--temperature TEMPERATURE- 0 = greedy (exact); >0 = speculative sampling (paper setup, lossless wrt target@T)
--top-p TOP_P- nucleus sampling (temperature > 0)
--top-k TOP_K- top-k sampling (temperature > 0)
--seed SEED- Random seed for sampled decoding (
--temperatureabove 0). --confidence-threshold CONFIDENCE_THRESHOLD- DSpark: stop drafting early in a round when the confidence head's cumulative survival estimate drops below this (0 = off). Pays only where the verify curve still rises inside the cap; see Draft caps.
--drafter-bits DRAFTER_BITS- Quantize the drafter to this many bits at load (default 4). Doesn't change acceptance, only drafter speed.
--kv-bits KV_BITS- quantize the target KV cache (4/8 bits; 0 = off). Cuts the KV bandwidth bill on long contexts; mlx-lm text targets only
--lookup-drafts, --no-lookup-drafts- hybrid n-gram drafting inside dspark mode. Unset = the pair's measured default (ON for most targets — free extra speedup on copy-heavy spans; OFF where the registry stamped it a net loss: MoE targets and the 4-bit 27B hybrids). Lossless either way.
--wide-gemm-min N- prefill only: row count from which QuantizedLinear dequantizes the weight once and runs a plain GEMM instead of quantized_matmul (mlx re-dequantizes per output tile, which is redundant at prefill widths). Default: calibrate the crossover for this machine+model once and cache it; 0 disables. Bit-identical, ~1.09x prefill
--trust-remote-code- allow a checkpoint to ship Python the loader imports (config.json model_file / auto_map). Refused by default: a crafted model repo would run code as you. Also MLX_DSPARK_TRUST_REMOTE_CODE=1
--cpu-split FRAC- prefill only: hand this fraction of every wide QuantizedLinear's rows to the CPU stream, concurrently with the GPU (the CPU's matrix units are a second GEMM engine prefill never used). Default: measure the best fraction per width for this machine+model once and cache it; 0 disables. Not bit-identical (fp-tie class, like chunked prefill); 1.41x prefill on an M4 Pro with Qwen3.8-27B-4bit
--small-m, --no-small-m- small-M MMA verify kernel: dequantize each weight group once and reuse it across verify rows 6-8, where stock quantized_matmul re-pays the weight read per row (4-bit) or cliffs at width 6 (8-bit). Unset = on for shapes a one-time cached probe proves faster on this machine (4/8-bit gs64); --no-small-m disables. Output stays greedy-correct (verify-checked); ids can differ from the stock kernel at fp ties, like the batched path
--sdpa-split, --no-sdpa-split- split a wide-verify attention (q_len in mlx's multi-row SDPA cliff, long KV) into <=5-row sub-calls that each stay on the fast path, then concatenate. ~1.5-2x on the long-context verify attention; per-row-equivalent (fp-tie, verify-checked). Unset = on where a one-time per-chip probe finds a cliff; --no-sdpa-split off.
--multirow-attn, --no-multirow-attn- multi-row decode-attention kernel for the speculative shapes (verify widths 2-16, the DSpark drafter's block over its context): packs every GQA x query row onto the GPU matrix units so each KV tile is read once. 1.5-4x on those attention calls, ~1.2x on a 32k-deep 27B verify; fp-tie class, verify-checked. Unset = on for attention shapes a one-time cached probe proves faster (and numerically sound) on this machine; --no-multirow-attn uses mlx's kernels.
--drafter-window N- DSpark drafter context window: the draft block cross-attends only the last N context rows (0 = the whole context). DeepSpec-style heads lose acceptance with depth when they attend everything; a window restores it (DimInfer head at 32k: accept 1.98 -> 2.60). Unset = the pair's default (4096 unless the registry row says otherwise). Drafting-only: output is unchanged.
--lookup-long-draft LOOKUP_LONG_DRAFT- match-scaled long-draft ceiling for lookup drafts (dspark hybrid + lookup mode): a deep context match (real copy run) earns drafts up to this length (default 32 = the measured M-series verify-width plateau); set to the base (6) to disable. Speed-only, output unchanged
--no-chat-template- Send the prompt as raw text, without the model's chat template.
--no-stream- Print the answer once at the end instead of streaming it.
mlx-dspark benchmark#
A warm, reproducible speed sweep on this Mac: plain decoding against the speculative modes, over a chat, a code and a math prompt.
--model MODEL- target repo/path (see
mlx-dspark models) --drafter DRAFTER- Drafter repo or path, overriding auto-resolution.
--modes MODES- comma-separated: dspark, dflash, lookup (baseline always runs)
--caps CAPS- comma-separated caps for dspark/dflash: ints and/or 'auto'
--max-new-tokens MAX_NEW_TOKENS- Maximum tokens to generate.
--trials TRIALS- repeat each prompt N times and report the MEDIAN. Between-trial noise is ~14%% on an M4 Pro (machine state, not content), so a single trial is not quotable — use 3+ for README numbers.
--wired-limit- wire MLX's recommended working set (~75%% of RAM) so weights can't be paged out. OFF by default and rarely worth it: wired pages cannot be reclaimed, so on a machine whose working set is already large this can HANG macOS hard enough to need a power cycle (observed on an M4 Pro). Short of that it has corrupted the verify logits on the gemma-4/mlx-vlm route, which can commit wrong tokens rather than crash. It bought no measurable speed where tested (<1%%). Only consider it if the model nearly fills your RAM and you actually see paging stalls - and validate a long run before trusting it.
--lookup-drafts, --no-lookup-drafts- dspark mode only: the 4-gram HYBRID lookup draft, where an n-gram hit supplies a free draft instead of running the drafter that round. Unset = the pair's measured default (registry rows stamped with lookup off carry it — every MoE and the 4-bit 27B hybrids: a free draft still costs verify rows). Distinct from
--modes lookup, the standalone drafter-free mode. Pass --lookup-drafts / --no-lookup-drafts to force an arm for A/B. --small-m, --no-small-m- small-M MMA verify kernel (see
mlx-dspark generate --help). Unset = on where the cached probe proves it faster; --no-small-m forces the stock kernel for A/B. The setting prints in the header so arms can't be conflated. --sdpa-split, --no-sdpa-split- wide-verify SDPA split (see
mlx-dspark generate --help). Unset = on where the per-chip probe finds a cliff; --no-sdpa-split off. Prints in the header. --multirow-attn, --no-multirow-attn- multi-row decode-attention kernel (see
mlx-dspark generate --help). Unset = on where the per-shape probe proves it faster; --no-multirow-attn off. Prints in the header. --drafter-window N- DSpark drafter context window (see
mlx-dspark generate --help); 0 = whole context, unset = the pair default. --trust-remote-code- allow a checkpoint to ship Python the loader imports (config.json model_file / auto_map). Refused by default: a crafted model repo would run code as you. Also MLX_DSPARK_TRUST_REMOTE_CODE=1
--cpu-split FRAC- prefill CPU co-prefill row fraction (see
mlx-dspark generate -h). Unset = calibrated default; 0 forces it off for an A/B. Prints in the header (it moves the prefill column, not decode). --json JSON- also write results to this JSON file
mlx-dspark models#
List the measured models whose drafters resolve automatically, with their memory.
No options.
mlx-dspark doctor#
Check the environment: Apple Silicon, the MLX stack, memory, and every folder searched for models.