mlx-dspark

Reference

HTTP API reference

Every route the server answers, the control plane the Mac app is built on, and the fields it adds to responses.

The server listens on http://127.0.0.1:8080 by default (--host, --port). With --api-key, every route except GET /health requires Authorization: Bearer <key> or x-api-key: <key>, and a wrong key gets 401.

Every route also answers without the /v1 prefix (/chat/completions, /messages, …).

Generation#

Route Dialect Notes
POST /v1/chat/completions OpenAI Chat Completions Streaming and not; n; tools; reasoning_content split; stream_options.include_usage.
POST /v1/completions OpenAI Completions Raw prompt, no chat template.
POST /v1/messages Anthropic Messages Streaming and not; tools; thinking blocks; stop_sequences.
POST /v1/messages/count_tokens Anthropic Token count for a would-be request.
POST /v1/responses OpenAI Responses Streaming and not; stateless (no previous_response_id).
GET /v1/models OpenAI / Anthropic The one loaded model, with a display_name.

Generation routes answer 503 while no model is loaded (--no-model, or during a swap). A prompt longer than the context window is refused with "prompt is too long", the wording agent clients recognize and compact on. Unknown request fields are ignored, never rejected.

Accepted request parameters: temperature, top_p, top_k, max_tokens, stop / stop_sequences, seed, presence_penalty, frequency_penalty, logprobs, top_logprobs, tools, enable_thinking, thinking, reasoning_effort. Unset sampling values fall back to --default-temperature / --default-top-p / --default-top-k, then to the model's own generation_config. max_tokens defaults to --default-max-tokens (2048) and is capped at --max-tokens-cap (32768).

The x_mlx_dspark block#

Every response carries this non-standard block. Clients ignore it; it's how you see the speedup.

Field Meaning
mode dspark, dflash, lookup or baseline.
cap Draft cap used for this request, after any depth adjustment.
accept_len Tokens committed per round.
tokens_per_sec End to end, including prompt reading.
decode_tokens_per_sec Generation only.
target_forwards Passes through the model.
lookup_rounds Rounds whose draft came free from lookup (when any).
prompt_tokens, cached_tokens, completion_tokens, context_tokens Token accounting; cached_tokens is what the prefix cache served.
prefill_seconds, decode_seconds, ttft_seconds Where the time went.
prefill_tokens_per_sec Prompt-reading speed (when at least 16 new tokens were read).
ceiling_tokens_per_sec, roofline_ratio The plain-decoding ceiling this Mac's measured bandwidth allows at this depth, and decode speed as a multiple of it.
decay_ratio, swap_delta_bytes, cold Diagnostics: late-vs-early speed within the request, swap growth during it, and whether it was the first request after a load without warm-up.

The OpenAI endpoints also fill the standard usage.prompt_tokens_details.cached_tokens.

Status and telemetry#

Route Returns
GET /health Always open, even with --api-key and mid-swap. status is ok, loading or no_model. When ready: model, mode, target and drafter, max_draft, lookup_drafts, confidence_threshold, context_window, kv_bits, the kernel and prefill switches (small_m, multirow_attn, sdpa_split, cpu_split, drafter_window), warmup, memory_guard, max_output_tokens, supports_reasoning_effort, thinking_default, reasoning_effort, and a warnings list (memory pressure, context-window RAM). While loading: phase (loading or warming_up) and download progress.
GET /metrics Aggregate counters, allocator memory, macOS memory pressure and swap, the last roofline verdict, memory-guard state.
GET /machine Chip, measured memory bandwidth, the loaded model's bytes per token and single-stream ceiling, and a verdict. Answers without a model too.
GET /calibration This Mac's measured cost curves for the loaded pair.
GET /events Server-sent events: one event per speculation round (drafted, accepted, committed, cap, timing) and prefill progress events.
GET /rounds?limit=N Recent rounds and per-position acceptance stats, for clients that would rather poll.
GET /doctor The same report as mlx-dspark doctor --json. Answers without a model.

/events#

Round events have no type key: {"drafted": 7, "accepted": 5, "committed": 6, "cap": 7, "source": "drafter", …}, where source is drafter, lookup (a free lookup draft) or plain (a parked or plain step). Prefill events do: {"type": "prefill", "req": …, "mode": …, "processed": 4096, "total": 20000}, with positions counted from the start of the prompt (so they begin at the cached length), and a final {"type": "prefill", …, "done": true} on every exit path, so a progress display always clears. The stream ends when the model is swapped.

Control plane#

These are what the Mac app drives. Each is a plain HTTP endpoint first, so a script can do anything the app can.

Route What it does
POST /admin/load Load or swap the model in place; the port stays the same, so connected clients keep working. Body below. A first-time load reports download progress in /health and resumes partial downloads.
POST /admin/load/cancel Cancel a download in progress.
POST /admin/unload Free the model; the server stays up.
GET /admin/status Loading state and errors.
GET /admin/models Registry models, what's on disk (Hugging Face cache, LM Studio, your model folders), disk usage, and this Mac's bandwidth relative to the M4 Pro the published numbers came from. Answers without a model.
GET /admin/integrations Ready-to-paste configs for Claude Code, Codex, OpenCode, pi and any OpenAI-compatible app, filled in with this server's real address, model and key.
POST /admin/race Run several decoders on one prompt and stream the result as SSE, ending with a token-by-token identical-output verdict. Body: {"prompt": "…", "arms": [{"mode": "dspark", "cap": 4}, {"mode": "baseline"}], "max_tokens": 200}.

POST /admin/load body#

Only model is required. Every other key is an override for this load; omit it to keep the pair's measured default or the server's setting.

Key Type Meaning
model string Repo id or local path.
mode string auto, dspark, dflash, lookup or baseline.
max_draft int or "auto" Pin the cap, or adapt per round.
lookup_drafts bool Hybrid lookup drafts on or off.
confidence_threshold number, 0–1 0 disables.
context_window int Cap the prompt length (a RAM lever).
kv_bits 0, 4 or 8 Quantized KV cache; 0 = full precision.
drafter_window int ≥ 0 DSpark drafter context window; 0 = whole context.
small_m, multirow_attn, sdpa_split bool Kernel switches, for A/B runs.
cpu_split number, 0 ≤ x < 1 CPU co-prefill fraction; 0 = off.
warmup bool Run the warm-up generation after loading.
memory_guard bool The memory-pressure guard.
enable_thinking bool The thinking default for clients that don't say (sticky across swaps).
reasoning_effort string low, medium, high or xhigh: the default depth (sticky across swaps).

There's deliberately no per-request way to allow remote code: that's --trust-remote-code for the whole process.

curl -X POST http://127.0.0.1:8080/admin/load \
  -H "Content-Type: application/json" \
  -d '{"model": "mlx-community/Qwen3.8-27B-4bit", "context_window": 65536, "enable_thinking": false}'