mlx-dspark

Get started

Quickstart

From install to a faster local model in four commands, then three things worth knowing before you stop reading.

This page uses Qwen3-8B (8-bit, about 11 GB), a solid all-rounder. Swap in any model from the models page, or let Choose a model pick one for your Mac.

1. Ask it something#

mlx-dspark generate --model mlx-community/Qwen3-8B-8bit \
  --prompt "Explain how rainbows form."

The first run downloads the model and its matched drafter, then measures your Mac once for this pair (a few seconds). The answer streams, and a summary line reports how many tokens were accepted per round and the speed.

2. See that nothing changed but the speed#

Run the same prompt without speculation:

mlx-dspark generate --model mlx-community/Qwen3-8B-8bit --mode baseline \
  --prompt "Explain how rainbows form." --max-new-tokens 300

Same text, fewer tokens per second. The model verifies every drafted token, so its output is exactly what it would write on its own. mlx-dspark benchmark does this comparison properly, over several prompts and trials: see Command line.

3. Serve it#

mlx-dspark serve --model mlx-community/Qwen3-8B-8bit

One port, three dialects: the OpenAI API at http://127.0.0.1:8080/v1, Anthropic's Messages API at /v1/messages, and OpenAI's Responses API at /v1/responses.

4. Talk to it#

curl http://127.0.0.1:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model": "Qwen3-8B-8bit",
       "messages": [{"role": "user", "content": "Explain rainbows briefly."}]}'
from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8080/v1", api_key="not-needed")
reply = client.chat.completions.create(
    model="Qwen3-8B-8bit",
    messages=[{"role": "user", "content": "Explain rainbows briefly."}],
)
print(reply.choices[0].message.content)
from anthropic import Anthropic

client = Anthropic(base_url="http://127.0.0.1:8080", api_key="not-needed")
msg = client.messages.create(
    model="Qwen3-8B-8bit",
    max_tokens=512,
    messages=[{"role": "user", "content": "Explain rainbows briefly."}],
)
print(msg.content[-1].text)
# leave the server running, then in a second terminal:
mlx-dspark claude

For agent work, start the server with --no-thinking: see Claude Code.

Any app that talks to an OpenAI-compatible server works the same way: point it at http://127.0.0.1:8080/v1 with any API key. See Agents and apps.

Three things worth knowing#

Pick a model, not a configuration. --model takes any Hugging Face repo or local path, exactly like mlx-lm. For the measured models, the matching drafter and that pair's best settings resolve automatically, for any quantization of the model. Anything else still gets drafter-free speculation with the default --mode auto, or name a drafter with --drafter <repo>.

Don't set the draft cap. The published speedups were measured on one M4 Pro. Your Mac's optimum is different, so mlx-dspark measures it on first run and derives its own. If you only remember one flag, --max-draft auto adapts the cap every round while it generates. Pin --max-draft N only if you've measured a better value. Why.

It's lossless by construction. The model checks every drafted token, so output matches plain decoding in every mode and at every cap. Sampling with --temperature stays lossless too: you get an exact sample from the model at that temperature.

Prefer clicking to typing? The Mac app does all of this, including the model picker and a live race that checks the output token by token.