mlx-dspark

Use it

Python

Load a model and its drafter in one call, generate, and read back acceptance and speed. For the full tuned stack, talk to the server instead.

DSpark#

from mlx_dspark import load_pair, speculative_generate

target, tok, drafter, cfg = load_pair("mlx-community/Qwen3-8B-8bit")   # drafter auto-resolved
res = speculative_generate(target, tok, drafter, "Explain how rainbows form.",
                           max_new_tokens=300)
print(res.text)
print(res.mean_accept_len, res.decode_tokens_per_sec)

load_pair takes any repo or local path. For measured models the drafter resolves automatically; otherwise pass drafter="…".

DFlash and DFlash 2#

from mlx_dspark import load_dflash_pair, dflash_generate

target, tok, drafter, cfg = load_dflash_pair("mlx-community/Qwen3.8-27B-4bit")   # DFlash 2
res = dflash_generate(target, tok, drafter, "Write a binary search in Python.")
print(res.text, res.mean_accept_len)

max_draft_tokens=None (the default for dflash_generate) drafts the full block.

Drafter-free lookup and plain decoding#

from mlx_dspark import load_target, lookup_generate, greedy_generate

target, tok = load_target("mlx-community/Qwen3-8B-8bit")
res = lookup_generate(target, tok, "Rewrite this function with type hints: …")
base = greedy_generate(target, tok, "Explain how rainbows form.")

Choosing the draft cap#

The library doesn't measure your Mac for you: speculative_generate defaults to max_draft_tokens=2, and None means the full block. To get the same measured, per-machine behaviour the CLI has, calibrate once (the result is cached on disk) and pass the controller:

from mlx_dspark import calibrate

ctrl = calibrate(target, drafter, mode="dspark",
                 target_repo="mlx-community/Qwen3-8B-8bit",
                 drafter_repo="deepseek-ai/dspark_qwen3_8b_block7")
res = speculative_generate(target, tok, drafter, "…", cap_controller=ctrl)

The controller picks the cap each round from the measured cost curves and live acceptance. That's --max-draft auto on the command line. See Draft caps.

Sampling, stops and streaming#

res = speculative_generate(
    target, tok, drafter, "Write a short poem.",
    temperature=1.0, top_p=0.95, seed=0,          # lossless speculative sampling
    stop=["\n\n"],
    on_text=lambda s: print(s, end="", flush=True),  # stream text as it's committed
)

temperature=0 (the default) is greedy. With a temperature, drafts are accepted by the speculative-sampling rule, so the output is an exact sample from the model at that temperature, not an approximation.

What you get back#

GenResult fields worth knowing:

Field Meaning
text, token_ids The output.
num_tokens, num_rounds, target_forwards Tokens produced, speculation rounds, and passes through the model.
mean_accept_len Tokens committed per round, on average.
tokens_per_sec End to end, including reading the prompt.
decode_tokens_per_sec Generation only, prompt excluded.
prefill_seconds, ttft_seconds Time spent reading the prompt, and time to the first text.
finish_reason "stop" or "length".
accept_lengths Per-round acceptance, for your own analysis.

The library vs the server#

The command line and server switch on extras that the library leaves off by default: the custom verify and attention kernels (each enabled per shape only after a one-time probe proves it faster on your Mac), CPU co-prefill, prefix caching, the warm-up pass and the memory guard. If you want all of that from Python, run mlx-dspark serve and use any OpenAI or Anthropic client. See Run a server.