mlx-dspark

Models

Ornith-1.0-9B

The mid-size sweet spot for agentic coding: 2.40× mean at ~59–68 tok/s, and the first target here above 2× on chat.

Made by
DeepReinforce
Size
9B
Architecture
Hybrid: linear + full attention
Features
Tool callingThinking

Run it

mlx-dspark serve --model mlx-community/Ornith-1.0-9B-8bit

Serves an OpenAI + Anthropic API on http://127.0.0.1:8080. For a one-off answer, use mlx-dspark generate --model mlx-community/Ornith-1.0-9B-8bit --prompt "…".

2.40× faster than plain decoding, mean of chat, code and math
59–68 tokens per second (plain: 26.7)
3.64 tokens accepted per round
~13 GB peak memory at chat length
Chat 2.21×
Code 2.53×
Math 2.48×
Dashed: plain decoding (1×). Solid: mlx-dspark on the same prompt, draft cap 4.

Drafter

The drafter downloads with the model the first time you run it, and loads 4-bit quantized (acceptance doesn't depend on the drafter's precision).

Measured 2026-07-18 on an M4 Pro, 48 GB with mlx 0.32.0. This predates the September 2026 verify kernels, which raised every 8-bit model re-measured since, so it is probably conservative.

Ornith-1.0-9B is an agentic-coding model on the Qwen3.5-style hybrid backbone. Its community drafter by stanleyphoong was rigorously qualified by its author (17 of 17 gates, 95% of the DSpark paper's reference acceptance), and it produces the best chat speedups in the table.

Good to know#

  • Acceptance on tool-call JSON is the highest this project has measured anywhere: 5.07 tokens per round in a Claude Code session.
  • Copy-heavy editing: a rename refactor runs at 3.57× (93 tok/s) with lookup drafts.
  • 4-bit trades ratio for speed: about 1.4–1.55× but 60–76 tok/s (plain baseline 49.3). Pick 4-bit for peak tok/s, 8-bit for quality and the headline ratio. The same drafter resolves for both.
  • A bf16 target is slower in both ratio and absolute speed (1.54× on code at 22.9 tok/s). Don't bother.
  • Tool calls use the XML <function=…> form, translated to standard tool calls.

Numbers not matching your Mac? They shouldn't be identical: the draft cap is derived from your machine's measured curves. Run mlx-dspark benchmark --model mlx-community/Ornith-1.0-9B-8bit --trials 3 to get your own. See how these are measured.