Models
Muse-Glimmer-30B
A dense 30B at 2.47× on 8-bit (8-bit quality at about 4-bit speed), or 1.7× on the 4-bit build in ~26 GB.
- Made by
- Meta
- Size
- 30B
- Architecture
- Dense multimodal, sliding + full attention
- Features
- Tool callingThinking
Run it (4-bit)
mlx-dspark serve --model mlx-community/Muse-Glimmer-30B-4bit
Serves an OpenAI + Anthropic API on http://127.0.0.1:8080. For a one-off answer, use mlx-dspark generate --model mlx-community/Muse-Glimmer-30B-4bit --prompt "…".
Drafter
DaoCloud/Muse-Glimmer-30B-DSparkresolves automatically with--mode auto(the default).
The drafter downloads with the model the first time you run it, and loads 4-bit quantized (acceptance doesn't depend on the drafter's precision).
Lookup drafts are off by default for this pair (measured as a net loss); --lookup-drafts turns them on.
Measured 2026-08-12 on an M4 Pro, 48 GB with mlx 0.32.0. This predates the September 2026 verify kernels, which raised every 8-bit model re-measured since, so it is probably conservative.
Run it (8-bit)
mlx-dspark serve --model mlx-community/Muse-Glimmer-30B-8bit
Serves an OpenAI + Anthropic API on http://127.0.0.1:8080. For a one-off answer, use mlx-dspark generate --model mlx-community/Muse-Glimmer-30B-8bit --prompt "…".
Resolves the same drafter as the 4-bit build.
Drafter
DaoCloud/Muse-Glimmer-30B-DSparkresolves automatically with--mode auto(the default).
The drafter downloads with the model the first time you run it, and loads 4-bit quantized (acceptance doesn't depend on the drafter's precision).
Measured 2026-08-12 on an M4 Pro, 48 GB with mlx 0.32.0. This predates the September 2026 verify kernels, which raised every 8-bit model re-measured since, so it is probably conservative.
Meta's Muse-Glimmer-30B is a dense, multimodal 30B with 3:1 sliding/full attention. It needs mlx-vlm ≥ 0.6.12, which a fresh install of mlx-dspark already requires. The community drafter by DaoCloud reuses both the target's embedding and its output head.
Which quant#
- 8-bit (~40 GB, fits a 48 GB Mac but tight) roughly doubles the 4-bit ratio: 2.47× mean at cap 4. It sits closer to the drafter's bf16 training verifier, and its verify curve stays flat to width 5. But 8-bit decode reads twice the bytes, so absolute speed is about the same as 4-bit on code. The better ratio buys 8-bit quality at about 4-bit speed.
- 4-bit (~26 GB) is the registry default for smaller Macs: 1.74× mean at cap 2.
- bf16 (~60 GB) doesn't fit a 48 GB Mac.
Lookup drafts are off by default for this pair. Its output scaling and logit softcap make near-ties more frequent than on a typical model, so it diverges from single-row greedy at more positions. Every one of those divergences is a sub-ulp floating-point tie.
Numbers not matching your Mac? They shouldn't be identical: the draft cap is derived from your machine's measured curves. Run mlx-dspark benchmark --model mlx-community/Muse-Glimmer-30B-4bit --trials 3 to get your own. See how these are measured.