Muse-Glimmer 30B

Meta's Muse-Glimmer 30B, a dense multimodal model that reads images as well as text. Runs across four Arc Pro B70 cards with a draft model speeding up generation.

100.37tok/s
measured
≈ 75 words a secondmuse-glimmer-30b · UD-Q8_K_XL target / BF16 draft · llama.cpp SYCL patched Muse stack · 4× B70 · B70-verified
Muse-Glimmer 30B · current measured weights
B70 fit
Multi card
Quality evidence
Strong but variant scoped
Measured results
11

Packets and recipes

The deployment variants of this family, at every maturity.

What has been classified

Dense scorecards summarize every declared combination; measured slices retain exact evidence links.

UD-Q8_K_XL · DFlash

Muse-Glimmer UD-Q8_K_XL · f16 KV · DFlash target verification · MTP is not used. TP2 is older upstream research; TP4 is the separate promoted Q8/WOQ record stack.

  • ✓ Measured2 cards, no speculative decoding TP2·MTP038 tok/sevidence
  • ✓ Measured4 cards, no speculative decoding TP4·MTP0100.65 tok/sevidence
  • ■ Closed1 card, no speculative decoding TP1·MTP0 — The near-lossless target plus drafter exceeds the measured one-card envelope.fit blockedevidence

BF16 · DFlash

Muse-Glimmer BF16 target · f16 KV · DFlash · MTP is not used. TP2 covers the upstream/fleet-shaped profile; TP4 is a default-off kernel-campaign closeout, not the banked two-replica service identity.

  • ✓ Measured2 cards, no speculative decoding TP2·MTP028.7 tok/sevidence
  • ✓ Measured4 cards, no speculative decoding TP4·MTP064.012 closeoutevidence
  • × Unsupported1 card, no speculative decoding TP1·MTP0 — The 55.7 GB BF16 target requires at least two 32 GiB cards in the tested layouts.target does not fitevidence

kquant-dynamic · DFlash

Muse-Glimmer kquant-dynamic · one-card replicas · DFlash n4 · MTP is not used. The four-card fleet figure is four independent TP1 replicas, not TP4.

  • ✓ Measured1 card, no speculative decoding TP1·MTP027.3 tok/sevidence
2 untested combinations
TP2·MTP0TP4·MTP0
Full matrix and exact selectors

Measured results

Every number links to its proof.

family data

Q8 prompt classes

TP4 · full 256 tokens · x=1 prose, 2 code, 3 JSON · two fresh servers; prompt arithmetic means are stored separately

034.468.8103.1137.5123prompt class index · tok/sfresh run 1 · Decode (decode_tok_s) · prompt class index=1.0 · value=71.583 tok/sfresh run 1 · Decode (decode_tok_s) · prompt class index=2.0 · value=106.436 tok/sfresh run 1 · Decode (decode_tok_s) · prompt class index=3.0 · value=122.246 tok/sfresh run 2 · Decode (decode_tok_s) · prompt class index=1.0 · value=72.487 tok/sfresh run 2 · Decode (decode_tok_s) · prompt class index=2.0 · value=106.673 tok/sfresh run 2 · Decode (decode_tok_s) · prompt class index=3.0 · value=122.786 tok/s
fresh run 1fresh run 2
fresh run 1 1: 71.58 · 2: 106.44 · 3: 122.25
fresh run 2 1: 72.49 · 2: 106.67 · 3: 122.79
evidence 1

Compression and card-count options

Older upstream 256-token three-class per-replica means; distinct quantization and speculation identities, shown as discrete points

010.621.331.942.612tensor parallel cards per replica · tok/sQ8 + DFlash n5 · Decode (decode_tok_s) · tensor parallel cards per replica=2.0 · value=38.0 tok/sBF16 target-only · Decode (decode_tok_s) · tensor parallel cards per replica=2.0 · value=9.847 tok/sBF16 + DFlash n5 · Decode (decode_tok_s) · tensor parallel cards per replica=2.0 · value=28.7 tok/sdynamic target-only · Decode (decode_tok_s) · tensor parallel cards per replica=1.0 · value=24.2 tok/sdynamic + DFlash n4 · Decode (decode_tok_s) · tensor parallel cards per replica=1.0 · value=27.3 tok/s
Q8 + DFlash n5BF16 target-onlyBF16 + DFlash n5dynamic target-onlydynamic + DFlash n4
Q8 + DFlash n5 2: 38
BF16 target-only 2: 9.85
BF16 + DFlash n5 2: 28.7
dynamic target-only 1: 24.2
dynamic + DFlash n4 1: 27.3
evidence 1

Uncompressed (BF16) results

TP4 · fixed cold prose/code/JSON classes · arithmetic mean 64.012; default-off research stack, not the banked service identity

022.54567.489.9123prompt class index · tok/sBF16 closeout · Decode (decode_tok_s) · prompt class index=1.0 · value=45.728 tok/sBF16 closeout · Decode (decode_tok_s) · prompt class index=2.0 · value=66.028 tok/sBF16 closeout · Decode (decode_tok_s) · prompt class index=3.0 · value=80.281 tok/s
BF16 closeout
BF16 closeout 1: 45.73 · 2: 66.03 · 3: 80.28
evidence 1
Fine print

dense causal transformer plus perception encoder, 52 layers, 6656 hidden. shared architecture; each target quantization, draft artifact, topology, and modality requires its own gate. Measurements, artifact hashes, outputs, quality decisions, and speed stay pinned to their exact recorded identity.

2 combinations the lab has stopped pursuing: The near-lossless target and DFlash buffers exceed the measured single-card memory envelope; reopening requires a source-level memory reduction. evidence The final exact kernel stack reached a 64.012 tok/s three-class mean. The original lossless BF16 century objective remained unmet and was not retroactively satisfied by the Q8 successor. evidence Exact selectors are in the family data.

Many people at once Not measured

Multi-user (aggregate) throughput has not been measured for this family. The lab’s one measured sweep so far is in the multi-user report; the projection block below includes a projected users curve where a projection exists.

How much faster could this get? Projected — not measured

The ML Bottleneck physics engine projects a tuned-run target and the physical ceiling for the headline setup (muse-glimmer-30b · UD-Q8_K_XL target / BF16 draft · llama.cpp SYCL patched Muse stack · 4× B70). The grade is optimization headroom against the tuned-run target, not model quality.