Packets and recipes
The deployment variants of this family, at every maturity.
Muse-Glimmer 30B Q8/WOQ + DFlash · TP4
Reproduce 100.37 tok/s on 4× B70 · closed-banked-record
Pooled arithmetic mean across two fresh canonical full-256 three-prompt runs; individual prompts and full natural completions can be slower.Muse-Glimmer 30B BF16 + DFlash fleet
Evidence packet · banked-service-runbook-inactive-at-last-workspace-check
lab report — documents the result; not a step-by-step install guideMuse-Glimmer 30B kquant-dynamic fleet
Evidence packet · research-only-external-quality-tier
lab report — documents the result; not a step-by-step install guideWhat has been classified
Dense scorecards summarize every declared combination; measured slices retain exact evidence links.
UD-Q8_K_XL · DFlash
Muse-Glimmer UD-Q8_K_XL · f16 KV · DFlash target verification · MTP is not used. TP2 is older upstream research; TP4 is the separate promoted Q8/WOQ record stack.
BF16 · DFlash
Muse-Glimmer BF16 target · f16 KV · DFlash · MTP is not used. TP2 covers the upstream/fleet-shaped profile; TP4 is a default-off kernel-campaign closeout, not the banked two-replica service identity.
- ✓ Measured2 cards, no speculative decoding
TP2·MTP028.7 tok/sevidence - ✓ Measured4 cards, no speculative decoding
TP4·MTP064.012 closeoutevidence - × Unsupported1 card, no speculative decoding
TP1·MTP0— The 55.7 GB BF16 target requires at least two 32 GiB cards in the tested layouts.target does not fitevidence
kquant-dynamic · DFlash
Muse-Glimmer kquant-dynamic · one-card replicas · DFlash n4 · MTP is not used. The four-card fleet figure is four independent TP1 replicas, not TP4.
- ✓ Measured1 card, no speculative decoding
TP1·MTP027.3 tok/sevidence
2 untested combinations
TP2·MTP0TP4·MTP0Measured results
Every number links to its proof.
Q8 prompt classes
TP4 · full 256 tokens · x=1 prose, 2 code, 3 JSON · two fresh servers; prompt arithmetic means are stored separately
fresh run 2 1: 72.49 · 2: 106.67 · 3: 122.79
Compression and card-count options
Older upstream 256-token three-class per-replica means; distinct quantization and speculation identities, shown as discrete points
BF16 target-only 2: 9.85
BF16 + DFlash n5 2: 28.7
dynamic target-only 1: 24.2
dynamic + DFlash n4 1: 27.3
Uncompressed (BF16) results
TP4 · fixed cold prose/code/JSON classes · arithmetic mean 64.012; default-off research stack, not the banked service identity
Fine print
dense causal transformer plus perception encoder, 52 layers, 6656 hidden. shared architecture; each target quantization, draft artifact, topology, and modality requires its own gate. Measurements, artifact hashes, outputs, quality decisions, and speed stay pinned to their exact recorded identity.
2 combinations the lab has stopped pursuing: The near-lossless target and DFlash buffers exceed the measured single-card memory envelope; reopening requires a source-level memory reduction. evidence The final exact kernel stack reached a 64.012 tok/s three-class mean. The original lossless BF16 century objective remained unmet and was not retroactively satisfied by the Q8 successor. evidence Exact selectors are in the family data.
Many people at once Not measured
Multi-user (aggregate) throughput has not been measured for this family. The lab’s one measured sweep so far is in the multi-user report; the projection block below includes a projected users curve where a projection exists.
How much faster could this get? Projected — not measured
The ML Bottleneck physics engine projects a tuned-run target and the physical ceiling for the headline setup (muse-glimmer-30b · UD-Q8_K_XL target / BF16 draft · llama.cpp SYCL patched Muse stack · 4× B70). The grade is optimization headroom against the tuned-run target, not model quality.
Loading projections from mlbottleneck.com…