Packets and recipes
The deployment variants of this family, at every maturity.
Qwen3 30B-A3B Instruct 2507 UD-Q4_K_XL
Measured evidence: 107.48 tok/s on 1× B70 · rapid research snapshot
12 unique cold prompts, 128 output tokens, median tokens 1-100 after TTFT, request prompt cache disabled lab report — documents the result; not a step-by-step install guideQwen3-Coder 30B-A3B Instruct UD-Q4_K_XL
Measured evidence: 108.12 tok/s on 1× B70 · rapid research snapshot
12 unique cold prompts, 128 output tokens, median tokens 1-100 after TTFT, server and request prompt caches disabled lab report — documents the result; not a step-by-step install guideWhat has been classified
Dense scorecards summarize every declared combination; measured slices retain exact evidence links.
sibling × TP
UD-Q4_K_XL, exact measured GGUF revisions and rapid runtime; sibling measurements never substitute for one another. Fixed: variant=UD-Q4_K_XL · runtime=llama.cpp SYCL fdb1db877 · mtp=0.
- ✓ Measured1 card, General 2507
TP1·General 2507107.48 tok/sreport · evidence - ✓ Measured1 card, Coder
TP1·Coder108.12 tok/sreport · evidence
4 untested combinations
TP2·General 2507TP4·General 2507TP2·CoderTP4·CoderMeasured results
Every number links to its proof.
Sibling models compared
1 = General Instruct 2507; 2 = Coder. Shared core geometry does not make these speed or quality results transferable.
Fine print
Qwen3MoeForCausalLM, 48 layers, 2048 hidden. partial transfer only; General Instruct 2507 and Coder are sibling model variants, not interchangeable revisions. Measurements, artifact hashes, outputs, quality decisions, and speed stay pinned to their exact recorded identity.
Many people at once Not measured
Multi-user (aggregate) throughput has not been measured for this family. The lab’s one measured sweep so far is in the multi-user report; the projection block below includes a projected users curve where a projection exists.
How much faster could this get? Projected — not measured
The ML Bottleneck physics engine projects a tuned-run target and the physical ceiling for the headline setup (qwen3-coder-30b-a3b-instruct · UD-Q4_K_XL · llama.cpp SYCL · 1× B70). The grade is optimization headroom against the tuned-run target, not model quality.
Loading projections from mlbottleneck.com…