Qwen3 30B-A3B · General + Coder

Alibaba's Qwen3 30B mixture-of-experts (about 3B active per word) in its General Instruct and Coder forms. Fast on one Arc Pro B70; the two siblings share a design but not their weights.

108.12tok/s
measured
≈ 81 words a secondqwen3-coder-30b-a3b-instruct · UD-Q4_K_XL · llama.cpp SYCL fdb1db877 · 1× B70 · Evidence D of A–D
Qwen3 30B-A3B generation · shared publisher/architecture family
Read the lab reportAll packets and recipesNo step-by-step install guide is published for this model yet.
B70 fit
Two siblings one card measured
Quality evidence
Rapid snapshot only
Measured results
2

Packets and recipes

The deployment variants of this family, at every maturity.

What has been classified

Dense scorecards summarize every declared combination; measured slices retain exact evidence links.

sibling × TP

UD-Q4_K_XL, exact measured GGUF revisions and rapid runtime; sibling measurements never substitute for one another. Fixed: variant=UD-Q4_K_XL · runtime=llama.cpp SYCL fdb1db877 · mtp=0.

  • ✓ Measured1 card, General 2507 TP1·General 2507107.48 tok/sreport · evidence
  • ✓ Measured1 card, Coder TP1·Coder108.12 tok/sreport · evidence
4 untested combinations
TP2·General 2507TP4·General 2507TP2·CoderTP4·Coder
Full matrix and exact selectors

Measured results

Every number links to its proof.

family data

Sibling models compared

1 = General Instruct 2507; 2 = Coder. Shared core geometry does not make these speed or quality results transferable.

General Instruct 2507 · model sibling 1107.48tok/sDecode166.95msTTFTstrict rapid snapshot · promoted representative row · evidence
Coder · model sibling 2108.12tok/sDecode164.13msTTFTstrict rapid snapshot · promoted representative row · evidence
General Instruct 2507Coder
evidence 1 · evidence 2
Fine print

Qwen3MoeForCausalLM, 48 layers, 2048 hidden. partial transfer only; General Instruct 2507 and Coder are sibling model variants, not interchangeable revisions. Measurements, artifact hashes, outputs, quality decisions, and speed stay pinned to their exact recorded identity.

Many people at once Not measured

Multi-user (aggregate) throughput has not been measured for this family. The lab’s one measured sweep so far is in the multi-user report; the projection block below includes a projected users curve where a projection exists.

How much faster could this get? Projected — not measured

The ML Bottleneck physics engine projects a tuned-run target and the physical ceiling for the headline setup (qwen3-coder-30b-a3b-instruct · UD-Q4_K_XL · llama.cpp SYCL · 1× B70). The grade is optimization headroom against the tuned-run target, not model quality.