GLM 4.7 · Flash 30B-A3B

Zhipu's GLM-4.7 Flash, a 30B mixture-of-experts with about 3B active per word. A fast general assistant that fits one Arc Pro B70 in 4-bit form.

40.77tok/s
measured
≈ 31 words a secondglm-4.7-flash-0d32489 · UD-Q4_K_XL · llama.cpp SYCL fdb1db877 · 1× B70 · Evidence D of A–D
GLM-4.7-Flash · measured GGUF weights
Read the lab reportAll packets and recipesNo step-by-step install guide is published for this model yet.
B70 fit
Runs on 1 card (measured)
Quality evidence
Rapid snapshot only
Measured results
1

Packets and recipes

The deployment variants of this family, at every maturity.

What has been classified

Dense scorecards summarize every declared combination; measured slices retain exact evidence links.

quant × TP

Exact measured revision and rapid runtime; TP2/4 are unmeasured gaps. Fixed: revision=glm-4.7-flash-0d32489 · runtime=llama.cpp SYCL fdb1db877 · mtp=0.

  • ✓ Measured1 card, UD-Q4_K_XL TP1·UD-Q4_K_XL40.77 tok/sreport · evidence
2 untested combinations
TP2·UD-Q4_K_XLTP4·UD-Q4_K_XL
Full matrix and exact selectors

Measured results

Every number links to its proof.

family data

Measured speed (quick lab check)

UD-Q4_K_XL · one B70 · f16 KV · ctx=4096 configured; faster concurrent-screen rows are not promoted

UD-Q4_K_XL · tensor parallel cards 140.77tok/sDecode206.21msTTFTstrict rapid snapshot · promoted conservative standalone row · evidence
UD-Q4_K_XLmissing
evidence 1
Fine print

Glm4MoeLiteForCausalLM. one measured GLM-4.7-Flash GGUF revision. Measurements, artifact hashes, outputs, quality decisions, and speed stay pinned to their exact recorded identity.

Many people at once Not measured

Multi-user (aggregate) throughput has not been measured for this family. The lab’s one measured sweep so far is in the multi-user report; the projection block below includes a projected users curve where a projection exists.

How much faster could this get? Projected — not measured

The ML Bottleneck physics engine projects a tuned-run target and the physical ceiling for the headline setup (glm-4.7-flash- · UD-Q4_K_XL · llama.cpp SYCL · 1× B70). The grade is optimization headroom against the tuned-run target, not model quality.