Packets and recipes
The deployment variants of this family, at every maturity.
What has been classified
Dense scorecards summarize every declared combination; measured slices retain exact evidence links.
quant × TP
Exact measured revision and rapid runtime; only TP1 has stored rows. Fixed: revision=mistral-small-3.2-24b-instruct-2506-b750ec2 · runtime=llama.cpp SYCL fdb1db877 · mtp=0.
- ✓ Measured1 card, UD-Q4_K_XL
TP1·UD-Q4_K_XL27.3 tok/sreport · evidence - ✓ Measured1 card, UD-Q8_K_XL
TP1·UD-Q8_K_XL16.38 tok/sevidence
4 untested combinations
TP2·UD-Q4_K_XLTP4·UD-Q4_K_XLTP2·UD-Q8_K_XLTP4·UD-Q8_K_XLMeasured results
Every number links to its proof.
Compression options compared
1 = UD-Q4_K_XL; 2 = UD-Q8_K_XL. Same weight revision and rapid suite; each quant retains its own measured decode and TTFT.
Fine print
Mistral Small 3.2 dense 24B. one weight revision; Q4 and Q8 are deployment variants with separate measurements. Measurements, artifact hashes, outputs, quality decisions, and speed stay pinned to their exact recorded identity.
Many people at once Not measured
Multi-user (aggregate) throughput has not been measured for this family. The lab’s one measured sweep so far is in the multi-user report; the projection block below includes a projected users curve where a projection exists.
How much faster could this get? Projected — not measured
The ML Bottleneck physics engine projects a tuned-run target and the physical ceiling for the headline setup (mistral-small-3.2-24b-instruct-2506- · UD-Q4_K_XL · llama.cpp SYCL · 1× B70). The grade is optimization headroom against the tuned-run target, not model quality.
Loading projections from mlbottleneck.com…