Mistral Small 3.2 · 24B Instruct 2506

Mistral Small 3.2, a dense 24B instruct model from Mistral AI. A solid general assistant that fits one Arc Pro B70 in 4-bit form.

27.3tok/s
measured
≈ 20 words a secondmistral-small-3.2-24b-instruct-2506-b750ec2 · UD-Q4_K_XL · llama.cpp SYCL fdb1db877 · 1× B70 · Evidence D of A–D
Mistral Small 3.2 24B Instruct 2506 · measured weights
Read the lab reportAll packets and recipesNo step-by-step install guide is published for this model yet.
B70 fit
Runs on 1 card (measured)
Quality evidence
Rapid snapshot only
Measured results
2

Packets and recipes

The deployment variants of this family, at every maturity.

What has been classified

Dense scorecards summarize every declared combination; measured slices retain exact evidence links.

quant × TP

Exact measured revision and rapid runtime; only TP1 has stored rows. Fixed: revision=mistral-small-3.2-24b-instruct-2506-b750ec2 · runtime=llama.cpp SYCL fdb1db877 · mtp=0.

  • ✓ Measured1 card, UD-Q4_K_XL TP1·UD-Q4_K_XL27.3 tok/sreport · evidence
  • ✓ Measured1 card, UD-Q8_K_XL TP1·UD-Q8_K_XL16.38 tok/sevidence
4 untested combinations
TP2·UD-Q4_K_XLTP4·UD-Q4_K_XLTP2·UD-Q8_K_XLTP4·UD-Q8_K_XL
Full matrix and exact selectors

Measured results

Every number links to its proof.

family data

Compression options compared

1 = UD-Q4_K_XL; 2 = UD-Q8_K_XL. Same weight revision and rapid suite; each quant retains its own measured decode and TTFT.

UD-Q4_K_XL · quantization variant 127.3tok/sDecode1,502msTTFTstrict rapid snapshot · promoted representative row · evidence
UD-Q8_K_XL · quantization variant 216.38tok/sDecode2,686msTTFTstrict rapid snapshot · measured higher precision reference · evidence
UD-Q4_K_XLUD-Q8_K_XL
evidence 1 · evidence 2
Fine print

Mistral Small 3.2 dense 24B. one weight revision; Q4 and Q8 are deployment variants with separate measurements. Measurements, artifact hashes, outputs, quality decisions, and speed stay pinned to their exact recorded identity.

Many people at once Not measured

Multi-user (aggregate) throughput has not been measured for this family. The lab’s one measured sweep so far is in the multi-user report; the projection block below includes a projected users curve where a projection exists.

How much faster could this get? Projected — not measured

The ML Bottleneck physics engine projects a tuned-run target and the physical ceiling for the headline setup (mistral-small-3.2-24b-instruct-2506- · UD-Q4_K_XL · llama.cpp SYCL · 1× B70). The grade is optimization headroom against the tuned-run target, not model quality.