Phi-4 Mini · Instruct

Microsoft's Phi-4 Mini Instruct, a 3.8B model built to punch above its size on reasoning and instruction following. Very fast on one Arc Pro B70; the full-size Phi-4 is not covered yet.

96.55tok/s
measured
≈ 72 words a secondphi4-mini-instruct-7ff82c2 · Q4_K_M · llama.cpp SYCL fdb1db877 · 1× B70 · Evidence D of A–D
Phi-4 Mini Instruct · only measured Phi-4-family weights
Read the lab reportAll packets and recipesNo step-by-step install guide is published for this model yet.
B70 fit
Runs on 1 card (measured)
Quality evidence
Rapid snapshot only
Measured results
2

Packets and recipes

The deployment variants of this family, at every maturity.

What has been classified

Dense scorecards summarize every declared combination; measured slices retain exact evidence links.

quant × TP

Phi-4 Mini Instruct only; exact measured revision and rapid runtime. Fixed: revision=phi4-mini-instruct-7ff82c2 · runtime=llama.cpp SYCL fdb1db877 · mtp=0.

  • ✓ Measured1 card, Q4_K_M TP1·Q4_K_M96.55 tok/sreport · evidence
  • ✓ Measured1 card, Q8_0 TP1·Q8_072.25 tok/sevidence
4 untested combinations
TP2·Q4_K_MTP4·Q4_K_MTP2·Q8_0TP4·Q8_0
Full matrix and exact selectors

Measured results

Every number links to its proof.

family data

Compression options compared

1 = Q4_K_M; 2 = Q8_0. Same Mini weights and rapid suite; full-size Phi-4 is not represented.

Q4_K_M · quantization variant 196.55tok/sDecode69.94msTTFTstrict rapid snapshot · promoted speed row · evidence
Q8_0 · quantization variant 272.25tok/sDecode119.46msTTFTstrict rapid snapshot · measured higher quality quant reference · evidence
Q4_K_MQ8_0
evidence 1 · evidence 2
Fine print

Phi-4 Mini Instruct. Mini Instruct only; Q4 and Q8 are deployment variants of one measured revision. Measurements, artifact hashes, outputs, quality decisions, and speed stay pinned to their exact recorded identity.

Many people at once Not measured

Multi-user (aggregate) throughput has not been measured for this family. The lab’s one measured sweep so far is in the multi-user report; the projection block below includes a projected users curve where a projection exists.

How much faster could this get? Projected — not measured

The ML Bottleneck physics engine projects a tuned-run target and the physical ceiling for the headline setup (Q4_K_M · llama.cpp SYCL · 1× B70). The grade is optimization headroom against the tuned-run target, not model quality.