What we measured Lab-measured
Class-balanced median of per-input-class medians using conventional 99-interval rates on the fixed cold suite with the Q4_0 MTP draft assisting. The all-prompt median is 123.727369 tok/s; the historical 100-event compatibility figure is 124.977141 tok/s. proof file
99-interval median = the middle rate across 99 measured stretches of a long answer · MTP = multi-token prediction, a small draft the main model verifies
- Model
- Gemma 4 26B A4B
- Publisher
- Checkpoint
- unsloth/gemma-4-26B-A4B-it-GGUF
- Compression
- UD-Q8_K_XL + Q4_0 MTP
- Software
- llama.cpp SYCL
- Cards
- 1× Intel Arc Pro B70 32 GiB
- Model weight bytes
- 27.6 GB
- Operating systems
- Linux
- Delivery
- native
- Good for
- general coding long context
- Published
- 2026-08-22
- Clean-host replay
- not yet
Still missing before this becomes an install guide
- historical llama-server binary SHA-256
- historical locally quantized Q4_0 MTP draft SHA-256 and byte size (the reconstructed draft is repeat-stable at 1f6706e4… across three builds)
- beginner recovery flow exercised end-to-end on a fresh host (documented; the pinned oneAPI 2026.0 container path and two clean-rebuild gate replays were verified on the lab host on 2026-09-07)
Measured performance profiles Lab-measured
Decode over active context Lab-measured
Prefill over prompt length Lab-measured
Time to first token Lab-measured
Many people at once Not published
No qualified multi-user aggregate profile is published for this exact package. Diagnostic or unsupported boundaries may still appear under “What to know” or in the full guide. Nothing is interpolated or promoted from a different model, quantization, runtime, or card count; the projection below remains clearly labeled as projected.
How much faster could this get? Projected — not measured
The ML Bottleneck physics engine projects what stock software, a tuned run, and the physical ceiling look like for this exact model, compression, card count, and software. The grade is optimization headroom against the tuned-run target, not model quality.
Loading projections from mlbottleneck.com…