What we measured Lab-measured
99-interval median across the fixed 12-prompt cache-zero suite; target-only, reasoning off, 12/12 exact output hashes. proof file
target-only = no draft model assisting · cache-zero = a fresh start, nothing pre-computed · 99-interval median = the middle rate across 99 measured stretches of a long answer
- Model
- Qwen3.8 27B Q4_K_M target-only TP2
- Publisher
- Alibaba / Qwen
- Checkpoint
- ggml-org/Qwen3.8-27B-GGUF
- Compression
- Q4_K_M / F16 KV
- Software
- llama.cpp SYCL
- Cards
- 2× Intel Arc Pro B70 32 GiB
- Model weight bytes
- 19.0 GB
- Operating systems
- Linux
- Delivery
- native
- Good for
- general coding long context two-card throughput
- Published
- 2026-08-23
- Clean-host replay
- not yet
Still missing before this becomes an install guide
- tested platform installation
- self-contained model download helper
- clean-host replay
- beginner recovery flow
- active-context prefill and power profiles plus a full memory curve
What to know
- The package has not been replayed from a clean host and is not a beginner install guide; source-rebuilt binaries require an explicit preflight override and the full output-oracle gate.
- The 49.717503 tok/s headline is a TP2 target-only reasoning-off result and must not be compared as a TP1, speculative, or alternate-accounting row.
- The optional large-batch mode improves prefill on a short probe while slightly reducing decode; it is not the headline configuration.
- Multi-user greedy token identity is batch-shape-dependent. The qualified HTTP curve proves complete isolated responses with zero cross-base oracle collisions, not sequential byte identity.
- The exact F16 ffn_down+ffn_gate cache is aggregate-only and uses approximately 13 GiB of additional device memory per card. Its c96 endpoint reached 192.341954 tok/s but used about 30.48/30.35 GiB per card and matched the same-shape batch oracle 96/96 while matching isolated sequential references only 50/96.
- The c96 endpoint includes peak VRAM samples, but no active-context prefill, full memory, or power curve has been captured for this TP2 identity.
Measured performance profiles Lab-measured
Qualified TP2 HTTP decode over exact active context Lab-measured
Qualified TP2 HTTP TTFT over exact active context Lab-measured
Many people at once Lab-measured
Output-audited TP2 HTTP aggregate decode Lab-measured
How much faster could this get? Projected — not measured
The ML Bottleneck physics engine projects what stock software, a tuned run, and the physical ceiling look like for this exact model, compression, card count, and software. The grade is optimization headroom against the tuned-run target, not model quality.
Loading projections from mlbottleneck.com…