What we measured Lab-measured
Median of two fresh-server class-balanced medians over the full 12-prompt/six-class, 512-cap, cache-zero raw-completion suite; target-only TP2, MTP0, packaged --reasoning off launcher, 12/12 complete token arrays exact and all objective canaries passed. proof file
target-only = no draft model assisting · cache-zero = a fresh start, nothing pre-computed · MTP = multi-token prediction, a small draft the main model verifies
- Model
- Qwen3.8 27B Q8_0 target-only
- Publisher
- Alibaba / Qwen
- Checkpoint
- ggml-org/Qwen3.8-27B-GGUF
- Compression
- Q8_0 / F16 KV
- Software
- llama.cpp SYCL
- Cards
- 2× Intel Arc Pro B70 32 GiB
- Model weight bytes
- 28.6 GB
- Operating systems
- Linux
- Delivery
- native
- Good for
- general coding long context quality conservative
- Published
- 2026-08-27
- Clean-host replay
- not yet
Still missing before this becomes an install guide
- tested platform installation
- clean-host replay
- beginner recovery flow
- natural-prompt HTTP speed and TTFT context sweep beyond the qualified short-context headline
- queued TTFT and per-request latency profile
What to know
- The 36.726447 tok/s headline uses raw untemplated completions on the packaged --reasoning off launcher. Those outputs match the historical raw-completions oracle 12/12, but the result must not be relabeled as chat-template service throughput.
- Multi-user greedy token identity is batch-shape-dependent. The qualified curve proves complete isolated responses with zero cross-base oracle collisions, not sequential byte identity.
- The c8-to-c16 aggregate drop reproduced on both fresh servers. It is retained as measured and is not interpolated away.
- The concurrency profile measures aggregate batch-wall throughput, not per-request TTFT or latency under queueing.
Measured performance profiles Lab-measured
Qualified Q8_0 TP2 HTTP decode over exact active context Lab-measured
Q8_0 TP2 server prompt evaluation Lab-measured
Qualified Q8_0 TP2 HTTP TTFT over exact active context Lab-measured
Many people at once Lab-measured
Output-audited Q8_0 TP2 HTTP aggregate decode Lab-measured
How much faster could this get? No projection
This package's measured workload does not map cleanly onto a single model + compression + card-count shape, so no like-for-like projection is shown. The measured numbers above stand on their own.