What we measured Lab-measured
Median of two fresh-server class-balanced medians over the fixed 12-prompt/six-class, 512-cap native HTTP suite; one B70, TP1, MTP2, F16 target/draft KV, cache zero, 12/12 complete arrays exact between replicas and against same-build MTP0. proof file
MTP = multi-token prediction, a small draft the main model verifies
- Model
- Qwen3.8 27B
- Publisher
- Alibaba / Qwen
- Checkpoint
- ggml-org/Qwen3.8-27B-GGUF
- Compression
- Q4_K_M target + Q4_0 MTP draft
- Software
- llama.cpp SYCL
- Cards
- 1× Intel Arc Pro B70 32 GiB
- Model weight bytes
- 20.3 GB
- Operating systems
- Linux
- Delivery
- native
- Good for
- general coding interactive
- Published
- 2026-08-27
- Clean-host replay
- not yet
Still missing before this becomes an install guide
- natural retrieval/task long-context HTTP suite beyond the measured raw-document continuation profile
- tested clean-host Intel driver and oneAPI installation
- clean-host source build and endpoint replay
- beginner recovery flow
What to know
- The 42.636988 tok/s headline is the fixed short-context realistic suite. The separate Grade B context profile uses raw continuations of unrepeated technical prose, Python code, and structured documentation; it measured 36.505065 tok/s and 39.538 s TTFT at exact 32K. It is not a natural retrieval/task suite.
- All 36 outputs in the new real-content MTP2 profile matched the fresh MTP0 oracle, but the older repeated-token diagnostic still reproducibly diverged at 2K/generated token 23. Target parity remains workload-scoped rather than universal.
- Output-qualified concurrency reaches 68.341 aggregate tok/s at 16 users only in the measured 16-slot/8K-total service profile. Larger 32- and 64-slot MTP2 profiles failed startup with device OOM; this is not a 32K-per-user or 64-user claim.
- MTP5 is explicitly unsafe for this identity: it changed all 12 complete target outputs and is not a supported speed mode.
- The tested 16 GiB host used swap and a 13 GiB process memory cap. A clean-host Intel/oneAPI install and source-build replay remain pending.
- The target and MTP draft come from two pinned repositories/revisions and both downloads are required.
Measured performance profiles Lab-measured
Strict decode over MTP depth Lab-measured
Three-class real-content HTTP decode over active context Lab-measured
Three-class real-content HTTP TTFT over active context Lab-measured
Many people at once Lab-measured
Output-audited MTP2 HTTP aggregate decode Lab-measured
How much faster could this get? No projection
This package's measured workload does not map cleanly onto a single model + compression + card-count shape, so no like-for-like projection is shown. The measured numbers above stand on their own.