What we measured Lab-measured
Median of two fresh-server class-balanced medians over the fixed 12-prompt/six-class, 512-cap native HTTP suite; one B70, TP1, 1024-token configured context, MTP2, F16 target/draft KV, cache zero, and 24/24 complete MTP2 arrays exact to matched MTP0. proof file
MTP = multi-token prediction, a small draft the main model verifies
- Model
- Qwen3.8 27B
- Publisher
- Alibaba / Qwen
- Checkpoint
- ggml-org/Qwen3.8-27B-GGUF plus unsloth/Qwen3.8-27B-GGUF MTP
- Compression
- Q8_0 target + Q4_0 MTP draft
- Software
- llama.cpp SYCL
- Cards
- 1× Intel Arc Pro B70 32 GiB
- Model weight bytes
- 30.0 GB
- Operating systems
- Linux
- Delivery
- native
- Good for
- general coding high quality interactive
- Published
- 2026-08-27
- Clean-host replay
- not yet
Still missing before this becomes an install guide
- Q8 plus MTP2 realistic-content HTTP speed and TTFT from 2K through 32K
- Q8 plus MTP2 output-qualified concurrency
- tested clean-host Intel driver and oneAPI installation
- clean-host source build and endpoint replay
- beginner recovery flow
What to know
- The 37.062028 tok/s headline is a one-slot short-context result with a 1024-token configured context. No Q8+MTP2 long-context or concurrency value has been measured.
- All 24 MTP2 outputs matched the matched MTP0 oracle in the strict suite, but this does not prove universal parity for every content or context shape.
- The tested 16 GiB host used swap and a 13 GiB process memory cap. A 32 GiB or larger host is simpler.
- A clean-host Intel/oneAPI installation and source-build replay remain pending.
- The target and MTP draft come from two pinned repositories/revisions and both downloads are required.
Measured performance profiles Lab-measured
Strict Q8 decode over MTP depth Lab-measured
Many people at once Not published
No qualified multi-user aggregate profile is published for this exact package. Diagnostic or unsupported boundaries may still appear under “What to know” or in the full guide. Nothing is interpolated or promoted from a different model, quantization, runtime, or card count.
How much faster could this get? No projection
This package's measured workload does not map cleanly onto a single model + compression + card-count shape, so no like-for-like projection is shown. The measured numbers above stand on their own.