What we measured Lab-measured
Median of two fresh-server class-balanced medians over the fixed 12-prompt/six-class, 512-cap native HTTP suite; two B70s, equal target split, draft on SYCL0, MTP2, F16 target/draft KV, cache zero, and 24/24 complete candidate arrays exact to the fresh target-only oracle. proof file
target-only = no draft model assisting · MTP = multi-token prediction, a small draft the main model verifies · oracle = the fixed prompt set whose exact outputs are checked
- Model
- Qwen3.8 27B Q4_K_M target + Q4_0 MTP2 TP2
- Publisher
- Alibaba / Qwen
- Checkpoint
- ggml-org/Qwen3.8-27B-GGUF
- Compression
- Q4_K_M target + Q4_0 MTP draft / F16 KV
- Software
- llama.cpp SYCL
- Cards
- 2× Intel Arc Pro B70 32 GiB
- Model weight bytes
- 20.3 GB
- Operating systems
- Linux
- Delivery
- native
- Good for
- general coding interactive two-card throughput
- Published
- 2026-08-30
- Clean-host replay
- not yet
Still missing before this becomes an install guide
- measured 32K active-context profile for this exact MTP2 deployment
- output-qualified HTTP concurrency profile for this exact MTP2 deployment
- standard 512-token prompt TTFT and prefill measurement
- tested clean-host Intel driver and oneAPI installation
- clean-host source build and endpoint replay
- beginner recovery replay
What to know
- The 64.237301 tok/s headline is short-context, one-user evidence with 8K configured context. Target-only TP2 and one-card MTP2 context or concurrency values do not transfer.
- Only MTP0 and MTP2 were measured under this exact TP2 campaign. No deeper MTP mode is supported by this package.
- The tested 16 GiB host used swap and a 13 GiB process memory cap. A clean-host Intel/oneAPI install and source-build replay remain pending.
- The target and MTP draft come from two pinned repositories/revisions and both downloads are required.
- A locally rebuilt binary is a distinct identity until the paired full-suite oracle and MTP2 validation pass.
Measured performance profiles Lab-measured
Strict TP2 decode with and without MTP2 Lab-measured
Many people at once Not published
No qualified multi-user aggregate profile is published for this exact package. Diagnostic or unsupported boundaries may still appear under “What to know” or in the full guide. Nothing is interpolated or promoted from a different model, quantization, runtime, or card count.
How much faster could this get? No projection
This package's measured workload does not map cleanly onto a single model + compression + card-count shape, so no like-for-like projection is shown. The measured numbers above stand on their own.