Candidate package · llama.cpp SYCL · 1× Intel Arc Pro B70

Qwen3.8 27B Q8_0 + MTP2 on one Intel Arc Pro B70

The quality-conservative Q8_0 target with a separately downloaded Q4_0 MTP draft at depth 2. Two fresh servers measured 37.062 tok/s with all 24 speculative outputs exact to a matched MTP0 control.

What we measured Lab-measured

37.1tok/s strict varied-prompt decode

Median of two fresh-server class-balanced medians over the fixed 12-prompt/six-class, 512-cap native HTTP suite; one B70, TP1, 1024-token configured context, MTP2, F16 target/draft KV, cache zero, and 24/24 complete MTP2 arrays exact to matched MTP0. proof file

MTP = multi-token prediction, a small draft the main model verifies

Model
Qwen3.8 27B
Publisher
Alibaba / Qwen
Checkpoint
ggml-org/Qwen3.8-27B-GGUF plus unsloth/Qwen3.8-27B-GGUF MTP
Compression
Q8_0 target + Q4_0 MTP draft
Software
llama.cpp SYCL
Cards
1× Intel Arc Pro B70 32 GiB
Model weight bytes
30.0 GB
Operating systems
Linux
Delivery
native
Good for
general coding high quality interactive
Published
2026-08-27
Clean-host replay
not yet

Still missing before this becomes an install guide

  • Q8 plus MTP2 realistic-content HTTP speed and TTFT from 2K through 32K
  • Q8 plus MTP2 output-qualified concurrency
  • tested clean-host Intel driver and oneAPI installation
  • clean-host source build and endpoint replay
  • beginner recovery flow

What to know

  • The 37.062028 tok/s headline is a one-slot short-context result with a 1024-token configured context. No Q8+MTP2 long-context or concurrency value has been measured.
  • All 24 MTP2 outputs matched the matched MTP0 oracle in the strict suite, but this does not prove universal parity for every content or context shape.
  • The tested 16 GiB host used swap and a 13 GiB process memory cap. A 32 GiB or larger host is simpler.
  • A clean-host Intel/oneAPI installation and source-build replay remain pending.
  • The target and MTP draft come from two pinned repositories/revisions and both downloads are required.

Measured performance profiles Lab-measured

Strict Q8 decode over MTP depth Lab-measured

19.630.337.1MTP0MTP1MTP2Maximum MTP draft depth · y: tok/s
One-B70 fixed full-suite screen at 1024 configured context under one target/runtime identity. MTP0 and MTP1 are one fresh-server values; MTP2 is the median of two fresh servers. All speculative outputs were exact to the matched MTP0 oracle. No value is interpolated or extrapolated. evidence

Many people at once Not published

No qualified multi-user aggregate profile is published for this exact package. Diagnostic or unsupported boundaries may still appear under “What to know” or in the full guide. Nothing is interpolated or promoted from a different model, quantization, runtime, or card count.

How much faster could this get? No projection

This package's measured workload does not map cleanly onto a single model + compression + card-count shape, so no like-for-like projection is shown. The measured numbers above stand on their own.