Candidate package · llama.cpp SYCL · 2× Intel Arc Pro B70

Qwen3.8 27B Q4_K_M + MTP2 on two Intel Arc Pro B70 cards

Two-B70 Qwen3.8 27B Q4_K_M with target-exact Q4_0 MTP2.

What we measured Lab-measured

64.2tok/s strict varied-prompt decode

Median of two fresh-server class-balanced medians over the fixed 12-prompt/six-class, 512-cap native HTTP suite; two B70s, equal target split, draft on SYCL0, MTP2, F16 target/draft KV, cache zero, and 24/24 complete candidate arrays exact to the fresh target-only oracle. proof file

target-only = no draft model assisting · MTP = multi-token prediction, a small draft the main model verifies · oracle = the fixed prompt set whose exact outputs are checked

Model
Qwen3.8 27B Q4_K_M target + Q4_0 MTP2 TP2
Publisher
Alibaba / Qwen
Checkpoint
ggml-org/Qwen3.8-27B-GGUF
Compression
Q4_K_M target + Q4_0 MTP draft / F16 KV
Software
llama.cpp SYCL
Cards
2× Intel Arc Pro B70 32 GiB
Model weight bytes
20.3 GB
Operating systems
Linux
Delivery
native
Good for
general coding interactive two-card throughput
Published
2026-08-30
Clean-host replay
not yet

Still missing before this becomes an install guide

  • measured 32K active-context profile for this exact MTP2 deployment
  • output-qualified HTTP concurrency profile for this exact MTP2 deployment
  • standard 512-token prompt TTFT and prefill measurement
  • tested clean-host Intel driver and oneAPI installation
  • clean-host source build and endpoint replay
  • beginner recovery replay

What to know

  • The 64.237301 tok/s headline is short-context, one-user evidence with 8K configured context. Target-only TP2 and one-card MTP2 context or concurrency values do not transfer.
  • Only MTP0 and MTP2 were measured under this exact TP2 campaign. No deeper MTP mode is supported by this package.
  • The tested 16 GiB host used swap and a 13 GiB process memory cap. A clean-host Intel/oneAPI install and source-build replay remain pending.
  • The target and MTP draft come from two pinned repositories/revisions and both downloads are required.
  • A locally rebuilt binary is a distinct identity until the paired full-suite oracle and MTP2 validation pass.

Measured performance profiles Lab-measured

Strict TP2 decode with and without MTP2 Lab-measured

49.864.2MTP0MTP2Maximum MTP draft depth · y: tok/s
Same target, runtime, two-card equal split, F16 KV, cache-zero suite, and short-context contract. MTP0 is one fresh oracle; MTP2 is the median of two fresh servers and all 24 candidate arrays were target-exact. Only depths 0 and 2 were measured in this TP2 campaign; no point is interpolated or extrapolated. evidence

Many people at once Not published

No qualified multi-user aggregate profile is published for this exact package. Diagnostic or unsupported boundaries may still appear under “What to know” or in the full guide. Nothing is interpolated or promoted from a different model, quantization, runtime, or card count.

How much faster could this get? No projection

This package's measured workload does not map cleanly onto a single model + compression + card-count shape, so no like-for-like projection is shown. The measured numbers above stand on their own.