Candidate package · llama.cpp SYCL · 1× Intel Arc Pro B70

Qwen3.8 27B Q4_K_M + MTP2 on one Intel Arc Pro B70

The one-card Qwen3.8 Q4_K_M target with its separately downloaded Q4_0 MTP draft at depth 2. Two fresh servers measured a 55.75% strict decode gain with target-exact output.

What we measured Lab-measured

42.6tok/s strict varied-prompt decode

Median of two fresh-server class-balanced medians over the fixed 12-prompt/six-class, 512-cap native HTTP suite; one B70, TP1, MTP2, F16 target/draft KV, cache zero, 12/12 complete arrays exact between replicas and against same-build MTP0. proof file

MTP = multi-token prediction, a small draft the main model verifies

Model
Qwen3.8 27B
Publisher
Alibaba / Qwen
Checkpoint
ggml-org/Qwen3.8-27B-GGUF
Compression
Q4_K_M target + Q4_0 MTP draft
Software
llama.cpp SYCL
Cards
1× Intel Arc Pro B70 32 GiB
Model weight bytes
20.3 GB
Operating systems
Linux
Delivery
native
Good for
general coding interactive
Published
2026-08-27
Clean-host replay
not yet

Still missing before this becomes an install guide

  • natural retrieval/task long-context HTTP suite beyond the measured raw-document continuation profile
  • tested clean-host Intel driver and oneAPI installation
  • clean-host source build and endpoint replay
  • beginner recovery flow

What to know

  • The 42.636988 tok/s headline is the fixed short-context realistic suite. The separate Grade B context profile uses raw continuations of unrepeated technical prose, Python code, and structured documentation; it measured 36.505065 tok/s and 39.538 s TTFT at exact 32K. It is not a natural retrieval/task suite.
  • All 36 outputs in the new real-content MTP2 profile matched the fresh MTP0 oracle, but the older repeated-token diagnostic still reproducibly diverged at 2K/generated token 23. Target parity remains workload-scoped rather than universal.
  • Output-qualified concurrency reaches 68.341 aggregate tok/s at 16 users only in the measured 16-slot/8K-total service profile. Larger 32- and 64-slot MTP2 profiles failed startup with device OOM; this is not a 32K-per-user or 64-user claim.
  • MTP5 is explicitly unsafe for this identity: it changed all 12 complete target outputs and is not a supported speed mode.
  • The tested 16 GiB host used swap and a 13 GiB process memory cap. A clean-host Intel/oneAPI install and source-build replay remain pending.
  • The target and MTP draft come from two pinned repositories/revisions and both downloads are required.

Measured performance profiles Lab-measured

Strict decode over MTP depth Lab-measured

27.438.342.642.1MTP0MTP1MTP2MTP3Maximum MTP draft depth · y: tok/s
One-B70 fixed full-suite screen under one target/runtime identity. MTP0/1/3 are one fresh-server screen values; MTP2 is the qualified median of two fresh servers. All displayed speculative arms are 12/12 target-exact. MTP5 measured 32.241 tok/s but is excluded because it matched 0/12 target arrays. No value is interpolated or extrapolated. evidence

Three-class real-content HTTP decode over active context Lab-measured

41.441.741.534.732.336.52K4K8K16K24K32KExact active prompt tokens · y: tok/s
One-slot native HTTP completions on the exact MTP2 package identity, F16 target/draft KV, cache zero, and 128 returned token IDs. Each point is the median of two fresh-server class medians over unrepeated technical prose, Python code, and structured documentation; all 36 MTP2 outputs were exact to the fresh matched MTP0 oracle. Raw document continuations are representative real-content context shapes, not a natural retrieval/task suite. No point is interpolated or extrapolated. evidence

Three-class real-content HTTP TTFT over active context Lab-measured

2,0844,1808,57018,03028,34839,5382K4K8K16K24K32KExact active prompt tokens · y: ms
Direct TTFT from the same two-fresh-server, three-class, target-oracle-exact receipts as the real-content decode curve. Cache was zero; all displayed prompt depths and token counts are exact; no point is interpolated or extrapolated. evidence

Many people at once Lab-measured

Output-audited MTP2 HTTP aggregate decode Lab-measured

34.941.352.447.968.3124816Concurrent HTTP users · log2 spacing · y: tok/s
llama.cpp native HTTP /completion on one B70 with Q4_K_M target, Q4_0 MTP draft at depth 2, 16 slots, 8K total F16 target/draft context (512 nominal tokens per slot), prompt cache and slot similarity disabled, unique short prompts, and 128 returned raw token IDs per throughput request. Each marker is the median of two fresh-server attempts whose relative range was at most 3.04%; 256/256 separate concurrent exact-answer canaries passed. Multi-user greedy output is batch-shape-dependent. No point is interpolated or extrapolated. evidence

How much faster could this get? No projection

This package's measured workload does not map cleanly onto a single model + compression + card-count shape, so no like-for-like projection is shown. The measured numbers above stand on their own.