Candidate package · llama.cpp SYCL · 1× Intel Arc Pro B70

Qwen3.8 27B Q8_0 on one Intel Arc Pro B70

A quality-conservative one-card Qwen3.8 lane using Q8_0 weights, F16 KV, llama.cpp/SYCL, and no speculative decoding.

What we measured Lab-measured

19.6tok/s strict varied-prompt decode

Median of two fresh-server class-balanced medians over the full 12-prompt/six-class, 512-cap, cache-zero raw-completion suite; target-only TP1, MTP0, 12/12 complete token arrays exact within TP1 and all objective canaries passed. proof file

target-only = no draft model assisting · cache-zero = a fresh start, nothing pre-computed · MTP = multi-token prediction, a small draft the main model verifies

Model
Qwen3.8 27B
Publisher
Alibaba / Qwen
Checkpoint
ggml-org/Qwen3.8-27B-GGUF
Compression
Q8_0
Software
llama.cpp SYCL
Cards
1× Intel Arc Pro B70 32 GiB
Model weight bytes
28.6 GB
Operating systems
Linux
Delivery
native
Good for
general coding long context
Published
2026-08-27
Clean-host replay
not yet

Still missing before this becomes an install guide

  • tested clean-host Intel driver and oneAPI installation
  • clean-host source build and endpoint replay
  • beginner recovery flow
  • realistic-prompt HTTP speed and TTFT context sweep beyond the qualified short-context headline
  • queued TTFT and per-request latency profile

What to know

  • The 19.619240 tok/s headline is a short-context varied-prompt HTTP result. The separate depth curve remains raw llama-bench pp2048/tg128 and must not be substituted for that workload.
  • TP1 is deterministic 12/12 across its two fresh servers but differs from the TP2 raw oracle 12/12, with first divergences at generated tokens 59-444. It is a separately quality-gated arithmetic identity, not an exact cross-card-count output claim.
  • The low-latency launcher is one slot at 8K. The throughput launcher uses eight active slots at 4K total context and queues excess requests; its aggregate curve does not qualify queued TTFT or per-request latency.
  • Exact 64-slot/32K and 32-slot/16K F16-KV profiles do not fit one B70. A 16-slot/8K profile fits but is slower than queued p8 above eight users.
  • The tested host had 16 GB nominal RAM plus swap. Clean-host installation and replay remain pending.

Measured performance profiles Lab-measured

Q8_0 raw decode with F16 KV Lab-measured

19.719.619.519.318.818.418.002K4K8K16K24K32KExisting context depth before tg128 · y: tok/s
Direct llama-bench raw-engine tg128, exact Q8_0 TP1 tuple, one B70, F16 K/V, flash attention on, and five repetitions per displayed depth. Every marker is measured; no point is interpolated or extrapolated. evidence

Q8_0 raw pp2048 with F16 KV Lab-measured

99798796091483877371902K4K8K16K24K32KExisting context depth before pp2048 · y: tok/s
Direct llama-bench raw-engine pp2048, exact Q8_0 TP1 tuple, one B70, F16 K/V, flash attention on, and five repetitions per displayed depth. Every marker is measured; no point is interpolated or extrapolated. evidence

Many people at once Lab-measured

Queued Q8_0 TP1 HTTP aggregate decode Lab-measured

18.129.147.567.168.168.368.61248163264Simultaneous HTTP requests · log2 spacing · y: tok/s
llama.cpp native HTTP /completion on one B70 with eight active slots, 4K total F16 KV context, and excess requests queued at 16/32/64 incoming concurrency. Prompt caching and slot similarity were disabled; every response returned 128 raw token IDs. Each marker is the median of two preregistered fresh-server attempts whose relative range was at most 0.96%. Output isolation passed with no cross-base oracle collision, but multi-user greedy token identity is batch-shape-dependent. This is aggregate batch-wall throughput, not queued TTFT or per-request latency. No point is interpolated or extrapolated. evidence

How much faster could this get? No projection

This package's measured workload does not map cleanly onto a single model + compression + card-count shape, so no like-for-like projection is shown. The measured numbers above stand on their own.