What we measured Lab-measured
Median of two fresh-server class-balanced medians over the full 12-prompt/six-class, 512-cap, cache-zero raw-completion suite; target-only TP1, MTP0, 12/12 complete token arrays exact within TP1 and all objective canaries passed. proof file
target-only = no draft model assisting · cache-zero = a fresh start, nothing pre-computed · MTP = multi-token prediction, a small draft the main model verifies
- Model
- Qwen3.8 27B
- Publisher
- Alibaba / Qwen
- Checkpoint
- ggml-org/Qwen3.8-27B-GGUF
- Compression
- Q8_0
- Software
- llama.cpp SYCL
- Cards
- 1× Intel Arc Pro B70 32 GiB
- Model weight bytes
- 28.6 GB
- Operating systems
- Linux
- Delivery
- native
- Good for
- general coding long context
- Published
- 2026-08-27
- Clean-host replay
- not yet
Still missing before this becomes an install guide
- tested clean-host Intel driver and oneAPI installation
- clean-host source build and endpoint replay
- beginner recovery flow
- realistic-prompt HTTP speed and TTFT context sweep beyond the qualified short-context headline
- queued TTFT and per-request latency profile
What to know
- The 19.619240 tok/s headline is a short-context varied-prompt HTTP result. The separate depth curve remains raw llama-bench pp2048/tg128 and must not be substituted for that workload.
- TP1 is deterministic 12/12 across its two fresh servers but differs from the TP2 raw oracle 12/12, with first divergences at generated tokens 59-444. It is a separately quality-gated arithmetic identity, not an exact cross-card-count output claim.
- The low-latency launcher is one slot at 8K. The throughput launcher uses eight active slots at 4K total context and queues excess requests; its aggregate curve does not qualify queued TTFT or per-request latency.
- Exact 64-slot/32K and 32-slot/16K F16-KV profiles do not fit one B70. A 16-slot/8K profile fits but is slower than queued p8 above eight users.
- The tested host had 16 GB nominal RAM plus swap. Clean-host installation and replay remain pending.
Measured performance profiles Lab-measured
Q8_0 raw decode with F16 KV Lab-measured
Q8_0 raw pp2048 with F16 KV Lab-measured
Many people at once Lab-measured
Queued Q8_0 TP1 HTTP aggregate decode Lab-measured
How much faster could this get? No projection
This package's measured workload does not map cleanly onto a single model + compression + card-count shape, so no like-for-like projection is shown. The measured numbers above stand on their own.