Candidate package · vLLM XPU · 4× Intel Arc Pro B70

Qwen3.8 Flash-Next FP8 with lossless MTP1 on four Intel Arc Pro B70 cards

Qwen's 125B-A6B hybrid-attention MoE, served from its official FP8 weights across four Arc Pro B70 cards with the expert and n-gram tables that do not fit placed in host memory. One speculative token per step with every output identical to the no-speculation line. A replay of the lab's certified run.

What we measured Lab-measured

27.0tok/s class-balanced decode median

Median of prompt-class medians over 99 inter-token intervals after TTFT on the fixed cold 12-prompt realistic suite, sent once (A189, 2026-09-05). proof file

MTP = multi-token prediction, a small draft the main model verifies

Model
Qwen3.8 Flash-Next official FP8 export, TP4/EP4, deterministic full-decode graph, lossless MTP1
Publisher
Qwen
Checkpoint
Qwen/Qwen3.8-Flash-Next-FP8
Compression
FP8 block-128 weights / BF16 KV
Software
vLLM XPU
Cards
4× Intel Arc Pro B70 32 GiB
Model weight bytes
Operating systems
Linux
Delivery
native
Good for
general coding deterministic serving research
Published
2026-09-06
Clean-host replay
not yet

Still missing before this becomes an install guide

  • tested platform installation (venv, oneAPI runtime, xe driver)
  • installable dependency hash lock
  • record-specific model acquisition helper
  • non-originating-host replay
  • built and replayed container image
  • beginner recovery flow
  • decode, prefill, and TTFT context sweep

Measured performance profiles Not published

No qualified structured context or depth profile is published for this package. Diagnostic evidence may still be linked under “What to know” or in the full guide; nothing is estimated in its place.

Many people at once Not published

No qualified multi-user aggregate profile is published for this exact package. Diagnostic or unsupported boundaries may still appear under “What to know” or in the full guide. Nothing is interpolated or promoted from a different model, quantization, runtime, or card count.

How much faster could this get? No projection

This package's measured workload does not map cleanly onto a single model + compression + card-count shape, so no like-for-like projection is shown. The measured numbers above stand on their own.