Candidate package · vLLM XPU · 4× Intel Arc Pro B70

Qwen3.8 Flash-Next FP8 with lossless MTP1, never-routed experts host-placed and both reference Triton kernels restored, on four Intel Arc Pro B70 cards

Qwen's 125B-A6B hybrid-attention MoE, served from its official FP8 weights across four Arc Pro B70 cards. The n-gram table and embeddings live in host memory, every hot expert stays on the cards, the experts a routing census never selects are parked in host memory behind a per-expert table in the MoE kernel, and the two Triton kernels the XPU port had replaced, the hyper-connection glue and the QSA pre-indexer, are restored to the model's own reference implementations. One speculative token per step, lossless within the lineage. The outputs are a new authority whose difference is at depth: exact-2K coincides with the certified stream and exact-4K does not, while the quality profile is preserved byte for byte. A replay of the lab's certified run.

What we measured Lab-measured

37.8tok/s class-balanced decode median

Median of prompt-class medians over 99 inter-token intervals after TTFT on the fixed cold 12-prompt realistic suite, sent once (A306, 2026-09-07). proof file

MTP = multi-token prediction, a small draft the main model verifies

Model
Qwen3.8 Flash-Next official FP8 export, TP4/EP4, deterministic full-decode graph, lossless MTP1, never-routed experts host-placed, hyper-connection glue and QSA pre-indexer restored to the reference kernels
Publisher
Qwen
Checkpoint
Qwen/Qwen3.8-Flash-Next-FP8
Compression
FP8 block-128 weights / BF16 KV
Software
vLLM XPU
Cards
4× Intel Arc Pro B70 32 GiB
Model weight bytes
Operating systems
Linux
Delivery
native
Good for
general coding deterministic serving research
Published
2026-09-07
Clean-host replay
not yet

Still missing before this becomes an install guide

  • tested platform installation (venv, oneAPI runtime, xe driver)
  • installable dependency hash lock
  • record-specific model acquisition helper
  • non-originating-host replay
  • built and replayed container image
  • beginner recovery flow
  • decode, prefill, and TTFT context sweep

Measured performance profiles Not published

No qualified structured context or depth profile is published for this package. Diagnostic evidence may still be linked under “What to know” or in the full guide; nothing is estimated in its place.

Many people at once Not published

No qualified multi-user aggregate profile is published for this exact package. Diagnostic or unsupported boundaries may still appear under “What to know” or in the full guide. Nothing is interpolated or promoted from a different model, quantization, runtime, or card count.

How much faster could this get? No projection

This package's measured workload does not map cleanly onto a single model + compression + card-count shape, so no like-for-like projection is shown. The measured numbers above stand on their own.