What we measured Lab-measured
Center of two fresh-server class-balanced medians (98.251 / 98.027) over the strict fixed 12-prompt six-class suite through the completions API, 512-token cap, separate empty compile caches, 12/12 complete token arrays exact versus a same-configuration MTP0 oracle (50.165 / 50.173), canaries on every server, cache zero; one B70, R276 image, qwen3_5_mtp depth 3, draft-only INT4 lm_head (VLLM_XPU_DRAFT_LM_HEAD_INT4=1), full decode-only XPU graph capture. proof file
MTP = multi-token prediction, a small draft the main model verifies · oracle = the fixed prompt set whose exact outputs are checked
- Model
- Qwen3.5 9B FP8-dynamic (compressed-tensors, per-channel FP8 weights, dynamic activations) with the publisher's MTP head
- Publisher
- RedHatAI (FP8-dynamic of Alibaba / Qwen)
- Checkpoint
- RedHatAI/Qwen3.5-9B-FP8-dynamic
- Compression
- FP8-dynamic (compressed-tensors, per-channel FP8 weights, dynamic activations)
- Software
- vLLM XPU
- Cards
- 1× Intel Arc Pro B70 32 GiB
- Model weight bytes
- 14.0 GB
- Operating systems
- Linux
- Delivery
- Docker / container
- Good for
- general coding
- Published
- 2026-09-07
- Clean-host replay
- not yet
Still missing before this becomes an install guide
- clean-host replay
Measured performance profiles Lab-measured
Qwen3.5-9B FP8 one card, MTP depth 3 with the draft INT4 head: decode over active context (2K-32K) Lab-measured
Qwen3.5-9B FP8 one card, no speculation: decode over active context (2K-32K) Lab-measured
Qwen3.5-9B FP8 two cards (TP2), MTP depth 3 with the draft INT4 head: decode over active context (2K-32K) Lab-measured
Many people at once Lab-measured
Qwen3.5-9B FP8 one card, MTP depth 3 with the draft INT4 head and XPU graph capture, identity-qualified aggregate decode vs concurrent users (c1-c8; warm pass) Lab-measured
Qwen3.5-9B FP8 one card, no speculation, XPU graph capture, identity-qualified aggregate decode vs concurrent users (c1-c8; warm pass) Lab-measured
Qwen3.5-9B FP8 two cards (TP2), MTP depth 3 with the draft INT4 head and XPU graph capture, identity-qualified aggregate decode vs concurrent users (c1-c8; warm pass) Lab-measured
Qwen3.5-9B FP8 two cards (TP2), no speculation, XPU graph capture, identity-qualified aggregate decode vs concurrent users (c1-c8; warm pass) Lab-measured
How much faster could this get? Projected — not measured
The ML Bottleneck physics engine projects what stock software, a tuned run, and the physical ceiling look like for this exact model, compression, card count, and software. The grade is optimization headroom against the tuned-run target, not model quality.
Loading projections from mlbottleneck.com…