Packets and recipes
The deployment variants of this family, at every maturity.
Qwen3.5 9B FP8-dynamic · one B70 · MTP depth 3 with the draft INT4 head
Reproduce 98.14 tok/s · candidate package
Center of two fresh-server class-balanced medians (98.251 / 98.027) over the strict fixed 12-prompt six-class suite through the completions API, 512-token cap, separate empty compile caches, 12/12 complete token arrays exact versus a same-configuration MTP0 oracle (50.165 / 50.173), canaries on every server, cache zero; one B70, R276 image, qwen3_5_mtp depth 3, draft-only INT4 lm_head (VLLM_XPU_DRAFT_LM_HEAD_INT4=1), full decode-only XPU graph capture.Qwen3.5 9B W4A16 · one or two B70 · MTP depth 3 with the draft INT4 head
Reproduce 113.27 tok/s · candidate package
Center of two fresh-server class-balanced medians over the strict fixed 12-prompt six-class suite through the completions API, 512-token cap, separate empty compile caches, 12/12 complete token arrays exact versus a same-configuration MTP0 oracle (64.332 / 64.338), canaries on every server, cache zero; one B70, R276 image, qwen3_5_mtp depth 3, draft-only INT4 lm_head, full decode-only XPU graph capture.What has been classified
Dense scorecards summarize every declared combination; measured slices retain exact evidence links.
MTP × TP (FP8 route)
FP8-dynamic revision, R276 vLLM XPU, graph capture on, draft INT4 head where applicable; depth 4 rows are withheld (not lossless). Fixed: revision=qwen3.5-9b-fp8-dynamic-790f0576 · runtime=vLLM XPU 0.27.2rc1.dev77 (R276 image 521eb277).
- ✓ Measured1 card, no speculative decoding
TP1·MTP050.18 tok/sguide · evidence - ✓ Measured2 cards, no speculative decoding
TP2·MTP079.46 tok/sguide · evidence - ✓ Measured1 card, + speculative decoding (depth 3)
TP1·MTP398.25 tok/sguide · evidence - ✓ Measured2 cards, + speculative decoding (depth 3)
TP2·MTP3147.89 tok/sguide · evidence - ✓ Measured1 card, + speculative decoding (depth 4)
TP1·MTP498.55 tok/sguide · evidence - ✓ Measured2 cards, + speculative decoding (depth 4)
TP2·MTP4151.74 tok/sguide · evidence - ✓ Measured1 card, + speculative decoding (depth 5)
TP1·MTP591.3 tok/sguide · evidence - ✓ Measured1 card, + speculative decoding (depth 6)
TP1·MTP688.68 tok/sguide · evidence
2 untested combinations
TP2·MTP5TP2·MTP6quantization × MTP
One B70, R276 vLLM XPU, graph capture on, draft INT4 head at depth 3. Fixed: runtime=vLLM XPU 0.27.2rc1.dev77 (R276 image 521eb277) · tp=1.
- ✓ Measuredno speculative decoding, FP8-dynamic
MTP0·FP8-dynamic50.18 tok/sguide · evidence - ✓ Measured+ speculative decoding (depth 3), FP8-dynamic
MTP3·FP8-dynamic98.25 tok/sguide · evidence - ✓ Measuredno speculative decoding, W4A16
MTP0·W4A1664.34 tok/sguide · evidence - ✓ Measured+ speculative decoding (depth 3), W4A16
MTP3·W4A16113.63 tok/sguide · evidence
Measured results
Every number links to its proof.
Speculative depth on one B70
FP8-dynamic · strict completions suite · graph capture on · MTP0 oracle vs depth 3 with the FP8 draft head vs the draft INT4 head · depth 4 with the INT4 head is withheld (8/12 vs the oracle)
draft INT4 head 0: 50.15–50.18 · 3: 98.03–98.25 · 4: 98.42–98.55 · 5: 91.29–91.3 · 6: 88.6–88.68
draft INT4 head 0: 53.71–54.33 · 3: 73.89–76.46 · 4: 80.74–81.01
One card vs two cards
FP8-dynamic · strict completions suite · graph capture on · MTP0 and depth 3 with the draft INT4 head
depth 3 + draft INT4 head 1: 98.03–98.25 · 2: 147.71–147.89
depth 3 + draft INT4 head 1: 73.89–76.46 · 2: 49.65–50.05
INT4 versus FP8 on one B70
Same model, same launcher, same suite. W4A16 routes to the row-invariant INT4 kernel; FP8 does not.
FP8-dynamic 0: 50.15–50.18 · 3: 98.03–98.25
FP8-dynamic 0: 53.71–54.33 · 3: 73.89–76.46
Fine print
Qwen3_5ForConditionalGeneration, 32 layers, 4096 hidden. one measured FP8-dynamic revision on one and two cards, depths 0-6 strict, 2K-32K ladders on both card counts. Measurements, artifact hashes, outputs, quality decisions, and speed stay pinned to their exact recorded identity.
Many people at once Not measured
Multi-user (aggregate) throughput has not been measured for this family. The lab’s one measured sweep so far is in the multi-user report; the projection block below includes a projected users curve where a projection exists.
How much faster could this get? No projection
No projection is shown: this model is not in the ML Bottleneck catalog yet. The measured numbers above stand on their own.