Qwen3.5 9B

Alibaba's dense-hybrid Qwen3.5 9B (linear-attention layers with a full-attention layer every fourth, and a shipped MTP head), served on one Arc Pro B70 as RedHatAI's FP8-dynamic checkpoint through the lab's deterministic vLLM XPU stack with lossless speculative decoding. Two routes are measured: RedHatAI's FP8-dynamic and its W4A16 INT4, the latter faster and byte-exact at every concurrency through 64 users.

113.27tok/s
measured
≈ 85 words a secondqwen3.5-9b-w4a16-a398088c · W4A16 · vLLM XPU 0.27.2rc1.dev77 (R276 image 521eb277) · Evidence B of A–D
Qwen3.5 9B FP8-dynamic (RedHatAI) · measured FP8 weights with the publisher MTP headQwen3.5 9B W4A16 (RedHatAI) · measured INT4 weights with the publisher MTP head
B70 fit
Runs on 1 card (measured)
Quality evidence
Deployment identity evidence
Measured results
12

Packets and recipes

The deployment variants of this family, at every maturity.

What has been classified

Dense scorecards summarize every declared combination; measured slices retain exact evidence links.

MTP × TP (FP8 route)

FP8-dynamic revision, R276 vLLM XPU, graph capture on, draft INT4 head where applicable; depth 4 rows are withheld (not lossless). Fixed: revision=qwen3.5-9b-fp8-dynamic-790f0576 · runtime=vLLM XPU 0.27.2rc1.dev77 (R276 image 521eb277).

  • ✓ Measured1 card, no speculative decoding TP1·MTP050.18 tok/sguide · evidence
  • ✓ Measured2 cards, no speculative decoding TP2·MTP079.46 tok/sguide · evidence
  • ✓ Measured1 card, + speculative decoding (depth 3) TP1·MTP398.25 tok/sguide · evidence
  • ✓ Measured2 cards, + speculative decoding (depth 3) TP2·MTP3147.89 tok/sguide · evidence
  • ✓ Measured1 card, + speculative decoding (depth 4) TP1·MTP498.55 tok/sguide · evidence
  • ✓ Measured2 cards, + speculative decoding (depth 4) TP2·MTP4151.74 tok/sguide · evidence
  • ✓ Measured1 card, + speculative decoding (depth 5) TP1·MTP591.3 tok/sguide · evidence
  • ✓ Measured1 card, + speculative decoding (depth 6) TP1·MTP688.68 tok/sguide · evidence
2 untested combinations
TP2·MTP5TP2·MTP6
Full matrix and exact selectors

quantization × MTP

One B70, R276 vLLM XPU, graph capture on, draft INT4 head at depth 3. Fixed: runtime=vLLM XPU 0.27.2rc1.dev77 (R276 image 521eb277) · tp=1.

  • ✓ Measuredno speculative decoding, FP8-dynamic MTP0·FP8-dynamic50.18 tok/sguide · evidence
  • ✓ Measured+ speculative decoding (depth 3), FP8-dynamic MTP3·FP8-dynamic98.25 tok/sguide · evidence
  • ✓ Measuredno speculative decoding, W4A16 MTP0·W4A1664.34 tok/sguide · evidence
  • ✓ Measured+ speculative decoding (depth 3), W4A16 MTP3·W4A16113.63 tok/sguide · evidence

Measured results

Every number links to its proof.

family data

Speculative depth on one B70

FP8-dynamic · strict completions suite · graph capture on · MTP0 oracle vs depth 3 with the FP8 draft head vs the draft INT4 head · depth 4 with the INT4 head is withheld (8/12 vs the oracle)

027.655.282.8110.403456MTP depth · tok/sFP8 draft head · Decode (decode_tok_s) · MTP depth=0.0 · values=50.183, 50.15 tok/sFP8 draft head · Decode (decode_tok_s) · MTP depth=3.0 · values=76.917, 76.879 tok/sdraft INT4 head · Decode (decode_tok_s) · MTP depth=0.0 · values=50.183, 50.15 tok/sdraft INT4 head · Decode (decode_tok_s) · MTP depth=3.0 · values=98.251, 98.027 tok/sdraft INT4 head · Decode (decode_tok_s) · MTP depth=4.0 · values=98.418, 98.55 tok/sdraft INT4 head · Decode (decode_tok_s) · MTP depth=5.0 · values=91.301, 91.294 tok/sdraft INT4 head · Decode (decode_tok_s) · MTP depth=6.0 · values=88.602, 88.677 tok/s
FP8 draft headdraft INT4 head
FP8 draft head 0: 50.15–50.18 · 3: 76.88–76.92
draft INT4 head 0: 50.15–50.18 · 3: 98.03–98.25 · 4: 98.42–98.55 · 5: 91.29–91.3 · 6: 88.6–88.68
evidence 1

One card vs two cards

FP8-dynamic · strict completions suite · graph capture on · MTP0 and depth 3 with the draft INT4 head

041.482.8124.2165.612tensor parallel cards · tok/sMTP0 · Decode (decode_tok_s) · tensor parallel cards=1.0 · values=50.183, 50.15 tok/sMTP0 · Decode (decode_tok_s) · tensor parallel cards=2.0 · values=79.456, 79.4 tok/sdepth 3 + draft INT4 head · Decode (decode_tok_s) · tensor parallel cards=1.0 · values=98.251, 98.027 tok/sdepth 3 + draft INT4 head · Decode (decode_tok_s) · tensor parallel cards=2.0 · values=147.712, 147.89 tok/s
MTP0depth 3 + draft INT4 head
MTP0 1: 50.15–50.18 · 2: 79.4–79.46
depth 3 + draft INT4 head 1: 98.03–98.25 · 2: 147.71–147.89
evidence 1

INT4 versus FP8 on one B70

Same model, same launcher, same suite. W4A16 routes to the row-invariant INT4 kernel; FP8 does not.

031.863.695.4127.303MTP depth · tok/sW4A16 · Decode (decode_tok_s) · MTP depth=0.0 · values=64.332, 64.338 tok/sW4A16 · Decode (decode_tok_s) · MTP depth=3.0 · values=113.627, 112.904 tok/sFP8-dynamic · Decode (decode_tok_s) · MTP depth=0.0 · values=50.183, 50.15 tok/sFP8-dynamic · Decode (decode_tok_s) · MTP depth=3.0 · values=98.251, 98.027 tok/s
W4A16FP8-dynamic
W4A16 0: 64.33–64.34 · 3: 112.9–113.63
FP8-dynamic 0: 50.15–50.18 · 3: 98.03–98.25
evidence 1 · evidence 2
Fine print

Qwen3_5ForConditionalGeneration, 32 layers, 4096 hidden. one measured FP8-dynamic revision on one and two cards, depths 0-6 strict, 2K-32K ladders on both card counts. Measurements, artifact hashes, outputs, quality decisions, and speed stay pinned to their exact recorded identity.

Many people at once Not measured

Multi-user (aggregate) throughput has not been measured for this family. The lab’s one measured sweep so far is in the multi-user report; the projection block below includes a projected users curve where a projection exists.

How much faster could this get? No projection

No projection is shown: this model is not in the ML Bottleneck catalog yet. The measured numbers above stand on their own.