Qwen3.5 4B

Alibaba's smallest dense-hybrid Qwen3.5 with a shipped MTP head, served on one Arc Pro B70. Only the INT4 W4A16 build is packaged: the FP8 build of the same model is not repeat-exact on this stack, while the INT4 route passes every identity gate.

177.29tok/s
measured
≈ 133 words a secondqwen3.5-4b-w4a16-7a613872 · W4A16 · vLLM XPU 0.27.2rc1.dev77 (R276 image 521eb277) · Evidence B of A–D
Qwen3.5 4B W4A16 (RedHatAI) · measured INT4 weights with the publisher MTP headQwen3.5 4B FP8-dynamic (RedHatAI) · measured but not packaged: fails the repeat-exactness gate
B70 fit
Runs on 1 card (measured)
Quality evidence
Deployment identity evidence
Measured results
2

Packets and recipes

The deployment variants of this family, at every maturity.

What has been classified

Dense scorecards summarize every declared combination; measured slices retain exact evidence links.

quantization × MTP

One B70, R276 vLLM XPU, graph capture on. The FP8 row is measured but rejected: it fails the repeat-exactness gate. Fixed: runtime=vLLM XPU 0.27.2rc1.dev77 (R276 image 521eb277) · tp=1.

  • ✓ Measuredno speculative decoding, W4A16 MTP0·W4A16102.62 tok/sguide · evidence
  • ✓ Measured+ speculative decoding (depth 3), W4A16 MTP3·W4A16177.41 tok/sguide · evidence
2 untested combinations
MTP0·FP8-dynamicMTP3·FP8-dynamic
Full matrix and exact selectors

Measured results

Every number links to its proof.

family data

Speculative depth on one B70

W4A16 · strict completions suite · graph capture on · the draft-only INT4 head at depth 3

W4A16 · MTP depth 0102.38–102.62tok/sDecode30.21–30.32msTTFTstrict fresh server pair · promoted MTP0 oracle (v1) · evidence
W4A16 · MTP depth 3177.17–177.41tok/sDecode44.91–45.05msTTFTstrict fresh server pair · promoted headline (v1); LocalMaxxing cmtrj2tp3000hps01n3fadg9d · evidence
W4A16
evidence 1
Fine print

Qwen3_5ForConditionalGeneration. one packaged INT4 revision on one card; the FP8 revision is measured and rejected. Measurements, artifact hashes, outputs, quality decisions, and speed stay pinned to their exact recorded identity.

Many people at once Not measured

Multi-user (aggregate) throughput has not been measured for this family. The lab’s one measured sweep so far is in the multi-user report; the projection block below includes a projected users curve where a projection exists.

How much faster could this get? No projection

No projection is shown: this model is not in the ML Bottleneck catalog yet. The measured numbers above stand on their own.