Qwen3.6 35B A3B

Alibaba's Qwen3.6 35B mixture-of-experts (about 3B active per word), built for long documents. Measured on four cards under the lab's strictest quality gate, on a single card experimentally, and serving 64 people at once from one card.

93.55tok/s
measured
≈ 70 words a secondqwen3.6-35b-a3b-base · qwen36-35b-quark-w8a8-cced565 · Quark W8A8 INT8 · vLLM XPU 2026-06 strict program · TP4 · MTP0 · graph on · auto KV · max context 32,768 · Current strict deep-gate reference
Qwen3.6 35B A3B · shared base identity · conceptual base weight revision; retained evidence does not pin its upstream repository commit
2 exact artifacts · 2 quantizations
  • qwen36-35b-quark-w8a8-cced565 · Quark W8A8 INT8 · nameistoken/Qwen3.6-35B-A3B-Quark-W8A8-INT8@cced56592e8c8935f8220836b4baa04dfd389118
  • qwen36-35b-autoround-w4a16 · AutoRound W4A16 INT4 · abhinand/Qwen3.6-35B-A3B-int4-AutoRound@Hugging Face revision was not captured in the retained experiment evidence
Read the lab reportAll packets and recipesNo step-by-step install guide is published for this model yet.
B70 fit
Quant specific: Quark TP4 strict; AutoRound TP1 screened
Quality evidence
Strict Quark TP4; bounded AutoRound TP1
Measured results
2

Packets and recipes

The deployment variants of this family, at every maturity.

What has been classified

Dense scorecards summarize every declared combination; measured slices retain exact evidence links.

MTP × TP

Current June 2026 strict-program identity; the legacy TP4 high is shown separately. Fixed: revision=qwen3.6-35b-a3b-base · artifact_id=qwen36-35b-quark-w8a8-cced565 · variant=Quark W8A8 INT8 · runtime=vLLM XPU 2026-06 strict program · graph=on · kv=auto · configured_max_context_tokens=32768.

  • ✓ Measured4 cards, no speculative decoding TP4·MTP093.55 tok/sreport · evidence
  • ◇ Speed check only2 cards, no speculative decoding TP2·MTP085.87 tok/sevidence
  • ⚠ Quarantined4 cards, + speculative decoding (depth 1) TP4·MTP1 — output not usable. Observed: 107.77 tok/s.evidence
12 untested combinations
TP1·MTP0TP1·MTP1TP2·MTP1TP1·MTP2TP2·MTP2TP4·MTP2TP1·MTP3TP2·MTP3TP4·MTP3TP1·MTP4TP2·MTP4TP4·MTP4
Full matrix and exact selectors

Quark active context × TP

Canonical active prompt-depth grid. The retained Quark speed rows used p512/o512, and separate long-context quality checks are not relabeled as performance at any canonical depth. Fixed: revision=qwen3.6-35b-a3b-base · artifact_id=qwen36-35b-quark-w8a8-cced565 · variant=Quark W8A8 INT8 · runtime_family=vLLM XPU · mtp=0 · kv=auto.

No classified combinations in this slice yet.

21 untested combinations
TP1·0TP2·0TP4·0TP1·2KTP2·2KTP4·2KTP1·4KTP2·4KTP4·4KTP1·8KTP2·8KTP4·8KTP1·16KTP2·16KTP4·16KTP1·24KTP2·24KTP4·24KTP1·32KTP2·32KTP4·32K
Full matrix and exact selectors

AutoRound MTP × TP

AutoRound W4A16 r16, one-user p128/o1024 screen, 4K configured ceiling. Quark results and closures do not transfer into this slice. Fixed: revision=qwen3.6-35b-a3b-base · artifact_id=qwen36-35b-autoround-w4a16 · variant=AutoRound W4A16 INT4 · runtime=vLLM/XPU combined-runtime-guards r16 · graph=piecewise 1/32/64 · kv=auto · configured_max_context_tokens=4096 · natural_eos=False · simultaneous_users=1.

  • ◇ Speed check only1 card, no speculative decoding TP1·MTP0 — Direct speed and literal-canary evidence; not a deterministic deployment recipe.90.91 tok/sreport · evidence
14 untested combinations
TP2·MTP0TP4·MTP0TP1·MTP1TP2·MTP1TP4·MTP1TP1·MTP2TP2·MTP2TP4·MTP2TP1·MTP3TP2·MTP3TP4·MTP3TP1·MTP4TP2·MTP4TP4·MTP4
Full matrix and exact selectors

AutoRound active context × TP

Canonical active prompt-depth grid. The retained AutoRound run used p128/o1024 with max_model_len=4K, so it is not relabeled as a 0, 2K, or 4K active-context point. Fixed: revision=qwen3.6-35b-a3b-base · artifact_id=qwen36-35b-autoround-w4a16 · variant=AutoRound W4A16 INT4 · runtime_family=vLLM/XPU · mtp=0 · kv=auto.

No classified combinations in this slice yet.

21 untested combinations
TP1·0TP2·0TP4·0TP1·2KTP2·2KTP4·2KTP1·4KTP2·4KTP4·4KTP1·8KTP2·8KTP4·8KTP1·16KTP2·16KTP4·16KTP1·24KTP2·24KTP4·24KTP1·32KTP2·32KTP4·32K
Full matrix and exact selectors

Measured results

Every number links to its proof.

family data

Four-card results: current and earlier

1 = older LocalMaxxing-approved 99.428358 high; 2 = current 93.550542 deep-gate reference. Different gate generations are discrete history, not a regression A/B.

older approved high · gate generation 199.43tok/sDecodeolder gate LocalMaxxing approved result · historical high · evidence
current strict reference · gate generation 293.55tok/sDecode187.34msTTFTstrict deep gate · current strict reference · evidence
older approved highcurrent strict reference
evidence 1 · evidence 2

One card serving many users ◇ screened, experimental

Combined rate across all simultaneous users, one Arc Pro B70, AutoRound INT4 · experimental: the 2-16-user dip is the stack's graph-capture boundary, and 64-user determinism is still open

0291582.1873.11,1641248163264simultaneous users · tok/smeasured combined rate · Combined (aggregate_tok_s) · simultaneous users=1.0 · value=87.6382088209863 tok/smeasured combined rate · Combined (aggregate_tok_s) · simultaneous users=2.0 · value=77.18628015630631 tok/smeasured combined rate · Combined (aggregate_tok_s) · simultaneous users=4.0 · value=149.48350787682995 tok/smeasured combined rate · Combined (aggregate_tok_s) · simultaneous users=8.0 · value=276.7655048798564 tok/smeasured combined rate · Combined (aggregate_tok_s) · simultaneous users=16.0 · value=488.6474098238277 tok/smeasured combined rate · Combined (aggregate_tok_s) · simultaneous users=32.0 · value=823.1508450629442 tok/smeasured combined rate · Combined (aggregate_tok_s) · simultaneous users=64.0 · value=1039.408012834992 tok/s
measured combined rate
measured combined rate 1: 87.64 · 2: 77.19 · 4: 149.48 · 8: 276.77 · 16: 488.65 · 32: 823.15 · 64: 1,039
evidence 1
Fine print

Qwen3.6 35B A3B sparse MoE. one Qwen3.6 base weight revision; Quark and AutoRound are quantized deployment artifacts whose evidence stays separate. Measurements, artifact hashes, outputs, quality decisions, and speed stay pinned to their exact recorded identity.

1 combination the lab has stopped pursuing: The TP4 target-only lane is exhausted for now; reopen only for a controlled upstream bakeoff or a new graph-compatible speculative-state mechanism. evidence Exact selectors are in the family data.

Many people at once Lab-measured

64 simultaneous users share 1,039 combined tok/s on 1 card (AutoRound W4A16 INT4); the full curve is under Measured results and in the multi-user report.

How much faster could this get? Projected — not measured

The ML Bottleneck physics engine projects a tuned-run target and the physical ceiling for the headline setup (qwen3.6-35b-a3b-base · Quark W8A8 INT8 · vLLM XPU 2026-06 strict program · TP4 · MTP0 · graph on · auto KV · max context 32,768). The grade is optimization headroom against the tuned-run target, not model quality.