Packets and recipes
The deployment variants of this family, at every maturity.
Qwen3.6 35B Quark INT8 · 4 cards, fully checked
Measured evidence: 93.55 tok/s on 4× B70 · research, closed — the lab has stopped tuning this configuration for now
p512/o512, four post-warmup repeats, corrected output throughput after first streamed chunk lab report — documents the result; not a step-by-step install guideQwen3.6 35B AutoRound W4A16 · TP1 research evidence
Measured evidence: 90.91 tok/s on 1× B70 · research
i128/o1024, persistent raw-engine batching, single user, temperature 0, two repeats lab report — documents the result; not a step-by-step install guideWhat has been classified
Dense scorecards summarize every declared combination; measured slices retain exact evidence links.
MTP × TP
Current June 2026 strict-program identity; the legacy TP4 high is shown separately. Fixed: revision=qwen3.6-35b-a3b-base · artifact_id=qwen36-35b-quark-w8a8-cced565 · variant=Quark W8A8 INT8 · runtime=vLLM XPU 2026-06 strict program · graph=on · kv=auto · configured_max_context_tokens=32768.
- ✓ Measured4 cards, no speculative decoding
TP4·MTP093.55 tok/sreport · evidence - ◇ Speed check only2 cards, no speculative decoding
TP2·MTP085.87 tok/sevidence - ⚠ Quarantined4 cards, + speculative decoding (depth 1)
TP4·MTP1— output not usable. Observed: 107.77 tok/s.evidence
12 untested combinations
TP1·MTP0TP1·MTP1TP2·MTP1TP1·MTP2TP2·MTP2TP4·MTP2TP1·MTP3TP2·MTP3TP4·MTP3TP1·MTP4TP2·MTP4TP4·MTP4Quark active context × TP
Canonical active prompt-depth grid. The retained Quark speed rows used p512/o512, and separate long-context quality checks are not relabeled as performance at any canonical depth. Fixed: revision=qwen3.6-35b-a3b-base · artifact_id=qwen36-35b-quark-w8a8-cced565 · variant=Quark W8A8 INT8 · runtime_family=vLLM XPU · mtp=0 · kv=auto.
No classified combinations in this slice yet.
21 untested combinations
TP1·0TP2·0TP4·0TP1·2KTP2·2KTP4·2KTP1·4KTP2·4KTP4·4KTP1·8KTP2·8KTP4·8KTP1·16KTP2·16KTP4·16KTP1·24KTP2·24KTP4·24KTP1·32KTP2·32KTP4·32KAutoRound MTP × TP
AutoRound W4A16 r16, one-user p128/o1024 screen, 4K configured ceiling. Quark results and closures do not transfer into this slice. Fixed: revision=qwen3.6-35b-a3b-base · artifact_id=qwen36-35b-autoround-w4a16 · variant=AutoRound W4A16 INT4 · runtime=vLLM/XPU combined-runtime-guards r16 · graph=piecewise 1/32/64 · kv=auto · configured_max_context_tokens=4096 · natural_eos=False · simultaneous_users=1.
- ◇ Speed check only1 card, no speculative decoding
TP1·MTP0— Direct speed and literal-canary evidence; not a deterministic deployment recipe.90.91 tok/sreport · evidence
14 untested combinations
TP2·MTP0TP4·MTP0TP1·MTP1TP2·MTP1TP4·MTP1TP1·MTP2TP2·MTP2TP4·MTP2TP1·MTP3TP2·MTP3TP4·MTP3TP1·MTP4TP2·MTP4TP4·MTP4AutoRound active context × TP
Canonical active prompt-depth grid. The retained AutoRound run used p128/o1024 with max_model_len=4K, so it is not relabeled as a 0, 2K, or 4K active-context point. Fixed: revision=qwen3.6-35b-a3b-base · artifact_id=qwen36-35b-autoround-w4a16 · variant=AutoRound W4A16 INT4 · runtime_family=vLLM/XPU · mtp=0 · kv=auto.
No classified combinations in this slice yet.
21 untested combinations
TP1·0TP2·0TP4·0TP1·2KTP2·2KTP4·2KTP1·4KTP2·4KTP4·4KTP1·8KTP2·8KTP4·8KTP1·16KTP2·16KTP4·16KTP1·24KTP2·24KTP4·24KTP1·32KTP2·32KTP4·32KMeasured results
Every number links to its proof.
Four-card results: current and earlier
1 = older LocalMaxxing-approved 99.428358 high; 2 = current 93.550542 deep-gate reference. Different gate generations are discrete history, not a regression A/B.
One card serving many users ◇ screened, experimental
Combined rate across all simultaneous users, one Arc Pro B70, AutoRound INT4 · experimental: the 2-16-user dip is the stack's graph-capture boundary, and 64-user determinism is still open
Fine print
Qwen3.6 35B A3B sparse MoE. one Qwen3.6 base weight revision; Quark and AutoRound are quantized deployment artifacts whose evidence stays separate. Measurements, artifact hashes, outputs, quality decisions, and speed stay pinned to their exact recorded identity.
1 combination the lab has stopped pursuing: The TP4 target-only lane is exhausted for now; reopen only for a controlled upstream bakeoff or a new graph-compatible speculative-state mechanism. evidence Exact selectors are in the family data.
Many people at once Lab-measured
64 simultaneous users share 1,039 combined tok/s on 1 card (AutoRound W4A16 INT4); the full curve is under Measured results and in the multi-user report.
How much faster could this get? Projected — not measured
The ML Bottleneck physics engine projects a tuned-run target and the physical ceiling for the headline setup (qwen3.6-35b-a3b-base · Quark W8A8 INT8 · vLLM XPU 2026-06 strict program · TP4 · MTP0 · graph on · auto KV · max context 32,768). The grade is optimization headroom against the tuned-run target, not model quality.
Loading projections from mlbottleneck.com…