Packets and recipes
The deployment variants of this family, at every maturity.
Qwen3.8 AutoRound INT4 · stock a3561ef8 TP scale
Evidence packet · quality-qualified-stock-profile-performance-variable
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · 4 cards, tuned
Evidence packet · accepted-versioned-overlay
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · 2 cards, tuned
Evidence packet · quality-qualified-partial-not-promoted
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP2 b2dd exact HTTP context curve
Evidence packet · quality-battery-qualified-grade-c-http-curve
lab report — documents the result; not a step-by-step install guideQwen3.8 Q4_K_M · TP1
Reproduce 27.83 tok/s on 1× B70 · candidate package
Class-balanced median of per-input-class medians using conventional 99-interval rates on the fixed cold 12-prompt suite; target-only and cache-zero. The all-prompt median is 27.824790 tok/s.Qwen3.8 Q4_K_M + Q4_0 MTP2 · TP1
Reproduce 42.64 tok/s on 1× B70 · candidate package
Median of two fresh-server class-balanced medians over the fixed 12-prompt/six-class, 512-cap native HTTP suite; one B70, TP1, MTP2, F16 target/draft KV, cache zero, 12/12 complete arrays exact between replicas and against same-build MTP0.Qwen3.8 Q8_0 + Q4_0 MTP2 · TP1
Reproduce 37.06 tok/s on 1× B70 · candidate package
Median of two fresh-server class-balanced medians over the fixed 12-prompt/six-class, 512-cap native HTTP suite; one B70, TP1, 1024-token configured context, MTP2, F16 target/draft KV, cache zero, and 24/24 complete MTP2 arrays exact to matched MTP0.Qwen3.8 Q4_K_M · TP1 F16-KV cache64 SYCL-graph HTTP context curve
Evidence packet · quality-battery-qualified-grade-c-graph-http-curve
lab report — documents the result; not a step-by-step install guideQwen3.8 Q4_K_M · TP2
Reproduce 49.72 tok/s on 2× B70 · candidate package
99-interval median across the fixed 12-prompt cache-zero suite; target-only, reasoning off, 12/12 exact output hashes.Qwen3.8 Q4_K_M + Q4_0 MTP2 · TP2
Reproduce 64.24 tok/s on 2× B70 · candidate package
Median of two fresh-server class-balanced medians over the fixed 12-prompt/six-class, 512-cap native HTTP suite; two B70s, equal target split, draft on SYCL0, MTP2, F16 target/draft KV, cache zero, and 24/24 complete candidate arrays exact to the fresh target-only oracle.Qwen3.8 Q8_0 · TP1
Reproduce 19.62 tok/s on 1× B70 · candidate package
Median of two fresh-server class-balanced medians over the full 12-prompt/six-class, 512-cap, cache-zero raw-completion suite; target-only TP1, MTP0, 12/12 complete token arrays exact within TP1 and all objective canaries passed.Qwen3.8 Q8_0 · TP2
Reproduce 36.73 tok/s on 2× B70 · candidate package
Median of two fresh-server class-balanced medians over the full 12-prompt/six-class, 512-cap, cache-zero raw-completion suite; target-only TP2, MTP0, packaged --reasoning off launcher, 12/12 complete token arrays exact and all objective canaries passed.Qwen3.8 FP8 · TP2
Reproduce 86.18 tok/s on 2× B70 · candidate package
Median of two clean-boot fresh-server class-balanced medians over the complete fixed 12-prompt/six-class natural-512 suite (48-78-token prompts), separate empty compile caches, 12/12 complete token arrays exact versus a same-configuration MTP0 oracle (depth-5 target-verified speculative decoding; MTP1 54.935, depth 4 82.396, MTP0 33.097 on the same line), canaries before and after, cache zero; repeat-exact at 100-300-token prompts; output-identical through 16 concurrent users (MTP0: through 64).Qwen3.8 AutoRound INT4 · fixed-K · TP2 MTP4
Evidence packet · research-lossless-profile
Qwen3.8 AutoRound INT4 · TP1 official-image eager/PIECEWISE parity-adjudicated depth
Evidence packet · quality-battery-qualified-grade-c-http-curves-with-8k-piecewise-quarantine
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP1 current-image E4M3-KV eager/PIECEWISE depth
Evidence packet · quality-battery-qualified-grade-c-http-curves
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP1 current-image eager MTP3 partial depth
Evidence packet · mixed-partial-grade-d-human-adjudicated
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP1 current-image eager MTP2 partial depth
Evidence packet · mixed-partial-grade-d-human-adjudicated
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP1 current-image eager MTP1 partial depth
Evidence packet · mixed-partial-grade-d-human-adjudicated
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP1 eager/E4M3 MTP1 exact 4K
Evidence packet · passed-quality-clean-mtp1-e4m3-exact-4k-grade-c
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP1 PIECEWISE/F16 MTP1 exact 4K
Evidence packet · grade-c-exact-4k-mtp1-piecewise
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP1 PIECEWISE/F16 MTP2 exact 4K
Evidence packet · grade-c-exact-4k-mtp2-piecewise-postrun-cache-audited
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP1 current-image eager/F16 MTP4 screened depth
Evidence packet · mixed-partial-grade-d-human-adjudicated
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP2 eager/F16 MTP4 8K quarantine
Evidence packet · quarantined-target-parity-failed
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP2 eager/F16 MTP4 exact 4K
Evidence packet · passed-quality-clean-sentinel-human-adjudicated-grade-c
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP2 eager/F16 MTP4 exact 16K and 24K
Evidence packet · passed-quality-clean-two-depth-expansion-human-adjudicated-grade-c
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP4 eager/F16 MTP4 8K quarantine
Evidence packet · quarantined-target-parity-failed
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP4 current-image eager/F16 MTP4 recovery
Evidence packet · quality-clean-grade-c-recovery
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP4 eager/F16 MTP3 exact 8K
Evidence packet · grade-c-exact-8k-mtp3
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP4 current-image eager/F16 MTP3 partial depth
Evidence packet · partial-grade-c-human-adjudicated
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP4 eager/F16 MTP2 exact 8K
Evidence packet · grade-c-exact-8k-mtp2
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP4 current-image eager/F16 MTP2 partial depth
Evidence packet · partial-grade-c-human-adjudicated
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP4 eager/F16 MTP1 exact 8K
Evidence packet · grade-c-exact-8k-mtp1
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP4 current-image eager/F16 MTP1 partial depth
Evidence packet · partial-grade-c-human-adjudicated
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP4 PIECEWISE/F16 MTP1 exact 4K
Evidence packet · passed-quality-clean-graph-mtp1-exact-4k-human-adjudicated-grade-c
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP2 PIECEWISE/F16 MTP1 exact 4K
Evidence packet · passed-quality-clean-graph-mtp1-exact-4k-human-adjudicated-grade-c
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP2 PIECEWISE/F16 MTP2 exact 4K
Evidence packet · passed-quality-clean-graph-mtp2-exact-4k-human-adjudicated-grade-c
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP2 current-image eager/F16 depth expansion
Evidence packet · quality-battery-qualified-grade-c-http-depth-expansion
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP2 current-image PIECEWISE/F16 partial depth
Evidence packet · partial-grade-c-human-adjudicated
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP2 current-image eager/F16 MTP1 depth expansion
Evidence packet · quality-battery-qualified-grade-c-http-depth-expansion
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP2 current-image eager/F16 MTP2 partial depth
Evidence packet · partial-grade-c-human-adjudicated
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP2 current-image eager/F16 MTP3 partial depth
Evidence packet · quality-battery-qualified-grade-c-parent-scoped
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP4 eager/F16 exact-8K oracle
Evidence packet · grade-c-exact-8k-oracle
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP4 current-image eager/F16 depth expansion
Evidence packet · quality-battery-qualified-grade-c-http-depth-expansion
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP4 current-image PIECEWISE/F16 partial depth
Evidence packet · partial-grade-c-human-adjudicated
lab report — documents the result; not a step-by-step install guideQwen3.8 official FP8 · TP1 eager depth quick check
Evidence packet · grade-c-bounded-fit-depth
lab report — documents the result; not a step-by-step install guideQwen3.8 official FP8 · TP4 PIECEWISE depth
Evidence packet · grade-c-exact-depth
lab report — documents the result; not a step-by-step install guideQwen3.8 Q5_K_S · 256K + vision + MTP
Reproduce 26.67 tok/s on 1× B70 · candidate package
Conventional 99-interval median, cold 12-prompt suite, up to 512-token responses, cache-zero, at 262144 configured capacity with q8_0 K/V, vision mmproj, and MTP draft loaded; active prompts were 48–78 tokens and the second fresh-server run measured 26.640510. Speculation-assisted and labeled as such; target-only tg128 at depth 0 is 22.64 tok/s (raw engine).Qwen3.8 Q5_K_S · TP1 external-MTP 8K route quick check
Evidence packet · grade-c-diagnostic-route-divergence-no-speed-authority
lab report — documents the result; not a step-by-step install guideQwen3.8 Q5_K_S · TP1 no draft model HTTP context curve
Evidence packet · quality-battery-qualified-grade-c-http-curve
lab report — documents the result; not a step-by-step install guideQwen3.8 Q5_K_S · TP1 no draft model F16-KV HTTP context curve
Evidence packet · quality-battery-qualified-grade-c-http-curve
lab report — documents the result; not a step-by-step install guideQwen3.8 Q5_K_S · TP1 F16-KV cache64 SYCL-graph HTTP context curve
Evidence packet · quality-battery-qualified-grade-c-graph-http-curve
lab report — documents the result; not a step-by-step install guideQwen3.8 Q4_K_XL · TP1 no draft model F16-KV HTTP context curve
Evidence packet · quality-battery-qualified-grade-c-http-curve
lab report — documents the result; not a step-by-step install guideQwen3.8 Q4_K_XL · TP1 F16-KV cache64 SYCL-graph HTTP context curve
Evidence packet · quality-and-graph-mechanism-qualified-grade-c-http-curve
lab report — documents the result; not a step-by-step install guideQwen3.8 Q4_K_XL · TP1 no draft model Q8_0-KV HTTP context curve
Evidence packet · quality-battery-qualified-grade-c-http-curve
lab report — documents the result; not a step-by-step install guideQwen3.8 Q8_0 weights · TP1 no draft model F16-KV HTTP context curve
Evidence packet · quality-battery-qualified-grade-c-http-curve
lab report — documents the result; not a step-by-step install guideQwen3.8 Q8_0 weights · TP1 F16-KV cache64 SYCL-graph HTTP context curve
Evidence packet · quality-battery-qualified-grade-c-graph-http-curve
lab report — documents the result; not a step-by-step install guideQwen3.8 Q8_0 weights · TP1 Q8_0-KV cache64 SYCL-graph 8K closure
Evidence packet · closed-bounded-negative-long-quality-crash
lab report — documents the result; not a step-by-step install guideQwen3.8 Q8_0 weights · TP1 no draft model Q8_0-KV HTTP context curve
Evidence packet · quality-battery-qualified-grade-c-http-curve
lab report — documents the result; not a step-by-step install guideQwen3.8 AutoRound INT4 · TP2 + MTP5
Evidence packet · research-only-margin-free-with-determinism-caveat
Qwen3.6 AutoRound INT4 · TP2 + MTP3
Measured evidence: 95.38 tok/s on 2× B70 · historical-record-retained-under-original-narrow-quality-scope
12 unique cold prompts, up to 512 output tokens, cached_tokens=0; historical 100-event compatibility decode accounting lab report — documents the result; not a step-by-step install guideQwen3.6 Q8_0 · TP2
Measured evidence: 36.6 tok/s on 2× B70 · qualified-result
12 unique cold prompts, 512 output tokens each, cached_tokens=0; conventional 99-interval median after TTFT lab report — documents the result; not a step-by-step install guideQwen3.6 Q8_0 · TP1 SYCL graph depth profiles
Evidence packet · quality-qualified-family-research-profile
lab report — documents the result; not a step-by-step install guideQwen3.6 embedded-MTP Q8_0 · TP1 SYCL graph depth profiles
Evidence packet · quality-qualified-family-research-profile-negative-optimization-result
lab report — documents the result; not a step-by-step install guideQwen3.6 embedded-MTP Q8_0 · TP1 MTP1–4 Q8_0-KV quick check
Evidence packet · failed-r1-retained-as-grade-c-exploratory-screen
lab report — documents the result; not a step-by-step install guideQwen3.6 embedded-MTP Q4_K_M · TP1 SYCL graph depth profiles
Evidence packet · quality-qualified-family-research-profile-negative-optimization-result
lab report — documents the result; not a step-by-step install guideQwen3.6 Q8_0 · TP1 + integrated MTP3
Measured evidence: 36.05 tok/s on 1× B70 · qualified-scoped-result-with-long-context-support-evidence
12-prompt fixed cold realistic suite, ordinary EOS, up to 512 output tokens, cached_tokens=0; conventional 99-interval median after TTFT lab report — documents the result; not a step-by-step install guideQwen3.6 runtime FP8 (vLLM native path) · TP2 community recipe
Measured evidence: 30.17 tok/s on 2× B70 · b70-tested-community-recipe
one warmup discarded; 3 passes over five unique prompts; actual prompt lengths 29/30, 60, 210, 510, 2010; 256 output tokens; one active request lab report — documents the result; not a step-by-step install guideQwen3.6 block-FP8 checkpoint · historical XPU requant quick check
Evidence packet · historical-rd-screen-closed
lab report — documents the result; not a step-by-step install guideQwen3.6 UD-Q4_K_XL · intrinsic and policy-tuned MTP
Measured evidence: 31.17 tok/s on 1× B70 · closed-portability-result
fixed 12-prompt cold realistic suite, 128 output tokens, cached_tokens=0; conventional 99-interval median recomputed from the retained historical helper lab report — documents the result; not a step-by-step install guideQwen3.6 Q4_0 · intrinsic MTP3
Measured evidence: 48.43 tok/s on 1× B70 · qualified-scoped-result
fixed 12-prompt cold realistic suite, 128 output tokens, cached_tokens=0; conventional 99-interval median recomputed from the retained historical helper lab report — documents the result; not a step-by-step install guideQwen3.6 Q4_0 · native DFlash5 Q6_K top-1 fusion
Measured evidence: 47.34 tok/s on 1× B70 · closed-record-patch-preserved
fixed 12-prompt cold realistic suite, 128 output tokens, cached_tokens=0; conventional 99-interval median recomputed from the retained historical helper lab report — documents the result; not a step-by-step install guideWhat has been classified
Dense scorecards summarize every declared combination; measured slices retain exact evidence links.
Qwen3.6 1 card · vLLM XPU, no draft model
189 meaningful target-only vLLM XPU cells: three runtime-compatible artifacts × MTP0 × seven exact active-context depths × off/full+piecewise/piecewise graph × f16/E4M3/E5M2 KV. MTP depths are not blindly attached to FP8 artifacts without an exact speculator/package binding.
Fixed: runtime_family=vLLM XPU · runtime_profile_id=qwen36-tp1-vllm-xpu-target-only-v1 · speculation_method=none
Break down by axis
Qwen3.6 TP1 · vLLM XPU AutoRound MTP matrix
252 meaningful native-MTP cells: the measured AutoRound artifact × MTP1–4 × seven exact active-context depths × off/full+piecewise/piecewise graph × f16/E4M3/E5M2 KV. Other artifacts stay outside this profile until their speculator binding is established.
Fixed: runtime_family=vLLM XPU · runtime_profile_id=qwen36-tp1-vllm-xpu-autoround-native-mtp-v1 · speculation_method=native-mtp
Break down by axis
Qwen3.6 TP1 · llama.cpp SYCL MTP matrix
448 meaningful llama.cpp SYCL speculative cells: four MTP-bearing GGUF artifacts × MTP1–4 × seven exact active-context depths × graph off/SYCL × f16/q8_0 KV. MTP0 is represented in the separate target-only profile; vLLM compilation modes and FP8 KV are excluded.
Fixed: runtime_family=llama.cpp SYCL · runtime_profile_id=qwen36-tp1-llamacpp-sycl-embedded-mtp-v1 · speculation_method=embedded-mtp
Break down by axis
Qwen3.6 1 card · llama.cpp SYCL, no draft model
168 meaningful target-only cells: all six Qwen3.6 GGUF target artifacts × MTP0 × seven exact active-context depths × graph off/SYCL × f16/q8_0 KV. Intrinsic MTP1–4 is never manufactured for target-only files; those depths live only in the separately bound MTP-bearing profile.
Fixed: runtime_family=llama.cpp SYCL · runtime_profile_id=qwen36-tp1-llamacpp-sycl-target-only-v1 · speculation_method=none
Break down by axis
Qwen3.8 1 card · vLLM XPU, no draft model
126 meaningful target-only vLLM XPU cells: two runtime-compatible artifacts × MTP0 × seven exact active-context depths × off/full+piecewise/piecewise graph × f16/E4M3/E5M2 KV. MTP depths are not blindly attached to the official FP8 artifact without an exact speculator/package binding.
Fixed: runtime_family=vLLM XPU · runtime_profile_id=qwen38-tp1-vllm-xpu-target-only-v1 · speculation_method=none
Break down by axis
Qwen3.8 TP1 · vLLM XPU AutoRound MTP matrix
252 meaningful native-MTP cells: the measured AutoRound artifact × MTP1–4 × seven exact active-context depths × off/full+piecewise/piecewise graph × f16/E4M3/E5M2 KV. The official FP8 artifact stays outside this profile until its exact speculator binding is established.
Fixed: runtime_family=vLLM XPU · runtime_profile_id=qwen38-tp1-vllm-xpu-autoround-native-mtp-v1 · speculation_method=native-mtp
Break down by axis
Qwen3.8 TP1 · current f01e AutoRound eager MTP4 screened depth
Seven exact current-f01e TP1 native-MTP4 eager F16-KV selectors: 4K/16K/24K are Grade D lab-screened only, 2K and conflicting separate-boot 8K are quarantined, 32K is runtime-closed, and x0 is missing. There are zero lab-measured cells and structural cells expose no speed.
Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-autoround-w4a16-bce40ca · quantization=AutoRound INT4 · tp=1 · mtp=4 · graph_mode=off · kv=f16 · runtime_family=vLLM XPU · runtime_profile_id=qwen38-tp1-vllm-xpu-autoround-native-mtp4-v1 · speculation_method=native-mtp
Break down by axis
Qwen3.8 pinned e9d1398 · AutoRound decision closure
The original 96-cell TP/MTP/graph/KV campaign, pinned only to vLLM XPU nightly e9d1398d9. Measured target and TP1 eager-MTP anchors remain measured; unsupported, closed, and quarantined descendants are explicit. Nothing in this contract transfers to f01e or another runtime.
Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-autoround-w4a16-bce40ca · runtime=vLLM XPU nightly e9d1398d9 · configured_max_context_tokens=32768
Break down by axis
Qwen3.8 1 card · llama.cpp SYCL, no draft model
112 meaningful target-only cells: four GGUF artifacts × MTP0 × seven exact active-context depths × graph off/SYCL × f16/q8_0 KV. These cells preserve 84 exact measurements and one speedless exact-selector closure. All 14 superseded live estimates are retired; their frozen snapshots remain data-only calibration history, and no MTP speed is transferred.
Fixed: runtime_family=llama.cpp SYCL
Break down by axis
Qwen3.8 TP1 · llama.cpp SYCL MTP-package matrix
224 meaningful package cells: the two targets pinned with the inspected Unsloth Q4_0 MTP draft × MTP1–4 × seven exact active-context depths × graph off/SYCL × f16/q8_0 KV. This is an exact package backlog; no Qwen3.6 speed or quality transfers.
Fixed: runtime_family=llama.cpp SYCL · runtime_profile_id=qwen38-tp1-llamacpp-sycl-external-mtp-v1 · speculation_method=external-mtp-draft · speculator_artifact_id=qwen38-27b-mtp-q4-0-50d9ce5
Break down by axis
Qwen3.8 TP2 · llama.cpp SYCL Q4_K_M exact HTTP depth
Seven exact selector cells for the pinned Q4_K_M TP2 target-only F16-KV graph-off HTTP lane: six Grade C cache-zero measurements at 2K-32K and an explicit missing x=0 cell.
Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-ggmlorg-q4-k-m-0669b98 · tp=2 · mtp=0 · graph_mode=off · kv=f16 · runtime_family=llama.cpp SYCL · runtime_profile_id=qwen38-tp2-llamacpp-sycl-q4km-http-depth-v1 · speculation_method=none
Break down by axis
Qwen3.8 TP2 · llama.cpp SYCL Q8_0 exact HTTP depth
Seven exact selector cells for the pinned Q8_0-weight TP2 target-only F16-KV graph-off HTTP lane: six Grade C cache-zero measurements at 2K-32K and an explicit missing x=0 cell.
Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-ggmlorg-q8-0-0669b98 · tp=2 · mtp=0 · graph_mode=off · kv=f16 · runtime_family=llama.cpp SYCL · runtime_profile_id=qwen38-tp2-llamacpp-sycl-q8-http-depth-v1 · speculation_method=none
Break down by axis
Qwen3.8 TP2 · vLLM XPU official FP8 exact HTTP depth
Seven exact selector cells for the pinned official-FP8 TP2 target-only F16-KV PIECEWISE HTTP lane: six Grade C cache-zero measurements at 2K-32K and an explicit missing x=0 cell.
Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-official-fp8-017b9c7 · tp=2 · mtp=0 · graph_mode=PIECEWISE · kv=f16 · runtime_family=vLLM XPU · runtime_profile_id=qwen38-tp2-vllm-xpu-official-fp8-http-depth-v1 · speculation_method=none
Break down by axis
Qwen3.8 TP2 · vLLM XPU official FP8/W8A16 deterministic MTP1 exact HTTP depth
Seven exact selector cells for the qualified official-FP8/W8A16 TP2 deterministic-MTP1 F16-KV graph-off HTTP lane: six Grade C cache-zero, target-array-exact measurements at 2K-32K and an explicit missing x=0 cell.
Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-official-fp8-017b9c7 · tp=2 · mtp=1 · graph_mode=off · kv=f16 · runtime_family=vLLM XPU · runtime_profile_id=qwen38-tp2-vllm-xpu-official-fp8-w8a16-mtp1-http-depth-r33 · speculation_method=publisher MTP depth 1
Break down by axis
Qwen3.8 TP4 · vLLM XPU official FP8 exact HTTP depth
Seven exact selector cells for the pinned official-FP8 TP4 target-only F16-KV PIECEWISE HTTP lane: six Grade C cache-zero measurements at 2K-32K and an explicit missing x=0 cell. This slower profile does not replace the AutoRound TP4 frontier.
Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-official-fp8-017b9c7 · tp=4 · mtp=0 · graph_mode=PIECEWISE · kv=f16 · runtime_family=vLLM XPU · runtime_profile_id=qwen38-tp4-vllm-xpu-official-fp8-http-depth-v1 · speculation_method=none
Break down by axis
Qwen3.8 TP2 · vLLM XPU AutoRound exact HTTP depth
Seven exact selector cells for the pinned AutoRound INT4 TP2 target-only MTP0 F16-KV FULL_AND_PIECEWISE HTTP lane: six Grade C cache-zero measurements at 2K-32K and an explicit missing x=0 cell.
Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-autoround-w4a16-bce40ca · quantization=AutoRound INT4 · tp=2 · mtp=0 · graph_mode=FULL_AND_PIECEWISE · kv=f16 · runtime_family=vLLM XPU · runtime_profile_id=qwen38-tp2-vllm-xpu-autoround-http-depth-v1 · speculation_method=none
Break down by axis
Qwen3.8 TP4 · vLLM XPU AutoRound exact HTTP depth
Seven exact selector cells for the pinned AutoRound INT4 TP4 target-only MTP0 F16-KV FULL_AND_PIECEWISE HTTP lane: six Grade C cache-zero measurements at 2K-32K and an explicit missing x=0 cell.
Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-autoround-w4a16-bce40ca · quantization=AutoRound INT4 · tp=4 · mtp=0 · graph_mode=FULL_AND_PIECEWISE · kv=f16 · runtime_family=vLLM XPU · runtime_profile_id=qwen38-tp4-vllm-xpu-autoround-http-depth-v1 · speculation_method=none
Break down by axis
Qwen3.8 TP2 · current f01e AutoRound eager target depth
Seven exact selector cells for current-f01e AutoRound INT4 TP2 target-only MTP0 eager F16-KV: six Grade C cache-zero nonzero measurements with exact TP1 comparison parity and explicit missing x=0.
Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-autoround-w4a16-bce40ca · quantization=AutoRound INT4 · tp=2 · mtp=0 · graph_mode=off · kv=f16 · runtime_family=vLLM XPU · runtime_profile_id=qwen38-tp2-vllm-xpu-autoround-eager-target-v1 · speculation_method=none
Break down by axis
Qwen3.8 TP2 · current f01e AutoRound PIECEWISE partial target depth
Seven exact current-f01e AutoRound INT4 TP2 target-only MTP0 PIECEWISE F16-KV selectors: four Grade C target-identical measurements at 2K/4K/24K/32K, exact 8K/16K speedless target-divergence quarantines, and explicit missing x0. The dated FULL_AND_PIECEWISE graph contract remains separate.
Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-autoround-w4a16-bce40ca · quantization=AutoRound INT4 · tp=2 · mtp=0 · graph_mode=PIECEWISE · kv=f16 · runtime_family=vLLM XPU · runtime_profile_id=qwen38-tp2-vllm-xpu-autoround-f01e-piecewise-target-v1 · speculation_method=none
Break down by axis
Qwen3.8 TP2 · current f01e AutoRound PIECEWISE MTP1 depth
Seven exact selector cells for current-f01e AutoRound INT4 TP2 native-MTP1 PIECEWISE F16-KV: one Grade C cache-zero exact-4K measurement with positive isolated acceptance, dual-parent token parity, full quality, graph/topology/cache/model verification, and six explicit missing depths.
Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-autoround-w4a16-bce40ca · quantization=AutoRound INT4 · tp=2 · mtp=1 · graph_mode=PIECEWISE · kv=f16 · runtime_family=vLLM XPU · runtime_profile_id=qwen38-tp2-vllm-xpu-autoround-native-mtp1-piecewise-v1 · speculation_method=native-mtp
Break down by axis
Qwen3.8 TP2 · current f01e AutoRound PIECEWISE MTP2 depth
Seven exact selector cells for current-f01e AutoRound INT4 TP2 native-MTP2 PIECEWISE F16-KV: one Grade C cache-zero exact-4K measurement with positive isolated acceptance, dual-parent token parity, full quality, graph/topology/cache/model verification, and six explicit missing depths.
Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-autoround-w4a16-bce40ca · quantization=AutoRound INT4 · tp=2 · mtp=2 · graph_mode=PIECEWISE · kv=f16 · runtime_family=vLLM XPU · runtime_profile_id=qwen38-tp2-vllm-xpu-autoround-native-mtp2-piecewise-v1 · speculation_method=native-mtp
Break down by axis
Qwen3.8 TP2 · current f01e AutoRound eager MTP1 depth
Seven exact selector cells for current-f01e AutoRound INT4 TP2 native-MTP1 eager F16-KV: six Grade C cache-zero nonzero measurements with positive isolated acceptance, exact same-topology TP2/MTP0 target parity, and explicit missing x=0.
Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-autoround-w4a16-bce40ca · quantization=AutoRound INT4 · tp=2 · mtp=1 · graph_mode=off · kv=f16 · runtime_family=vLLM XPU · runtime_profile_id=qwen38-tp2-vllm-xpu-autoround-native-mtp1-v1 · speculation_method=native-mtp
Break down by axis
Qwen3.8 TP2 · current f01e AutoRound eager MTP2 partial depth
Seven exact selector cells for current-f01e AutoRound INT4 TP2 native-MTP2 eager F16-KV: five Grade C cache-zero target-parity measurements at 4K-32K, one exact/cache-zero/positive-acceptance 2K quarantine after token-90 target divergence, and explicit missing x=0.
Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-autoround-w4a16-bce40ca · quantization=AutoRound INT4 · tp=2 · mtp=2 · graph_mode=off · kv=f16 · runtime_family=vLLM XPU · runtime_profile_id=qwen38-tp2-vllm-xpu-autoround-native-mtp2-v1 · speculation_method=native-mtp
Break down by axis
Qwen3.8 TP2 · current f01e AutoRound eager MTP3 parent-scoped depth
Seven exact selector cells for current-f01e AutoRound INT4 TP2 native-MTP3 eager F16-KV: five Grade C cache-zero target-parity measurements at 4K-32K, one 2K cell structurally excluded by the parent MTP2 token-90 quarantine with no MTP3 speed, and explicit missing x=0.
Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-autoround-w4a16-bce40ca · quantization=AutoRound INT4 · tp=2 · mtp=3 · graph_mode=off · kv=f16 · runtime_family=vLLM XPU · runtime_profile_id=qwen38-tp2-vllm-xpu-autoround-native-mtp3-v1 · speculation_method=native-mtp
Break down by axis
Qwen3.8 TP2 · current f01e AutoRound eager MTP4 sentinel depth
Seven exact current-f01e AutoRound INT4 TP2 native-MTP4 eager F16-KV selectors: exact 4K, 16K, and 24K are Grade C target-parity measurements, exact 8K is a structural target-parity quarantine with no site speed, and x0/2K/32K remain missing. No value is inferred from another MTP depth or topology.
Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-autoround-w4a16-bce40ca · quantization=AutoRound INT4 · tp=2 · mtp=4 · graph_mode=off · kv=f16 · runtime_family=vLLM XPU · runtime_profile_id=qwen38-tp2-vllm-xpu-autoround-native-mtp4-v1 · speculation_method=native-mtp
Break down by axis
Qwen3.8 TP4 · current f01e AutoRound eager oracle depth
Seven exact selector cells for the current-f01e AutoRound INT4 TP4 target-only MTP0 eager F16-KV lane: six Grade C cache-zero nonzero measurements and explicit missing x=0. The earlier exact-8K cell is retained; five cells are additive from the expansion.
Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-autoround-w4a16-bce40ca · quantization=AutoRound INT4 · tp=4 · mtp=0 · graph_mode=off · kv=f16 · runtime_family=vLLM XPU · runtime_profile_id=qwen38-tp4-vllm-xpu-autoround-eager-oracle-v1 · speculation_method=none
Break down by axis
Qwen3.8 TP4 · current f01e AutoRound PIECEWISE partial target depth
Seven exact current-f01e AutoRound INT4 TP4 target-only MTP0 PIECEWISE F16-KV selectors: five Grade C target-identical measurements at 2K/4K/16K/24K/32K, exact 8K speedless target-divergence quarantine, and explicit missing x0. The dated FULL_AND_PIECEWISE graph contract remains separate.
Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-autoround-w4a16-bce40ca · quantization=AutoRound INT4 · tp=4 · mtp=0 · graph_mode=PIECEWISE · kv=f16 · runtime_family=vLLM XPU · runtime_profile_id=qwen38-tp4-vllm-xpu-autoround-f01e-piecewise-target-v1 · speculation_method=none
Break down by axis
Qwen3.8 TP4 · current f01e AutoRound PIECEWISE MTP1 depth
Seven exact current-f01e AutoRound INT4 TP4 native-MTP1 PIECEWISE F16-KV selectors: exact 4K is a Grade C dual-parent-parity measurement and x0/2K/8K/16K/24K/32K remain missing. The known current-profile 8K graph-corruption signature is retained as a caveat, not transferred into an unrun MTP1 cell.
Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-autoround-w4a16-bce40ca · quantization=AutoRound INT4 · tp=4 · mtp=1 · graph_mode=PIECEWISE · kv=f16 · runtime_family=vLLM XPU · runtime_profile_id=qwen38-tp4-vllm-xpu-autoround-native-mtp1-piecewise-v1 · speculation_method=native-mtp
Break down by axis
Qwen3.8 TP4 · current f01e AutoRound eager MTP1 depth
Seven exact selector cells for current-f01e AutoRound INT4 TP4 native-MTP1 eager F16-KV: five Grade C target-parity measurements at 4K-32K, 2K quarantined at token 90, and x0 missing. The prior 8K cell is retained unchanged.
Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-autoround-w4a16-bce40ca · quantization=AutoRound INT4 · tp=4 · mtp=1 · graph_mode=off · kv=f16 · runtime_family=vLLM XPU · runtime_profile_id=qwen38-tp4-vllm-xpu-autoround-native-mtp1-v1 · speculation_method=native-mtp
Break down by axis
Qwen3.8 TP4 · current f01e AutoRound eager MTP2 depth
Seven exact current-f01e TP4 native-MTP2 eager F16-KV selectors: five Grade C target-parity measurements at 4K-32K, 2K quarantined at token 90, and x0 missing; prior 8K retained.
Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-autoround-w4a16-bce40ca · quantization=AutoRound INT4 · tp=4 · mtp=2 · graph_mode=off · kv=f16 · runtime_family=vLLM XPU · runtime_profile_id=qwen38-tp4-vllm-xpu-autoround-native-mtp2-v1 · speculation_method=native-mtp
Break down by axis
Qwen3.8 TP4 · current f01e AutoRound eager MTP3 depth
Seven exact current-f01e TP4 native-MTP3 eager F16-KV selectors: five Grade C target-parity measurements at 4K-32K, 2K quarantined at token 90, and x0 missing; prior 8K retained.
Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-autoround-w4a16-bce40ca · quantization=AutoRound INT4 · tp=4 · mtp=3 · graph_mode=off · kv=f16 · runtime_family=vLLM XPU · runtime_profile_id=qwen38-tp4-vllm-xpu-autoround-native-mtp3-v1 · speculation_method=native-mtp
Break down by axis
Qwen3.8 TP4 · current f01e AutoRound eager MTP4 depth
Seven exact current-f01e TP4 native-MTP4 eager F16-KV selectors: 4K/16K/24K are Grade C target-parity measurements, 2K/8K are quarantined target divergences, 32K is a runtime-fatal closure, and x0 is missing. Structural cells expose no speed.
Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-autoround-w4a16-bce40ca · quantization=AutoRound INT4 · tp=4 · mtp=4 · graph_mode=off · kv=f16 · runtime_family=vLLM XPU · runtime_profile_id=qwen38-tp4-vllm-xpu-autoround-native-mtp4-v1 · speculation_method=native-mtp
Break down by axis
Qwen3.8 TP4 · vLLM XPU AutoRound fully checked snapshot
One exact short-suite TP4 anchor for the frozen b2dd9ce73d + 1e90ffa672 source stack. This is not an active-context curve: 32K is configured capacity only, and no x=0/2K/4K/8K/16K/24K/32K throughput cell is claimed.
Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-autoround-w4a16-bce40ca · quantization=AutoRound INT4 · runtime_family=vLLM XPU · runtime=vLLM XPU b2dd9ce73d + XPU kernels 1e90ffa672 · measurement_class=strict-natural-eos · workload=two strict natural-EOS replays; 25 unique cold prompts; up to 512 output tokens; cached_tokens=0; conventional 99-interval median · tp=4 · mtp=0 · graph=on · graph_mode=FULL_AND_PIECEWISE · kv=f16 · configured_max_context_tokens=32768 · gpu_memory_utilization=0.6 · natural_eos=True
Break down by axis
Qwen3.6 Q4_0 MTP × TP
Qwen3.6 Q4_0 intrinsic-MTP artifact · llama.cpp SYCL · graph off · q8_0 target KV · fixed cold/cache-zero suite. Every MTP0–5 × TP1/2/3/4 cell is explicit. Fixed: revision=qwen3.6-27b · artifact_id=qwen36-27b-unsloth-mtp-q4-0-20c9c45 · quantization=Q4_0 · runtime_family=llama.cpp SYCL · profile_id=q36-q4-0-tp1-intrinsic-mtp-v1 · graph=off · kv=q8_0.
- ✓ Measured1 card, no speculative decoding
TP1·MTP025.68 tok/sreport · evidence - ✓ Measured1 card, + speculative decoding (depth 3)
TP1·MTP348.43 tok/sreport · evidence
22 untested combinations
TP2·MTP0TP3·MTP0TP4·MTP0TP1·MTP1TP2·MTP1TP3·MTP1TP4·MTP1TP1·MTP2TP2·MTP2TP3·MTP2TP4·MTP2TP2·MTP3TP3·MTP3TP4·MTP3TP1·MTP4TP2·MTP4TP3·MTP4TP4·MTP4TP1·MTP5TP2·MTP5TP3·MTP5TP4·MTP5Qwen3.6 UD-Q4_K_XL MTP × TP
Qwen3.6 UD-Q4_K_XL · llama.cpp SYCL · intrinsic MTP with n_min=0/p_min=0 · f16 KV · graph requested but replay unproven · fixed cold/cache-zero suite. Fixed: revision=qwen3.6-27b · artifact_id=qwen36-27b-unsloth-mtp-ud-q4-k-xl-4085665 · quantization=UD-Q4_K_XL · runtime_family=llama.cpp SYCL · profile_id=q36-udq4kxl-tp1-intrinsic-mtp-depth-v1 · graph=requested-not-proven-replay · kv=f16.
- ✓ Measured1 card, no speculative decoding
TP1·MTP023.44 tok/sreport · evidence - ✓ Measured1 card, + speculative decoding (depth 3)
TP1·MTP329.22 tok/sreport · evidence - ✓ Measured1 card, + speculative decoding (depth 4)
TP1·MTP428.31 tok/sreport · evidence - ✓ Measured1 card, + speculative decoding (depth 5)
TP1·MTP524.66 tok/sreport · evidence
20 untested combinations
TP2·MTP0TP3·MTP0TP4·MTP0TP1·MTP1TP2·MTP1TP3·MTP1TP4·MTP1TP1·MTP2TP2·MTP2TP3·MTP2TP4·MTP2TP2·MTP3TP3·MTP3TP4·MTP3TP2·MTP4TP3·MTP4TP4·MTP4TP2·MTP5TP3·MTP5TP4·MTP5Qwen3.6 Q4_0 speculation × TP
Qwen3.6 Q4_0 target on llama.cpp SYCL with graph off and q8_0 target KV. Rows distinguish target-only, bundled intrinsic MTP3, and a separate native Q8_0 DFlash5 draft. Fixed: revision=qwen3.6-27b · artifact_id=qwen36-27b-unsloth-mtp-q4-0-20c9c45 · quantization=Q4_0 · runtime_family=llama.cpp SYCL · graph=off · kv=q8_0.
- ✓ Measured1 card, target only
TP1·target only25.68 tok/sreport · evidence - ✓ Measured1 card, intrinsic MTP3
TP1·intrinsic MTP348.43 tok/sreport · evidence - ✓ Measured1 card, native DFlash5
TP1·native DFlash547.34 tok/sreport · evidence
9 untested combinations
TP2·target onlyTP3·target onlyTP4·target onlyTP2·intrinsic MTP3TP3·intrinsic MTP3TP4·intrinsic MTP3TP2·native DFlash5TP3·native DFlash5TP4·native DFlash5Qwen3.6 runtime FP8 × TP
Qwen3.6 base weights quantized at runtime through the vLLM native-FP8 path, replayed in the reference lab with 262,144 configured max, eager vLLM XPU, and fp8_e4m3 KV. This is not a prequantized FP8 checkpoint. Fixed: revision=qwen3.6-27b · artifact_id=qwen36-27b-base-runtime-fp8-6a9e13b · runtime_family=vLLM XPU · mtp=0 · graph=off · kv=fp8_e4m3 · configured_max_context_tokens=262144.
3 untested combinations
TP1·FP8TP3·FP8TP4·FP8Qwen3.6 Q8 MTP × TP
Qwen3.6 Q8_0 llama.cpp SYCL evidence with graph disabled and f16 KV. Every MTP0–5 × TP1/2/3/4 cell is explicit; independent replicas are not mislabeled as tensor parallel. Fixed: revision=qwen3.6-27b · quantization=Q8_0 · runtime_family=llama.cpp SYCL · graph=off · kv=f16.
- ✓ Measured1 card, no speculative decoding
TP1·MTP017.11 tok/sevidence - ✓ Measured2 cards, no speculative decoding
TP2·MTP036.6 tok/sreport · evidence - ✓ Measured1 card, + speculative decoding (depth 3)
TP1·MTP336.05 tok/sreport · evidence
21 untested combinations
TP3·MTP0TP4·MTP0TP1·MTP1TP2·MTP1TP3·MTP1TP4·MTP1TP1·MTP2TP2·MTP2TP3·MTP2TP4·MTP2TP2·MTP3TP3·MTP3TP4·MTP3TP1·MTP4TP2·MTP4TP3·MTP4TP4·MTP4TP1·MTP5TP2·MTP5TP3·MTP5TP4·MTP5Qwen3.6 Q8 speculation × serving slots
Qwen3.6 Q8_0 TP1 llama.cpp SYCL at 32K configured capacity per slot. The target c1 cell uses official isolated two-prompt band evidence; target-only c2 was actually launched near 32K, while embedded MTP3 c2 stopped at the safe-fit gate. Fixed: revision=qwen3.6-27b · quantization=Q8_0 · runtime_family=llama.cpp SYCL · tp=1 · graph=off · kv=f16 · context_tokens_per_slot=32768.
- ✓ Measuredc1, target only
c1·target onlyD15.077 @17K / 13.686 @31.8K · D511evidence - ✓ Measuredc2, target only
c2·target only10.14 tok/sevidence - ✓ Measuredc1, embedded MTP3
c1·embedded MTP3D44.49 @17K / 41.49 @31.8Kevidence - ■ Closedc2, embedded MTP3
c2·embedded MTP3— A second target and draft KV plus recurrent state exceed safe B70 capacity before useful headroom; retry only with a materially lower-memory, separately preregistered identity.32,683 MiB fit no-goevidence
Codes in the rows: D = decode tok/s · P = prefill tok/s · T = ms to first token · AR = share of drafted tokens accepted · A = combined tok/s.
Qwen3.6 AutoRound MTP × TP
Qwen3.6 Webhie AutoRound INT4 with the BF16-scale INT8 target LM-head, graph enabled, and f16 KV. Cells show normalized historical evidence already in the lab; they do not imply that dated runtimes and suites are directly rankable. Fixed: revision=qwen3.6-27b · artifact_id=qwen36-27b-autoround-w4a16-f5750c9 · quantization=AutoRound INT4 · runtime_family=vLLM XPU · graph=on · kv=f16.
- ✓ Measured1 card, + speculative decoding (depth 1)
TP1·MTP151.25 tok/sevidence - ✓ Measured1 card, + speculative decoding (depth 2)
TP1·MTP259.59 tok/sevidence - ✓ Measured1 card, + speculative decoding (depth 3)
TP1·MTP368.24 tok/sevidence - ✓ Measured2 cards, + speculative decoding (depth 3)
TP2·MTP395.38 tok/sevidence - ✓ Measured1 card, + speculative decoding (depth 4)
TP1·MTP459.89 tok/sevidence - ✓ Measured1 card, + speculative decoding (depth 5)
TP1·MTP5— Both strict/cache-zero MTP5 capture-size arms are recorded; neither beat the MTP3 control, so the depth/config-only lane is closed.59.82 tok/sevidence - × Unsupported3 cards, no speculative decoding
TP3·MTP0— Qwen3.6 and Qwen3.8 share this architecture; the vLLM tensor-parallel route cannot evenly partition 16 GDN key heads across TP3.16 heads cannot split by 3evidence - × Unsupported3 cards, + speculative decoding (depth 1)
TP3·MTP1— Qwen3.6 and Qwen3.8 share this architecture; the vLLM tensor-parallel route cannot evenly partition 16 GDN key heads across TP3.16 heads cannot split by 3evidence - × Unsupported3 cards, + speculative decoding (depth 2)
TP3·MTP2— Qwen3.6 and Qwen3.8 share this architecture; the vLLM tensor-parallel route cannot evenly partition 16 GDN key heads across TP3.16 heads cannot split by 3evidence - × Unsupported3 cards, + speculative decoding (depth 3)
TP3·MTP3— Qwen3.6 and Qwen3.8 share this architecture; the vLLM tensor-parallel route cannot evenly partition 16 GDN key heads across TP3.16 heads cannot split by 3evidence - × Unsupported3 cards, + speculative decoding (depth 4)
TP3·MTP4— Qwen3.6 and Qwen3.8 share this architecture; the vLLM tensor-parallel route cannot evenly partition 16 GDN key heads across TP3.16 heads cannot split by 3evidence - × Unsupported3 cards, + speculative decoding (depth 5)
TP3·MTP5— Qwen3.6 and Qwen3.8 share this architecture; the vLLM tensor-parallel route cannot evenly partition 16 GDN key heads across TP3.16 heads cannot split by 3evidence
12 untested combinations
TP1·MTP0TP2·MTP0TP4·MTP0TP2·MTP1TP4·MTP1TP2·MTP2TP4·MTP2TP4·MTP3TP2·MTP4TP4·MTP4TP2·MTP5TP4·MTP5pinned e9d1398 · graph off
Qwen3.8 AutoRound · pinned vLLM XPU e9d1398 · f16 KV · short fixed-output diagnostic. TP2/TP4 MTP1/3 were closed by the frozen MTP2 expansion gate; MTP4/5 were outside the 96-cell program and remain missing.
- ✓ Measured1 card, no speculative decoding
TP1·MTP024.25 tok/sevidence - ✓ Measured2 cards, no speculative decoding
TP2·MTP016.77 tok/sevidence - ✓ Measured4 cards, no speculative decoding
TP4·MTP017.38 tok/sevidence - ✓ Measured1 card, + speculative decoding (depth 1)
TP1·MTP14.51 tok/sevidence - ✓ Measured1 card, + speculative decoding (depth 2)
TP1·MTP24.41 tok/sevidence - ✓ Measured4 cards, + speculative decoding (depth 2)
TP4·MTP231.17 tok/sevidence - ✓ Measured1 card, + speculative decoding (depth 3)
TP1·MTP34.3 tok/sevidence - ◇ Speed check only2 cards, + speculative decoding (depth 2)
TP2·MTP2boot+canary onlyevidence - ■ Closed2 cards, + speculative decoding (depth 1)
TP2·MTP1gateevidence - ■ Closed4 cards, + speculative decoding (depth 1)
TP4·MTP1gateevidence - ■ Closed2 cards, + speculative decoding (depth 3)
TP2·MTP3MTP2 gateevidence - ■ Closed4 cards, + speculative decoding (depth 3)
TP4·MTP3MTP2 gateevidence - × Unsupported3 cards, no speculative decoding
TP3·MTP0head splitevidence - × Unsupported3 cards, + speculative decoding (depth 1)
TP3·MTP1head splitevidence - × Unsupported3 cards, + speculative decoding (depth 2)
TP3·MTP2head splitevidence - × Unsupported3 cards, + speculative decoding (depth 3)
TP3·MTP3head splitevidence - × Unsupported3 cards, + speculative decoding (depth 4)
TP3·MTP4head splitevidence - × Unsupported3 cards, + speculative decoding (depth 5)
TP3·MTP5head splitevidence
6 untested combinations
TP1·MTP4TP2·MTP4TP4·MTP4TP1·MTP5TP2·MTP5TP4·MTP5pinned e9d1398 · graph on
Qwen3.8 AutoRound · pinned vLLM XPU e9d1398 · f16 KV · short fixed-output diagnostic. Multi-GPU XPU Graph is runtime-labeled unsupported/experimental; the e9d1398 TP1 MTP1 corruption anchor quarantined MTP1/2/3 in the bounded program. MTP4/5 remain missing.
- ✓ Measured1 card, no speculative decoding
TP1·MTP030.26 tok/sevidence - ✓ Measured2 cards, no speculative decoding
TP2·MTP048.95 tok/sevidence - ✓ Measured4 cards, no speculative decoding
TP4·MTP071.67 tok/sevidence - × Unsupported3 cards, no speculative decoding
TP3·MTP0head splitevidence - × Unsupported3 cards, + speculative decoding (depth 1)
TP3·MTP1head splitevidence - × Unsupported3 cards, + speculative decoding (depth 2)
TP3·MTP2head splitevidence - × Unsupported3 cards, + speculative decoding (depth 3)
TP3·MTP3head splitevidence - × Unsupported3 cards, + speculative decoding (depth 4)
TP3·MTP4head splitevidence - × Unsupported3 cards, + speculative decoding (depth 5)
TP3·MTP5head splitevidence - ⚠ Quarantined1 card, + speculative decoding (depth 1)
TP1·MTP1— output not usable. Observed: 7.63 tok/s.evidence - ⚠ Quarantined2 cards, + speculative decoding (depth 1)
TP2·MTP1— output not usable. Observed: graph+MTP.evidence - ⚠ Quarantined4 cards, + speculative decoding (depth 1)
TP4·MTP1— output not usable. Observed: graph+MTP.evidence - ⚠ Quarantined1 card, + speculative decoding (depth 2)
TP1·MTP2— output not usable. Observed: MTP1 anchor.evidence - ⚠ Quarantined2 cards, + speculative decoding (depth 2)
TP2·MTP2— output not usable. Observed: graph+MTP.evidence - ⚠ Quarantined4 cards, + speculative decoding (depth 2)
TP4·MTP2— output not usable. Observed: graph+MTP.evidence - ⚠ Quarantined1 card, + speculative decoding (depth 3)
TP1·MTP3— output not usable. Observed: MTP1 anchor.evidence - ⚠ Quarantined2 cards, + speculative decoding (depth 3)
TP2·MTP3— output not usable. Observed: graph+MTP.evidence - ⚠ Quarantined4 cards, + speculative decoding (depth 3)
TP4·MTP3— output not usable. Observed: graph+MTP.evidence
6 untested combinations
TP1·MTP4TP2·MTP4TP4·MTP4TP1·MTP5TP2·MTP5TP4·MTP5dated current-main snapshots × TP
Qwen3.8 AutoRound · exact dated vLLM/XPU-kernel source stacks · strict natural-EOS profile · MTP0 · f16 KV · graph on. Each row is immutable historical evidence, never a moving 'current' alias. Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-autoround-w4a16-bce40ca · variant=AutoRound INT4 · runtime_family=vLLM XPU · measurement_class=strict-natural-eos · mtp=0 · graph=on · kv=f16 · configured_max_context_tokens=32768 · natural_eos=True.
- ✓ Measured1 card, 0ecc + stock kernel
TP1·0ecc + stock kernel30.33 tok/sevidence - ✓ Measured1 card, 0ecc + baaa
TP1·0ecc + baaa30.28 tok/sevidence - ✓ Measured1 card, 6a9 + baaa
TP1·6a9 + baaa30.27 tok/sevidence - ✓ Measured1 card, b2dd + 1e90
TP1·b2dd + 1e9030.28 tok/sevidence - ✓ Measured4 cards, b2dd + 1e90
TP4·b2dd + 1e90— All model and quality arms passed; deterministic aggregation was recovered offline after a mawk reserved-name wrapper failure.71.83 tok/sevidence - × Unsupported3 cards, 0ecc + stock kernel
TP3·0ecc + stock kernelhead split - × Unsupported3 cards, 0ecc + baaa
TP3·0ecc + baaahead split - × Unsupported3 cards, 6a9 + baaa
TP3·6a9 + baaahead split - × Unsupported3 cards, b2dd + 1e90
TP3·b2dd + 1e90head split
7 untested combinations
TP2·0ecc + stock kernelTP4·0ecc + stock kernelTP2·0ecc + baaaTP4·0ecc + baaaTP2·6a9 + baaaTP4·6a9 + baaaTP2·b2dd + 1e90stock a3561ef8 · graph on
Qwen3.8 AutoRound · stock vLLM XPU a3561ef8 · strict natural-EOS profile · f16 KV · graph on. No MTP1–5 value is inferred from e9d1398; diagnostic captures remain in their separate graph. Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-autoround-w4a16-bce40ca · variant=AutoRound INT4 · runtime=vLLM XPU nightly a3561ef8 stock · profile_id=a3561ef8-stock · measurement_class=strict-natural-eos · graph=on · kv=f16 · configured_max_context_tokens=32768 · natural_eos=True.
- ✓ Measured1 card, no speculative decoding
TP1·MTP030.24 tok/sevidence - ✓ Measured2 cards, no speculative decoding
TP2·MTP048.49 tok/sevidence - ✓ Measured4 cards, no speculative decoding
TP4·MTP0— The 71.900 strict capture passed quality but did not replicate and is not a promoted record.71.9 tok/sevidence - × Unsupported3 cards, no speculative decoding
TP3·MTP0head splitevidence - × Unsupported3 cards, + speculative decoding (depth 1)
TP3·MTP1head splitevidence - × Unsupported3 cards, + speculative decoding (depth 2)
TP3·MTP2head splitevidence - × Unsupported3 cards, + speculative decoding (depth 3)
TP3·MTP3head splitevidence - × Unsupported3 cards, + speculative decoding (depth 4)
TP3·MTP4head splitevidence - × Unsupported3 cards, + speculative decoding (depth 5)
TP3·MTP5head splitevidence
15 untested combinations
TP1·MTP1TP2·MTP1TP4·MTP1TP1·MTP2TP2·MTP2TP4·MTP2TP1·MTP3TP2·MTP3TP4·MTP3TP1·MTP4TP2·MTP4TP4·MTP4TP1·MTP5TP2·MTP5TP4·MTP5a3561ef8 TP2 overlay · graph on
Qwen3.8 AutoRound · vLLM XPU a3561ef8 plus the TP2 historical-winner decisions · strict natural-EOS profile · MTP0 · f16 KV · graph on. This is measured partial recovery and not promoted. Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-autoround-w4a16-bce40ca · variant=AutoRound INT4 · runtime=vLLM XPU nightly a3561ef8 + TP2 historical-winner overlay · profile_id=a3561ef8-tp2-e9d1398-best-config-overlay-v1 · measurement_class=strict-natural-eos · graph=on · kv=f16 · configured_max_context_tokens=32768 · natural_eos=True.
- ✓ Measured2 cards, no speculative decoding
TP2·MTP0— Full quality passed, but strict speed missed the frozen 49.019651 gate by 0.021%; not promoted.49.01 tok/sevidence - × Unsupported3 cards, no speculative decoding
TP3·MTP0head split - × Unsupported3 cards, + speculative decoding (depth 1)
TP3·MTP1head split - × Unsupported3 cards, + speculative decoding (depth 2)
TP3·MTP2head split - × Unsupported3 cards, + speculative decoding (depth 3)
TP3·MTP3head split - × Unsupported3 cards, + speculative decoding (depth 4)
TP3·MTP4head split - × Unsupported3 cards, + speculative decoding (depth 5)
TP3·MTP5head split
17 untested combinations
TP1·MTP0TP4·MTP0TP1·MTP1TP2·MTP1TP4·MTP1TP1·MTP2TP2·MTP2TP4·MTP2TP1·MTP3TP2·MTP3TP4·MTP3TP1·MTP4TP2·MTP4TP4·MTP4TP1·MTP5TP2·MTP5TP4·MTP5a3561ef8 TP4 overlay · graph on
Qwen3.8 AutoRound · vLLM XPU a3561ef8 plus the TP4 historical-winner decisions · strict natural-EOS profile · MTP0 · f16 KV · graph on. This is the accepted versioned overlay. Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-autoround-w4a16-bce40ca · variant=AutoRound INT4 · runtime=vLLM XPU nightly a3561ef8 + TP4 historical-winner overlay · profile_id=a3561ef8-tp4-e9d1398-best-config-overlay-v1 · measurement_class=strict-natural-eos · graph=on · kv=f16 · configured_max_context_tokens=32768 · natural_eos=True.
- ✓ Measured4 cards, no speculative decoding
TP4·MTP0— Accepted observed strict range; neither endpoint is independently replicated.71.45 tok/sevidence - × Unsupported3 cards, no speculative decoding
TP3·MTP0head split - × Unsupported3 cards, + speculative decoding (depth 1)
TP3·MTP1head split - × Unsupported3 cards, + speculative decoding (depth 2)
TP3·MTP2head split - × Unsupported3 cards, + speculative decoding (depth 3)
TP3·MTP3head split - × Unsupported3 cards, + speculative decoding (depth 4)
TP3·MTP4head split - × Unsupported3 cards, + speculative decoding (depth 5)
TP3·MTP5head split
17 untested combinations
TP1·MTP0TP2·MTP0TP1·MTP1TP2·MTP1TP4·MTP1TP1·MTP2TP2·MTP2TP4·MTP2TP1·MTP3TP2·MTP3TP4·MTP3TP1·MTP4TP2·MTP4TP4·MTP4TP1·MTP5TP2·MTP5TP4·MTP5Q4_K_M context × TP
Qwen3.8 Q4_K_M · llama.cpp SYCL · target only · f16 KV. Exact raw-engine decode/prefill sweep points are shown from the measured zero-depth anchor through 32K. Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-ggmlorg-q4-k-m-0669b98 · variant=Q4_K_M · runtime=llama.cpp SYCL · mtp=0 · graph=off · kv=f16.
- ✓ Measured1 card, 0
TP1·0D24.81 · P825.24guide · evidence - ✓ Measured1 card, 2K
TP1·2KD24.46 · P919.67guide · evidence - ✓ Measured1 card, 4K
TP1·4KD24.25 · P892.64guide · evidence - ✓ Measured1 card, 8K
TP1·8KD23.83 · P850.97guide · evidence - ✓ Measured1 card, 16K
TP1·16KD23.1 · P779.5guide · evidence - ✓ Measured1 card, 24K
TP1·24KD22.42 · P719.39guide · evidence - ✓ Measured1 card, 32K
TP1·32KD21.77 · P667.83guide · evidence
Codes in the rows: D = decode tok/s · P = prefill tok/s · T = ms to first token · AR = share of drafted tokens accepted · A = combined tok/s.
21 untested combinations
TP2·0TP3·0TP4·0TP2·2KTP3·2KTP4·2KTP2·4KTP3·4KTP4·4KTP2·8KTP3·8KTP4·8KTP2·16KTP3·16KTP4·16KTP2·24KTP3·24KTP4·24KTP2·32KTP3·32KTP4·32KQ4_K_M q8 KV context × TP
Qwen3.8 Q4_K_M · llama.cpp SYCL · target only · q8_0 KV. Exact raw-engine decode/prefill sweep points are shown from the measured zero-depth anchor through 32K. Fixed: revision=qwen3.8-27b · artifact_id=qwen38-27b-ggmlorg-q4-k-m-0669b98 · variant=Q4_K_M · runtime=llama.cpp SYCL · mtp=0 · graph=off · kv=q8_0.
- ✓ Measured1 card, 0
TP1·0D24.27 · P817.78guide · evidence - ✓ Measured1 card, 2K
TP1·2KD22.45 · P912.17guide · evidence - ✓ Measured1 card, 4K
TP1·4KD21.05 · P887.17guide · evidence - ✓ Measured1 card, 8K
TP1·8KD18.68 · P843.21guide · evidence - ✓ Measured1 card, 16K
TP1·16KD14.86 · P772.01guide · evidence - ✓ Measured1 card, 24K
TP1·24KD12.4 · P711.48guide · evidence - ✓ Measured1 card, 32K
TP1·32KD10.66 · P662.56guide · evidence
Codes in the rows: D = decode tok/s · P = prefill tok/s · T = ms to first token · AR = share of drafted tokens accepted · A = combined tok/s.
21 untested combinations
TP2·0TP3·0TP4·0TP2·2KTP3·2KTP4·2KTP2·4KTP3·4KTP4·4KTP2·8KTP3·8KTP4·8KTP2·16KTP3·16KTP4·16KTP2·24KTP3·24KTP4·24KTP2·32KTP3·32KTP4·32K256K package quant × TP
Qwen3.8 full-package TP1 serving profile with q8_0 target/draft KV, MTP3, and mmproj-F16 resident. 262144 is configured capacity; the measured serving prompts were 48–78 tokens, not 256K active context. Fixed: revision=qwen3.8-27b · runtime_family=llama.cpp SYCL · profile_id=q38-256k-vision-mtp3-q8kv-v1 · measurement_class=conventional-99-interval-cold-realistic-suite · mtp=3 · graph=off · kv=q8_0 · configured_max_context_tokens=262144 · vision=mmproj-F16 resident.
- ✓ Measured1 card, UD-Q5_K_S
TP1·UD-Q5_K_S26.67 tok/sguide · evidence - ✓ Measured1 card, UD-Q4_K_XL
TP1·UD-Q4_K_XL27.51 tok/sevidence
6 untested combinations
TP2·UD-Q5_K_STP3·UD-Q5_K_STP4·UD-Q5_K_STP2·UD-Q4_K_XLTP3·UD-Q4_K_XLTP4·UD-Q4_K_XLllama.cpp quant × TP
Qwen3.8 target-only llama.cpp/SYCL routes with f16 KV. Cells retain their recorded workload identity: some TP1 quant rows expose raw context curves, while Q4_K_M TP2 and Q8_0 TP1/TP2 expose conventional or strict varied-prompt suite headlines. Fixed: revision=qwen3.8-27b · runtime=llama.cpp SYCL · mtp=0 · graph=off · kv=f16.
- ✓ Measured1 card, Q4_K_M
TP1·Q4_K_Mmeasured context curveguide · evidence - ✓ Measured2 cards, Q4_K_M
TP2·Q4_K_M49.72 tok/sguide · evidence - ✓ Measured1 card, UD-Q5_K_S
TP1·UD-Q5_K_Smeasured context curveevidence - ✓ Measured1 card, UD-Q4_K_XL
TP1·UD-Q4_K_XLmeasured context curveevidence - ✓ Measured1 card, Q8_0
TP1·Q8_019.62 tok/sguide · evidence - ✓ Measured2 cards, Q8_0
TP2·Q8_036.73 tok/sguide · evidence
10 untested combinations
TP3·Q4_K_MTP4·Q4_K_MTP2·UD-Q5_K_STP3·UD-Q5_K_STP4·UD-Q5_K_STP2·UD-Q4_K_XLTP3·UD-Q4_K_XLTP4·UD-Q4_K_XLTP3·Q8_0TP4·Q8_0vLLM quant × TP
Qwen3.8 target-only vLLM XPU routes with f16 KV and graph enabled. AutoRound cells use the stock a3561ef8 strict profile; FP8 uses its separately pinned official-model runtime. Fixed: revision=qwen3.8-27b · runtime_family=vLLM XPU · mtp=0 · graph=on · kv=f16.
- ✓ Measured1 card, AutoRound INT4
TP1·AutoRound INT430.24 tok/sreport · evidence - ✓ Measured2 cards, AutoRound INT4
TP2·AutoRound INT448.49 tok/sreport · evidence - ✓ Measured4 cards, AutoRound INT4
TP4·AutoRound INT4— The 71.900 capture did not replicate and is not promoted.71.9 tok/sreport · evidence - ✓ Measured2 cards, FP8
TP2·FP821.71 tok/sguide · evidence - × Unsupported3 cards, AutoRound INT4
TP3·AutoRound INT4head split - × Unsupported3 cards, FP8
TP3·FP8head split
2 untested combinations
TP1·FP8TP4·FP8Measured results
Every number links to its proof.
What changes with card count
Qwen3.8 AutoRound · pinned vLLM XPU e9d1398 · MTP0 · f16 KV · fixed-output diagnostic; TP2/TP4 XPU Graph is runtime-labeled unsupported/experimental. No a3561ef8 eager, prefill, or TTFT values are inferred here.
pinned e9d1398 graph 1: 30.22–30.26 · 2: 48.83–48.95 · 4: 71.55–71.67
pinned e9d1398 graph 1: 276.3–277.9 · 2: 509 · 4: 849.5–864.8
pinned e9d1398 graph 1: 277.5–278.6 · 2: 150.31–152.2 · 4: 90.8–91.1
Draft depth (one pinned build)
Qwen3.8 AutoRound · TP1 eager · f16 KV · pinned vLLM XPU e9d1398; preregistered ladder ended at MTP3. No a3561ef8 MTP value is inferred.
Q4_K_M HTTP context × KV/graph
Qwen3.8 Q4_K_M · TP1 llama.cpp · graph-off F16 is Grade C cache-zero HTTP at 2K-32K and keeps x=0 as a separate raw point · exact cache64 graph-patched F16 covers 0-32K · Q8_0 remains the preserved raw-engine graph-off curve · profiles remain separately labeled
F16 KV · SYCL graph cache64 · HTTP 0-32K · Grade C 0: 26.72 · 2K: 26.27 · 4K: 26.02 · 8K: 25.54 · 16K: 24.58 · 24K: 23.74 · 32K: 23.22
Q8_0 KV · graph off · raw 0-32K 0: 24.27 · 2K: 22.45 · 4K: 21.05 · 8K: 18.68 · 16K: 14.86 · 24K: 12.4 · 32K: 10.66
Conversation length × compression
Qwen3.8 · TP1 llama.cpp · f16 KV · separately pinned builds and raw-engine shape sweeps; inspect each series identity before cross-quant comparison
UD-Q5_K_S 0: 22.72 · 2K: 22.37 · 4K: 22.13 · 8K: 21.79 · 16K: 21.19 · 24K: 20.61 · 32K: 20.09
UD-Q4_K_XL 0: 21.81 · 2K: 21.53 · 4K: 21.37 · 8K: 21.06 · 16K: 20.49 · 24K: 19.95 · 32K: 19.45
Q8_0 weights 0: 19.66 · 2K: 19.59 · 4K: 19.51 · 8K: 19.29 · 16K: 18.84 · 24K: 18.42 · 32K: 18.02
UD-Q5_K_S 0: 849.3 · 2K: 778.35 · 4K: 759.39 · 8K: 726.59 · 16K: 671.59 · 24K: 625.83 · 32K: 588.21
UD-Q4_K_XL 0: 762.57 · 2K: 735.45 · 4K: 719.47 · 8K: 689.58 · 16K: 639.46 · 24K: 596.49 · 32K: 564.86
Q8_0 weights 0: 996.89 · 2K: 987.39 · 4K: 959.56 · 8K: 914.06 · 16K: 837.51 · 24K: 773.15 · 32K: 719.14
33 more evidence views
Card count (fully checked runs)
Qwen3.8 AutoRound · MTP0 · f16 KV · graph on · strict natural-EOS. Pinned e9d1398, stock a3561ef8, and a3561ef8 winner overlays remain separate; stock TP4 71.900 is a captured high, not a replicated promotion; TP2 overlay is partial and TP4 overlay is accepted.
stock a3561ef8 strict 1: 30.24–30.24 · 2: 48.49 · 4: 71.25–71.9
a3561ef8 overlay · TP2 partial / TP4 accepted 2: 49.01 · 4: 71.35–71.45
Card count (quick checks)
Qwen3.8 AutoRound · MTP0 · f16 KV · graph on · fixed-output ignore-EOS diagnostic. These are measured diagnostic captures, not strict deployment headlines; the overlay was measured only at TP2/TP4.
stock a3561ef8 diagnostic 1: 30.33 · 2: 48.65 · 4: 71.34
a3561ef8 overlay diagnostic 2: 49.06 · 4: 71.72
Qwen3.6 draft-depth ladder
Webhie AutoRound · TP1 · graph capture 8 · same-window strict gate · historical generated-tokens-1–100 accounting; MTP0 was not rerun in this window
Qwen3.6 draft depth 5 (quick check)
Webhie AutoRound · TP1 · MTP5 · PIECEWISE graph · same-window strict/cache-zero screen · historical generated-tokens-1–100 accounting; neither capture size beat MTP3.
Qwen3.6 draft depth 3: then vs. now
AutoRound · TP1/TP2 · graph on · f16 KV. Historical highs and dated current support use disclosed different suites/accounting; whiskers are captured support bands, not a regression claim.
dated current support 1: 65.36–66.72 · 2: 94.32–94.79
Qwen3.6 Q4_0 with its built-in draft
TP1 llama.cpp SYCL · exact Q4_0 artifact · graph off · q8_0 KV · fixed 12-prompt cold/cache-zero suite · MTP3 is 48.434 conventional (48.923285 retained helper); MTP1/2/4/5 remain gaps
Qwen3.6 Q4_0 with a DFlash draft
TP1 native DFlash5 · Q4_0 target + Q8_0 draft · configuration-exact fusion-off/on A/B, not output-exact · confirmed candidate is +8.14% versus control at 47.341 conventional (47.818818 retained helper)
Qwen3.6 UD-Q4_K_XL draft depth
Four independent TP1 services in one matched wave · MTP0/3/4/5 · n_min=0 · p_min=0 · fixed 12-prompt cold/cache-zero suite · graph requested but replay unproven · conventional 99-interval accounting
Qwen3.6 UD-Q4_K_XL deep draft (closed)
TP1 policy screen · each point changes n_min and/or p_min as labeled · discrete closure, not a pure depth curve · best is 31.165 conventional (31.480049 retained helper) · fixed cold/cache-zero suite
MTP7 · n1 · p.75 7: 30.68
MTP7 · n1 · p.65 7: 31.17
MTP9 · n2 · p.75 9: 27.04
MTP7 · n1 · p.75 7: 427.97
MTP7 · n1 · p.65 7: 424.65
MTP9 · n2 · p.75 9: 427.32
Qwen3.8 by software build
Separate older TP2 runtime family · honest MTP5 three-arm median 101.17 (100.46–101.39) · only 21–22/25 self-agreement; MTP1–4 omitted, not interpolated
Qwen3.6 quant × KV context
TP1 llama.cpp SYCL · MTP0 · graph off · raw-engine exact-depth points only; quantizations and KV types remain separate series
Q4_K_M · q8_0 KV 0: 28.43 · 2K: 25.91 · 4K: 24 · 8K: 20.92 · 16K: 16.17 · 24K: 13.3 · 32K: 11.38
UD-Q4_K_XL · f16 KV 0: 28.2 · 2K: 27.73 · 4K: 27.4 · 8K: 26.9 · 16K: 25.97 · 24K: 25.11 · 32K: 24.31
UD-Q4_K_XL · q8_0 KV 0: 27.4 · 2K: 25.05 · 4K: 23.21 · 8K: 20.33 · 16K: 15.86 · 24K: 13.11 · 32K: 11.21
Q4_0 · f16 KV 0: 26.4 · 2K: 25.61 · 4K: 24.98 · 8K: 23.89 · 16K: 21.96 · 24K: 20.31 · 32K: 18.91
Q4_0 · q8_0 KV 0: 25.87 · 2K: 23.94 · 4K: 22.51 · 8K: 20.21 · 16K: 16.84 · 24K: 14.42 · 32K: 12.6
Q8_0 embedded-MTP artifact · f16 KV 0: 19.83 · 2K: 19.64 · 4K: 19.5 · 8K: 19.27 · 16K: 18.81 · 24K: 18.39 · 32K: 17.98
Q8_0 embedded-MTP artifact · q8_0 KV 0: 19.41 · 2K: 18.27 · 4K: 17.34 · 8K: 15.71 · 16K: 12.88 · 24K: 10.96 · 32K: 9.57
Q8_0 target-only Unsloth · f16 KV 0: 19.84 · 2K: 19.64 · 4K: 19.51 · 8K: 19.27 · 16K: 18.81 · 24K: 18.38 · 32K: 17.98
Q8_0 target-only Unsloth · q8_0 KV 0: 19.41 · 2K: 18.27 · 4K: 17.35 · 8K: 15.71 · 16K: 12.88 · 24K: 10.96 · 32K: 9.57
Q4_K_M · q8_0 KV 0: 824.62 · 2K: 884.56 · 4K: 865.76 · 8K: 823.1 · 16K: 756.11 · 24K: 698.56 · 32K: 648.26
UD-Q4_K_XL · f16 KV 0: 840.76 · 2K: 890.41 · 4K: 864.38 · 8K: 826.18 · 16K: 758.86 · 24K: 701.5 · 32K: 654.62
UD-Q4_K_XL · q8_0 KV 0: 837.72 · 2K: 883.05 · 4K: 859.58 · 8K: 822.39 · 16K: 751.83 · 24K: 696.25 · 32K: 645.98
Q4_0 · f16 KV 0: 635.64 · 2K: 300.21 · 4K: 284.81 · 8K: 255.79 · 16K: 203.91 · 24K: 168.95 · 32K: 144.91
Q4_0 · q8_0 KV 0: 634.03 · 2K: 299.3 · 4K: 284.23 · 8K: 256.07 · 16K: 203.93 · 24K: 168.64 · 32K: 144.96
Q8_0 embedded-MTP artifact · f16 KV 0: 908.1 · 2K: 897.5 · 4K: 871.16 · 8K: 830.35 · 16K: 763.17 · 24K: 707.36 · 32K: 658.01
Q8_0 embedded-MTP artifact · q8_0 KV 0: 901.98 · 2K: 890.55 · 4K: 863.3 · 8K: 823.43 · 16K: 753.25 · 24K: 701.09 · 32K: 653.03
Q8_0 target-only Unsloth · f16 KV 0: 907.79 · 2K: 897.08 · 4K: 872.57 · 8K: 830.75 · 16K: 762.67 · 24K: 708.32 · 32K: 660.85
Q8_0 target-only Unsloth · q8_0 KV 0: 901.91 · 2K: 891.93 · 4K: 862.37 · 8K: 823.05 · 16K: 757.05 · 24K: 702.84 · 32K: 653.73
Q4_K_XL HTTP context × KV/graph
Grade C · same Qwen3.8 UD-Q4_K_XL artifact · target-only TP1 · exact cache-zero HTTP serving through 32K · graph-off F16/Q8_0 and exact cache64 graph-patched F16 profiles remain separately labeled · all full quality batteries passed
F16 KV · SYCL graph cache64 · Grade C 0: 23.21 · 2K: 22.82 · 4K: 22.66 · 8K: 22.27 · 16K: 21.45 · 24K: 20.76 · 32K: 20.35
Q8_0 KV · graph off · Grade C 0: 21.77 · 2K: 20.27 · 4K: 19.09 · 8K: 17.12 · 16K: 13.83 · 24K: 11.65 · 32K: 10.08
Q8_0-weight HTTP context × KV/graph
Grade C · same Qwen3.8 Q8_0 artifact · target-only TP1 · graph-off F16/Q8_0 and cache64 graph-patched F16 remain separately labeled · the graph profile can be slower and does not replace graph-off · all full quality batteries passed
F16 KV · SYCL graph cache64 · HTTP · Grade C 0: 19.17 · 2K: 18.98 · 4K: 18.93 · 8K: 18.69 · 16K: 18.22 · 24K: 17.77 · 32K: 17.52
Q8_0 KV · graph off · HTTP · Grade C 0: 15.64 · 2K: 14.87 · 4K: 14.24 · 8K: 13.14 · 16K: 11.1 · 24K: 9.64 · 32K: 8.55
Q5_K_S HTTP context × KV/graph
Grade C · same Qwen3.8 UD-Q5_K_S artifact · target-only TP1 · exact cache-zero HTTP serving through 32K · graph-off F16/Q8_0 and exact cache64 graph-patched F16 profiles remain separately labeled · all full quality batteries passed
F16 KV · SYCL graph cache64 · Grade C 0: 23.99 · 2K: 23.64 · 4K: 23.43 · 8K: 23 · 16K: 22.22 · 24K: 21.51 · 32K: 21.02
Q8_0 KV · graph off · Grade C 0: 22.49 · 2K: 20.88 · 4K: 19.59 · 8K: 17.5 · 16K: 14.05 · 24K: 11.82 · 32K: 10.21
Qwen3.8 TP2 HTTP context × format
Grade C · exact cache-zero HTTP serving at 2K-32K for three separately pinned artifacts and runtimes; x=0 remains missing · Q4_K_M and Q8_0 use llama.cpp graph-off while official FP8 uses vLLM PIECEWISE · cross-format comparisons are descriptive, not causal, and quality scopes remain distinct
Q8_0 weights · F16 KV · llama.cpp graph off 2K: 36.8 · 4K: 36.55 · 8K: 36.01 · 16K: 35.09 · 24K: 34.52 · 32K: 33.85
Official FP8 · F16 KV · vLLM PIECEWISE 2K: 21.84 · 4K: 21.67 · 8K: 21.27 · 16K: 20.93 · 24K: 20.65 · 32K: 20.39
Q8_0 weights · F16 KV · llama.cpp graph off 2K: 2,003 · 4K: 3,957 · 8K: 8,047 · 16K: 16,681 · 24K: 25,958 · 32K: 35,832
Official FP8 · F16 KV · vLLM PIECEWISE 2K: 1,385 · 4K: 2,606 · 8K: 5,192 · 16K: 10,533 · 24K: 16,139 · 32K: 21,873
Official FP8 TP1/TP2/TP4 HTTP context
Grade C · same exact official Qwen3.8 FP8 weights · TP1 eager fit supports measured 2K-8K; target-only TP2/TP4 PIECEWISE and the separately qualified TP2 W8A16/MTP1 graph-off profile reach measured 2K-32K · x=0 and unmeasured TP1 depths remain missing · profiles are separately labeled and not a causal scaling A/B
Official FP8 · TP2 · F16 KV · vLLM PIECEWISE 2K: 21.84 · 4K: 21.67 · 8K: 21.27 · 16K: 20.93 · 24K: 20.65 · 32K: 20.39
Official FP8 + W8A16 · TP2 · MTP1 · deterministic graph off 2K: 44.78 · 4K: 54.93 · 8K: 51.31 · 16K: 51.29 · 24K: 43.72 · 32K: 46.64
Official FP8 · TP4 · F16 KV · vLLM PIECEWISE 2K: 35.53 · 4K: 34.9 · 8K: 33.97 · 16K: 33.64 · 24K: 33.27 · 32K: 33.11
Official FP8 · TP2 · F16 KV · vLLM PIECEWISE 2K: 1,385 · 4K: 2,606 · 8K: 5,192 · 16K: 10,533 · 24K: 16,139 · 32K: 21,873
Official FP8 + W8A16 · TP2 · MTP1 · deterministic graph off 2K: 799.1 · 4K: 1,187 · 8K: 2,343 · 16K: 4,869 · 24K: 7,586 · 32K: 10,487
Official FP8 · TP4 · F16 KV · vLLM PIECEWISE 2K: 1,973 · 4K: 3,663 · 8K: 7,200 · 16K: 14,424 · 24K: 21,792 · 32K: 29,236
Qwen3.8 AutoRound TP1 current-image KV/graph
Grade C · exact cache-zero HTTP serving · same current AutoRound weights and official f01e/ac7509e2 image · TP1 · MTP0 · F16 PIECEWISE plots target-identical 2K/4K/16K/24K/32K with 8K quarantined · F16 and FP8 E4M3 KV remain separately labeled · x=0 remains missing
FP8 E4M3 KV · eager 2K: 12.11 · 4K: 11.99 · 8K: 12.09 · 16K: 12.18 · 24K: 12.16 · 32K: 12.16
F16 KV · PIECEWISE parity-qualified partial 2K: 30.08 · 4K: 29.41 · 16K: 28.19 · 24K: 27.46 · 32K: 26.76
FP8 E4M3 KV · PIECEWISE 2K: 29.76 · 4K: 28.94 · 8K: 28.66 · 16K: 28.03 · 24K: 27.32 · 32K: 26.78
FP8 E4M3 KV · eager 2K: 1,532 · 4K: 2,805 · 8K: 5,817 · 16K: 12,452 · 24K: 19,854 · 32K: 27,973
F16 KV · PIECEWISE parity-qualified partial 2K: 1,363 · 4K: 2,475 · 16K: 10,545 · 24K: 16,544 · 32K: 22,672
FP8 E4M3 KV · PIECEWISE 2K: 1,383 · 4K: 2,535 · 8K: 5,269 · 16K: 11,340 · 24K: 18,166 · 32K: 25,756
Qwen3.8 AutoRound TP1 graph modes
Grade C · exact cache-zero HTTP serving · same AutoRound weights · MTP0 · F16 KV · current f01e eager spans 2K-32K, current f01e PIECEWISE plots target-identical 2K/4K/16K/24K/32K with 8K quarantined, and FULL_AND_PIECEWISE remains the separately labeled b2dd/1e90 profile · x=0 remains missing
PIECEWISE parity-qualified partial · f01e/ac7509e2 2K: 30.08 · 4K: 29.41 · 16K: 28.19 · 24K: 27.46 · 32K: 26.76
FULL_AND_PIECEWISE · b2dd/1e90 2K: 30.1 · 4K: 29.77 · 8K: 29.28 · 16K: 28.42 · 24K: 27.66 · 32K: 26.99
PIECEWISE parity-qualified partial · f01e/ac7509e2 2K: 1,363 · 4K: 2,475 · 16K: 10,545 · 24K: 16,544 · 32K: 22,672
Qwen3.8 AutoRound TP2 graph modes
Grade C · exact cache-zero HTTP serving · F16 KV · target-only MTP0 spans eager, current PIECEWISE partial, and dated FULL_AND_PIECEWISE profiles; native MTP1 and MTP2 each add one separately qualified current-PIECEWISE exact-4K point · all unplotted selectors remain missing
PIECEWISE partial · f01e/ac7509e2 2K: 39.68 · 4K: 46.64 · 24K: 42.17 · 32K: 41.14
MTP1 PIECEWISE · exact 4K Grade C 4K: 13.74
MTP2 PIECEWISE · exact 4K Grade C 4K: 18.41
FULL_AND_PIECEWISE · b2dd/1e90 2K: 48.15 · 4K: 47.93 · 8K: 46.83 · 16K: 45.2 · 24K: 43.7 · 32K: 42.34
PIECEWISE partial · f01e/ac7509e2 2K: 1,703 · 4K: 1,563 · 24K: 10,230 · 32K: 14,062
MTP1 PIECEWISE · exact 4K Grade C 4K: 2,651
MTP2 PIECEWISE · exact 4K Grade C 4K: 3,126
FULL_AND_PIECEWISE · b2dd/1e90 2K: 782.76 · 4K: 1,569 · 8K: 3,187 · 16K: 6,611 · 24K: 10,265 · 32K: 14,105
Qwen3.8 AutoRound TP4 graph modes
Grade C · exact cache-zero HTTP serving · target-only MTP0 · F16 KV · current f01e eager spans 2K-32K, current f01e PIECEWISE plots target-identical 2K/4K/16K/24K/32K with 8K quarantined, and dated b2dd/1e90 FULL_AND_PIECEWISE remains a separate fully certified six-depth profile · x0 missing
eager 8K oracle · f01e/ac7509e2 8K: 9.65
PIECEWISE partial · f01e/ac7509e2 2K: 51.07 · 4K: 64.42 · 16K: 62.78 · 24K: 62.09 · 32K: 60.51
FULL_AND_PIECEWISE · b2dd/1e90 2K: 71.17 · 4K: 70.15 · 8K: 69.87 · 16K: 69.07 · 24K: 67.87 · 32K: 66.65
eager 8K oracle · f01e/ac7509e2 8K: 4,768
PIECEWISE partial · f01e/ac7509e2 2K: 1,633 · 4K: 1,106 · 16K: 4,568 · 24K: 7,051 · 32K: 9,606
FULL_AND_PIECEWISE · b2dd/1e90 2K: 547.47 · 4K: 1,105 · 8K: 2,210 · 16K: 4,562 · 24K: 7,012 · 32K: 9,569
Qwen3.8 AutoRound TP1 eager MTP4 screened depth ◇ screened, experimental
Grade D lab-screened only · current f01e/ac7509e2 image · F16 KV · exact cache-zero target parity at 4K/16K/24K · 2K and cross-boot-conflict 8K quarantined · 32K runtime-closed · x=0 missing · zero lab-measured cells
Qwen3.8 AutoRound TP1 eager MTP3 partial depth
Grade D partial · current f01e/ac7509e2 image · F16 KV · exact target parity at 4K-32K · 2K quarantined at token 90 · x=0 missing · conventional 99-interval decode
Qwen3.8 AutoRound TP1 MTP2 depth
Current f01e/ac7509e2 image · F16 KV · eager Grade D partial plus a separate PIECEWISE exact-4K Grade C point with its original cache-gate omission disclosed and a passing post-run audit · 2K/8K/16K corruption caveats remain · conventional 99-interval decode
MTP2 PIECEWISE · exact 4K Grade C · cache defect disclosed 4K: 11.99
MTP2 PIECEWISE · exact 4K Grade C · cache defect disclosed 4K: 2,942
MTP2 PIECEWISE · exact 4K Grade C · cache defect disclosed 4K: 0.85
Qwen3.8 AutoRound TP1 MTP1 depth
Current f01e/ac7509e2 image · eager/F16 Grade D partial plus separate Grade C exact-4K PIECEWISE/F16 and eager/E4M3 singletons · every unplotted depth remains missing; PIECEWISE 8K retains the token-99 corruption caveat · conventional 99-interval decode
MTP1 PIECEWISE/F16 · exact 4K Grade C 4K: 8.69
MTP1 eager/E4M3 · exact 4K Grade C 4K: 8.38
MTP1 PIECEWISE/F16 · exact 4K Grade C 4K: 2,962
MTP1 eager/E4M3 · exact 4K Grade C 4K: 3,843
MTP1 PIECEWISE/F16 · exact 4K Grade C 4K: 0.79
MTP1 eager/E4M3 · exact 4K Grade C 4K: 0.94
Qwen3.8 AutoRound TP1/TP2/TP4 HTTP context
Grade C · exact cache-zero HTTP serving · current f01e eager TP2 MTP0/MTP1 and TP4 MTP0 span every nonzero depth; TP2 MTP2 and TP4 MTP1-3 have target-parity 4K-32K with 2K quarantined; TP4 prior 8K points are retained; TP4 MTP4 recovery measures 4K/16K/24K while 2K/8K remain quarantined and 32K is runtime-closed · F16 KV · x=0 missing
AutoRound INT4 · TP2 · F16 KV 2K: 48.15 · 4K: 47.93 · 8K: 46.83 · 16K: 45.2 · 24K: 43.7 · 32K: 42.34
AutoRound INT4 · TP2 · MTP0 · F16 KV · current f01e eager 2K: 9.65 · 4K: 10.04 · 8K: 10.11 · 16K: 10.12 · 24K: 10.15 · 32K: 10.2
AutoRound INT4 · TP2 · MTP1 · F16 KV · current f01e eager 2K: 11.88 · 4K: 14.95 · 8K: 15.22 · 16K: 13.78 · 24K: 13.74 · 32K: 13.92
AutoRound INT4 · TP2 · MTP2 · F16 KV · current f01e eager 4K: 20.36 · 8K: 20.89 · 16K: 18.29 · 24K: 17.76 · 32K: 17.7
AutoRound INT4 · TP2 · MTP3 · F16 KV · current f01e eager 4K: 19.08 · 8K: 25.12 · 16K: 20.28 · 24K: 20.17 · 32K: 20.16
AutoRound INT4 · TP4 · F16 KV · graph 2K: 71.17 · 4K: 70.15 · 8K: 69.87 · 16K: 69.07 · 24K: 67.87 · 32K: 66.65
AutoRound INT4 · TP4 · MTP0 · F16 KV · current f01e eager prior 8K 8K: 9.65
AutoRound INT4 · TP4 · MTP0 · F16 KV · current f01e eager expansion 2K: 9.83 · 4K: 10.05 · 16K: 10.26 · 24K: 10.17 · 32K: 10.21
AutoRound INT4 · TP4 · MTP0 · F16 KV · current f01e PIECEWISE partial 2K: 51.07 · 4K: 64.42 · 16K: 62.78 · 24K: 62.09 · 32K: 60.51
AutoRound INT4 · TP4 · MTP1 · F16 KV · current f01e eager 8K: 13.71
AutoRound INT4 · TP4 · MTP1 · F16 KV · current f01e eager expansion 4K: 15.96 · 16K: 15.12 · 24K: 14.88 · 32K: 14.77
AutoRound INT4 · TP4 · MTP2 · F16 KV · current f01e eager 8K: 18.08
AutoRound INT4 · TP4 · MTP2 · F16 KV · current f01e eager expansion 4K: 21.84 · 16K: 19.74 · 24K: 19.44 · 32K: 19.4
AutoRound INT4 · TP4 · MTP3 · F16 KV · current f01e eager 8K: 21.08
AutoRound INT4 · TP4 · MTP3 · F16 KV · current f01e eager expansion 4K: 25.32 · 16K: 21.94 · 24K: 22.09 · 32K: 22.55
AutoRound INT4 · TP4 · MTP4 · F16 KV · current f01e eager quality recovery 4K: 21.97 · 16K: 23.79 · 24K: 25.75
Q5_K_S external-MTP 8K route screen ◇ screened, experimental
Grade C diagnostic · TP1 · graph off · Q8_0 target/draft KV · MTP1–4 all diverge from MTP0 at token index 6; rates and draft counters have no certified or headline authority
Qwen3.6 Q8 target-only MTP0 SYCL graph
TP1 · raw-engine pp2048/tg128 · decode fully graph-certified; prefill mixed partial above depth 0 · both target-only and embedded-artifact graph lanes are slower than their matched graph-off controls and are evidence, not optimization wins
Q8_0 KV · SYCL graph 0: 19.01 · 2K: 17.93 · 4K: 17.03 · 8K: 15.47 · 16K: 12.71 · 24K: 10.84 · 32K: 9.48
Embedded-MTP artifact · F16 KV · SYCL graph (MTP0) 0: 19.36 · 2K: 19.18 · 4K: 19.08 · 8K: 18.85 · 16K: 18.41 · 24K: 18.01 · 32K: 17.62
Embedded-MTP artifact · Q8_0 KV · SYCL graph (MTP0) 0: 19.02 · 2K: 17.94 · 4K: 17.04 · 8K: 15.47 · 16K: 12.71 · 24K: 10.83 · 32K: 9.48
Q8_0 KV · SYCL graph 0: 896.09 · 2K: 871.93 · 4K: 850.23 · 8K: 809.77 · 16K: 747.45 · 24K: 691.4 · 32K: 645.09
Embedded-MTP artifact · F16 KV · SYCL graph (MTP0) 0: 900.48 · 2K: 880.57 · 4K: 857.1 · 8K: 819.59 · 16K: 755.79 · 24K: 700.16 · 32K: 650.65
Embedded-MTP artifact · Q8_0 KV · SYCL graph (MTP0) 0: 894.19 · 2K: 871.09 · 4K: 849.46 · 8K: 810.19 · 16K: 747.51 · 24K: 692.95 · 32K: 644.91
Qwen3.6 Q4_K_M KV × graph comparison
TP1 · target-only MTP0 · raw-engine pp2048/tg128 · F16 and Q8_0 graph decode are fully certified while prefill is mixed partial above depth 0 · graph-on is slower in both phases at all seven depths for both KV types, so the protected graph-off lanes remain recommended
F16 KV · SYCL graph · quality-qualified negative 0: 28.58 · 2K: 28.05 · 4K: 27.66 · 8K: 27.18 · 16K: 26.23 · 24K: 25.37 · 32K: 24.57
Q8_0 KV · graph off · protected control 0: 28.43 · 2K: 25.91 · 4K: 24 · 8K: 20.92 · 16K: 16.17 · 24K: 13.3 · 32K: 11.38
Q8_0 KV · SYCL graph · quality-qualified negative 0: 27.77 · 2K: 25.39 · 4K: 23.53 · 8K: 20.56 · 16K: 16.01 · 24K: 13.21 · 32K: 11.3
F16 KV · SYCL graph · quality-qualified negative 0: 818.2 · 2K: 800.16 · 4K: 780.49 · 8K: 746.47 · 16K: 688.17 · 24K: 644.83 · 32K: 606.64
Q8_0 KV · graph off · protected control 0: 824.62 · 2K: 884.56 · 4K: 865.76 · 8K: 823.1 · 16K: 756.11 · 24K: 698.56 · 32K: 648.26
Q8_0 KV · SYCL graph · quality-qualified negative 0: 815.88 · 2K: 795.43 · 4K: 775.66 · 8K: 738.53 · 16K: 684.94 · 24K: 639.66 · 32K: 600.52
Qwen3.6 embedded-Q8 MTP serving context
TP1 · F16 KV · graph off · HTTP serving · sealed MTP0/MTP3 R3 plus quality-R1 MTP1/2/4 · exact target-output parity · x=0 means zero prior context plus one prompt token
MTP1 · exact-output speculative 0: 28.48 · 2K: 27.55 · 4K: 27.73 · 8K: 27.57 · 16K: 24.9 · 24K: 24.74 · 32K: 23.73
MTP2 · exact-output speculative 0: 37.24 · 2K: 32.32 · 4K: 36.23 · 8K: 36.32 · 16K: 29.04 · 24K: 30.58 · 32K: 29.2
MTP3 · exact-output speculative 0: 39.51 · 2K: 34.66 · 4K: 39.73 · 8K: 41.39 · 16K: 31.84 · 24K: 34.25 · 32K: 31.81
MTP4 · exact-output speculative 0: 37.56 · 2K: 34.69 · 4K: 41.65 · 8K: 43.72 · 16K: 31.41 · 24K: 36.49 · 32K: 31.28
Qwen3.6 embedded-Q8 MTP Q8_0-KV screen ◇ screened, experimental
Grade C exploratory · TP1 · graph off · Q8_0 KV · deterministic 2K route divergence
MTP2 · Grade C screened 0: 31.98 · 2K: 31.56 · 4K: 35.08 · 8K: 34.49 · 16K: 26.72 · 24K: 30.42 · 32K: 25.46
MTP3 · Grade C screened 0: 33.04 · 2K: 32.9 · 4K: 38.34 · 8K: 39.29 · 16K: 27.53 · 24K: 35.84 · 32K: 27.32
MTP4 · Grade C screened 0: 32.53 · 2K: 31.8 · 4K: 40.28 · 8K: 39.87 · 16K: 27.68 · 24K: 39.25 · 32K: 26.7
MTP2 · Grade C screened 0: 0.88 · 2K: 0.66 · 4K: 0.85 · 8K: 0.87 · 16K: 0.7 · 24K: 0.94 · 32K: 0.81
MTP3 · Grade C screened 0: 0.73 · 2K: 0.54 · 4K: 0.77 · 8K: 0.84 · 16K: 0.59 · 24K: 0.93 · 32K: 0.73
MTP4 · Grade C screened 0: 0.59 · 2K: 0.45 · 4K: 0.69 · 8K: 0.75 · 16K: 0.52 · 24K: 0.92 · 32K: 0.63
Qwen3.6 Q8 (no draft) as the conversation grows
TP1 llama.cpp SYCL · discrete evidence, not a fitted curve · baseline x is one exact prompt count; each VDR2 x is the larger anchor of a disclosed two-prompt median with band-specific ubatch
VDR2 exact full-512 bands 4.4K: 16.59 · 17.3K: 15.14 · 31.8K: 13.69
Qwen3.6 Q8 with draft depth 3
TP1 llama.cpp SYCL · same embedded publisher-MTP artifact · fixed cold 12-prompt realistic suite · full token/content exact against the fresh control
Qwen3.6 Q8 at long context
TP1 llama.cpp VDR2 · exact control/MTP3 full outputs · two-wave functional screen, not an isolated headline
Fine print
Qwen3_5ForConditionalGeneration, 64 layers, 5120 hidden. architecture-matched; individual runtime paths require revision-specific exercise. Measurements, artifact hashes, outputs, quality decisions, and speed stay pinned to their exact recorded identity.
9 combinations the lab has stopped pursuing: The fixed cache-zero MTP5/7/9 policy screen peaked at 31.165 conventional tok/s and remained noncompetitive with the separately qualified vLLM AutoRound route. Reopen only for a material backend, kernel, or quantization change—not another generic flag sweep. evidence The intensive native-DFlash program ended at a confirmed 47.341 conventional tok/s on one B70, below its 100 tok/s TP1 objective. The patch and safety evidence remain preserved; reopen only for a new mechanism, not generic configuration churn. evidence The measured c1 MTP3 footprint was 29,911 MiB and projected c2 was 32,683 MiB, leaving no useful safe headroom. The c2 launch was not performed and CPU offload was not used. evidence The model has 16 GDN key heads, which cannot be evenly partitioned across TP3. evidence The XPU FlashAttention backend rejects fp8_e5m2 KV cache at engine initialization. evidence The current f01e/ac7509e2 engine explicitly rejects fp8_e5m2 KV at initialization; startup, canary, and quality did not run, so this closure does not transfer to graph-on, other TP/MTP, KV formats, or runtimes. evidence The least-compounded TP1 target-only anchor matched only 3/20 stable-set oracle outputs. evidence The exact 8K graph/Q8_0-KV request passed parity, but the required long-context quality case aborted at qptr->wait during command-graph recording. No speed is published; sibling depths remain missing and require their own evidence. evidence The checksum-pinned target-only Q8_0 file has mtp_included=false. Same-artifact intrinsic MTP1–4 therefore requires a different MTP-bearing artifact/package and is excluded from the target-only contract; this does not close separately bound external drafts. evidence Exact selectors are in the family data.
Many people at once Lab-measured
16 simultaneous users share 68 combined tok/s on 1 card (Q4_K_M target with external Q4_0 MTP2 · 16-slot/8K-total HTTP); the full curve is under Measured results and in the multi-user report.
How much faster could this get? Projected — not measured
The ML Bottleneck physics engine projects a tuned-run target and the physical ceiling for the headline setup (qwen3.8-27b · AutoRound INT4 · vLLM XPU nightly + TP4 historical-winner overlay · TP4 · MTP0 · graph on · f16 KV · max context 32,768 · GPU memory 0.6 · natural EOS). The grade is optimization headroom against the tuned-run target, not model quality.
Loading projections from mlbottleneck.com…