# Qwen3.8 27B AutoRound INT4 on 2x Intel Arc Pro B70 — lane setup

> **Certification: `research-status`.** Active or unresolved work, not a
> promoted reproduction; see its entry in
> [`repro/guide-catalog.json`](../guide-catalog.json).

<!-- replication-matrix:begin (generated by tools/build-qwen38-int4-replication-matrix.py; do not edit by hand) -->

**Replication matrix (class-balanced median decode tok/s at one request, strict 12-prompt six-class suite, 512-token completion cap; identity ladder = c1-c64 exactness vs a sequential oracle, 128 tokens per request):**

| image | TP | MTP depth | XPU graph | draft head | strict pair (tok/s) | gates | identity ladder | run |
|---|---|---|---|---|---|---|---|---|
| R228 | 2 | 0 | off | FP16 | 34.21 / 35.64 | G1 12/12 | c1 1/1, c2 2/2, c4 4/4, c8 8/8, c16 16/16, c32 32/32, c64 64/64 | R239 |
| R228 | 2 | 1 | off | FP16 | 51.10 / 50.09 | G2 12/12, G3 12/12 x2, probe exact | c1 1/1, c2 2/2, c4 4/4, c8 8/8, c16 16/16, c32 30/32, c64 58/64 | R239 |
| R228 | 2 | 2 | off | FP16 | 61.14 / 61.54 | G2 12/12, G3 12/12, G3 12/12 | - | R239 |
| R228 | 2 | 3 | off | FP16 | 67.61 / 67.83 | G2 12/12, G3 12/12, G3 12/12 | - | R239 |
| R228 | 2 | 4 | off | FP16 | 68.55 / 67.79 | G2 12/12, G3 12/12, G3 12/12 | - | R239 |
| R228 | 2 | 0 | on | FP16 | 49.83 / 49.89 | G1 12/12; 12/12 vs eager oracle | c1 1/1, c2 2/2, c4 4/4, c8 8/8, c16 16/16, c32 32/32, c64 64/64 | R253 |
| R228 | 2 | 1 | on | FP16 | 76.72 / 76.63 | G2 12/12, G3 12/12 x2, probe exact | c1 1/1, c2 2/2, c4 4/4, c8 8/8, c16 12/16, c32 30/32, c64 61/64 | R253 |
| R228 | 2 | 4 | on | FP16 | 91.00 / 91.01 | G2 12/12, G3 12/12 x2 | c1 1/1, c2 2/2, c4 3/4, c8 7/8, c16 16/16, c32 30/32, c64 58/64 | R247/R251 |
| R228 | 2 | 5 | on | FP16 | 88.84 / 89.12 | G2 12/12, G3 12/12 x2 | - | R250 |
| R228 | 2 | 6 | on | FP16 | 84.06 / 83.98 | G2 12/12, G3 12/12 x2 | - | R250 |
| R228 (R257 pair; R256 identical code path) | 2 | 4 | on | INT4 draft-only | 112.36 / 112.33 | G2 12/12, G3 12/12 x2 | c1 1/1, c2 2/2, c4 4/4, c8 8/8, c16 16/16, c32 31/32, c64 58/64 (warm pass; c32/c64 with the W4A16 pad forwarded off: R281 632.6 30/32, 603.7 62/64) | R257/R265b/R281 |
| R228 | 1 | 0 | off | FP16 | 32.96 / 32.95 | G1 12/12 | c1 1/1, c2 2/2, c4 4/4, c8 8/8, c16 16/16, c32 32/32, c64 64/64 | R239 |
| R228 | 1 | 1 | off | FP16 | 49.64 / 49.47 | G2 12/12, G3 12/12 x2, probe exact | c1 1/1, c2 2/2, c4 4/4, c8 8/8, c16 15/16, c32 30/32, c64 60/64 | R239 |
| R228 | 1 | 2 | off | FP16 | 56.51 / 56.45 | G2 12/12, G3 12/12, G3 12/12 | - | R239 |
| R228 | 1 | 3 | off | FP16 | 58.47 / 58.47 | G2 12/12, G3 12/12, G3 12/12 | - | R239 |
| R228 | 1 | 4 | off | FP16 | 56.29 / 56.25 | G2 12/12, G3 12/12, G3 12/12 | - | R239 |
| R228 | 1 | 4 | on | FP16 | 56.91 / 56.90 | 12/12 all | - | R246b |
| R256 | 2 | 5 | on | INT4 draft-only | 109.97 / 110.07 | G2 12/12, G3 12/12 x2 | - | R258 |
| R256 | 2 | 6 | on | INT4 draft-only | 108.34 / 108.37 | G2 12/12, G3 12/12 x2 | - | R258 |
| R256 | 2 | 0 | on | n/a | - | MTP0 ladder | c1 1/1, c2 2/2, c4 4/4, c8 8/8, c16 16/16, c32 31/32, c64 64/64 | R259 |
| R266 | 2 | 4 | on | INT4 draft-only (V2 runner, R266) | 112.70 / 112.96 | G2 12/12, G3 12/12 x2 | c1 1/1, c2 2/2, c4 3/4, c8 8/8, c16 16/16, c32 31/32, c64 62/64 (warm pass) | R269/R270 |
| R276 (served since 2026-09-06 pm) | 2 | 4 | on (sizes to 320) | INT4 draft-only | 112.90 / 113.00 | G2 12/12, G3 12/12 x2 | c1 1/1, c2 2/2, c4 4/4, c8 7/8, c16 16/16, c32 30/32, c64 60/64 (warm pass) | R283/R282 |
| R293 (classpad on) | 2 | 4 | on (sizes to 320) | INT4 draft-only | 111.69 / 111.33 | G2 12/12, G3 12/12 x2 (own MTP0 pair) | c1 1/1, c2 2/2, c4 4/4, c8 8/8, c16 16/16, c32 30/32, c64 59/64 (warm pass) | R295 |
| R293 (classpad on) | 2 | 0 | on (sizes to 320) | n/a | 49.39 / 49.39 | G1 12/12 | c1 1/1, c2 2/2, c4 4/4, c8 8/8, c16 16/16, c32 32/32, c64 64/64 (warm pass) | R295 |
| R293 (classpad on) | 1 | 4 | on (sizes to 320) | INT4 draft-only | 73.19 / 73.25 | G2 12/12, G3 12/12 x2 (own MTP0 pair) | c1 1/1, c2 2/2, c4 4/4, c8 8/8, c16 15/16, c32 29/32, c64 59/64 (warm pass) | R298 |
| R293 (classpad on) | 1 | 0 | on (sizes to 320) | n/a | 32.58 / 32.50 | G1 12/12 | c1 1/1, c2 2/2, c4 4/4, c8 8/8, c16 16/16, c32 32/32, c64 63/64 (warm pass) | R298 |

**Settings common to every row:** vLLM 0.27.2rc1.dev77+gac7509e2b (XPU), `--dtype float16 --quantization gptq --kv-cache-dtype auto --block-size 64 --no-enable-prefix-caching --language-model-only`, `VLLM_BATCH_INVARIANT=0` (vLLM's own switch is off: the strict launchers pin it and vLLM refuses to boot the GDN backend with it on; batch invariance on this lane comes from the kernels and the switches below), `TORCHINDUCTOR_DETERMINISTIC=1`, `VLLM_ENABLE_INDUCTOR_MAX_AUTOTUNE=0`, `VLLM_ENABLE_INDUCTOR_COORDINATE_DESCENT_TUNING=0`, `PYTHONHASHSEED=0`, `VLLM_XPU_GDN_SPEC_PERSISTENT_SCRATCH=1`, `VLLM_XPU_QWEN_GEMMA_RMSNORM_PACKED_SERIAL_EXACT=1`, `VLLM_XPU_GDN_NATIVE_FALLBACK=1`, `VLLM_XPU_FP8_BLOCK_W8A16=1` (inert on the gptq path), `VLLM_XPU_GDN_SPLIT_MIXED=1`, `VLLM_XPU_GDN_SPEC_GROUP=16`, `VLLM_XPU_FP16_LINEAR_ROWCHUNK=32` (R293 rows: `VLLM_XPU_FP16_LINEAR_CLASSPAD=1`, which replaces the 32-row pieces with one verified oneDNN M-class per weight shape), `VLLM_XPU_W4A16_DETERMINISM_PAD=0` (correction 2026-09-06: the launchers did not forward this switch until R278k, so every runner-launched ladder rung above 128 verify rows, i.e. c32/c64, ran with the R213b pad on; c1-c16 and the single-user headline were never affected; the R281 ladder below is the corrected measurement), `VLLM_XPU_ALLREDUCE_HOST_WAIT=1`, `VLLM_XPU_RMSNORM_TRITON=0`, `VLLM_XPU_GEMMA_RMSNORM_TRITON=0`, whole-graph `torch.compile` (`splitting_ops: []`) with `inductor_compile_config {deterministic: true, split_reductions: false, triton.autotune_pointwise: false, combo_kernels: false, benchmark_combo_kernel: false, benchmark_epilogue_fusion: false}`, oneCCL `CCL_ATL_TRANSPORT=ofi FI_PROVIDER=tcp CCL_ZE_IPC_EXCHANGE=pidfd CCL_SEND=direct CCL_RECV=direct CCL_TOPO_P2P_ACCESS=1` with the three `CCL_SYCL_*_SIMPLE_THRESHOLD=4294967296`, greedy decoding (`temperature 0`), speculative config `{"method":"qwen3_next_mtp","num_speculative_tokens":<depth>}` (omitted for MTP0). Model: `devan-carlin/Qwen3.8-27B-int4-AutoRound` bce40cac relabelled to plain gptq (manifest `model-gptq-relabel-r212.json`).

| setting | strict pairs (1 user) | identity ladders (c1-c64) |
|---|---|---|
| `--max-model-len` / context | 1024 | 256 |
| `--max-num-seqs` | 1 | 64 |
| `--max-num-batched-tokens` | 1024 | 512 |
| completion cap | 512 tokens (suite prompts 26-31 tokens, six classes) | 128 tokens, 64-prompt small-context suite |
| GPU memory utilization | TP2 0.95, TP1 0.96 | same |
| XPU graph on | `VLLM_XPU_ENABLE_XPU_GRAPH=1`, `cudagraph_mode FULL_DECODE_ONLY`, `cudagraph_capture_sizes [1,2,3,4,5,6,8]`, `max_cudagraph_capture_size 8` | same |
| XPU graph off | `VLLM_XPU_ENABLE_XPU_GRAPH=0`, `cudagraph_mode PIECEWISE`, `cudagraph_capture_sizes [1]` | same |
| draft head INT4 | `VLLM_XPU_DRAFT_LM_HEAD_INT4=1` (group 128, bf16 scales; R256 image or later) | same |
| TP1 | `TENSOR_PARALLEL_SIZE=1 XPU_DEVICE_MASK=0` (`ZE_AFFINITY_MASK=0`) | same |
| TP2 | `TENSOR_PARALLEL_SIZE=2 XPU_DEVICE_MASK=0,1` | same |

Images: R228 = `ghcr.io/steveseguin/vllm-openai-xpu-qwen38-int4@sha256:aaf920b04224cb3f4be881ae41dbef4fa7841f4ab26fbbe09e4e780fe361ff7d` (`_xpu_ops.py` sha256 c91d6b0d…, `_xpu_C.abi3.so` 271db0d4…); R256 = `ghcr.io/steveseguin/vllm-openai-xpu-qwen38-int4@sha256:f7696bcaefab1bc1c93e12cbde630b6e81bed8e00e41154ca2198e246c35dea3` (R228 + a draft-only INT4 head fallback whose branch is never taken on this model: the relabelled head is unquantized and already carries `make_xpu_int4_draft_copy`, so both images run the same code here; the container records show the R257 headline pair ran R228 and R258/R259/R260b/R265b ran R256); R276 = `ghcr.io/steveseguin/vllm-openai-xpu-qwen38-int4@sha256:521eb277c0733f8c2ce47aea1bb98ed576c6f1ad63bf5baf22d38fc07abf54ad` (`_xpu_ops.py` sha256 6ee6b8db…; R256 + the sync-free grouped GDN branch, served image since 2026-09-06 afternoon: graph capture to 320 tokens, lossless, +22%/+10% at c2/c4); R266 = `neural-download/vllm-openai-xpu:qwen38-int4-v2-draft-int4-head-r266 sha256:1d12b64e46f99a6092014319b2b66f14f380f24da82c6ce7b852db7ee6ebd10e (local; R256 + docker/r266-v2-draft-int4-head.py)` (V2 model runner variant, not the published path). Registry digest equals the local image id; pass it as `EXPECTED_IMAGE_ID` and the two file digests as `XPU_OPS_SHA256_OVERRIDE` / `XPU_EXTENSION_SHA256_OVERRIDE` (the launcher does this).

<!-- replication-matrix:end -->

## Fixed-K batch-invariant profile on the R187 stack (2026-09-05, research-status)

The 2026-09-04/05 refresh runs the same AutoRound tensors on the FP8 lane's R187 stack (whole-graph compile,
deterministic Inductor, no autotune, r152 harness) with two kernel-level changes and three runtime switches. Every
number below is class-balanced median decode at concurrency 1 unless a ladder is named; the two-run identity rule and
the MTP-vs-MTP0 oracle gate are the FP8 lane's.

**What changed and why**

1. *Kernel routing.* vLLM routes `quant_method: auto-round` to INC/ARK `woqgemm`, which is nondeterministic for 32-256
   rows, never batch-shape invariant (so speculative decoding can never be lossless on it) and 6-10x slower at two rows
   than one. The identical tensors relabelled as plain `gptq` select `XPUwNa16LinearKernel -> _xpu_C.int4_gemm_w4a16`
   (oneDNN). Builder: [`scripts/make-gptq-relabel.py`](scripts/make-gptq-relabel.py) (hard links plus rewritten
   `config.json`/`quantization_config.json`; the `mtp.fc` exclusion is spelled `mt[p]\.fc` so vLLM keeps the INT4 draft
   layers quantized); manifest [`manifests/model-gptq-relabel-r212.json`](manifests/model-gptq-relabel-r212.json).
2. *Fixed-K W4A16 GEMM.* The oneDNN catalog splits K 8-way for 1-8 token rows, 2-way to 128 and not at all above, so a
   request's rows get different bits depending on batch composition. The R221 kernel library pins a two-tier strategy
   (the natural 1-8-row entry for n <= 8, the 9-24 tile with the same 8-way K split above; bitwise equal on all 14 TP1/TP2
   shapes for n = 1..1024): decode unchanged, prefill GEMMs about 2x. Patches
   `experiments/qwen38-27b-b70/patches/onednn-qwen38-w4a16-{strategy-override-dump-r220,fixed-k-two-tier-r221}-20260905.patch`,
   build `experiments/qwen38-27b-b70/docker/{build-w4a16-strategy-r220-image.sh,rebuild-w4a16-incremental-r221.sh}`
   (the R139 flow: vllm-xpu-kernels 1e90ffa6 + r35 + r50, oneDNN 0e2a5bfe + r137a + r137b). `QWEN38_W4A16_FIXED_K=0`
   restores the catalog.
3. *FP16 linears in 32-row pieces* (R224, `docker/r224-fp16-linear-rowchunk.py`, `VLLM_XPU_FP16_LINEAR_ROWCHUNK`): the
   oneDNN f16 GEMM behind `lm_head` and `mtp.fc` keeps the single-row class only to 32 rows.
4. *Runtime switches:* Inductor `"split_reductions": false` (correction 2026-09-06: `VLLM_BATCH_INVARIANT` was recorded as 1 but every campaign container ran with 0, pinned by the strict launchers; vLLM refuses to boot the GDN backend with it set, so it is not part of this profile)
   in `inductor_compile_config` (size-independent reduction order for the compiled RMSNorm and other reductions);
   `VLLM_XPU_GDN_SPEC_GROUP` / `VLLM_XPU_GDN_PREFILL_GROUP` (R228/R236) group the GDN launches.

Images: `neural-download/vllm-openai-xpu:qwen38-int4-w4a16-fixed-k-r221` (699e2699, `_xpu_C` 271db0d4) -> ... ->
`...:qwen38-int4-draft-int4-head-r256` (f7696bca, the served image; GHCR
`vllm-openai-xpu-qwen38-int4@sha256:f7696bca...`). Earlier: `...:qwen38-int4-w4a16-fixed-k-r221` ->
`...:qwen38-int4-fp16-rowchunk-r224` (a23ff249) -> `...:qwen38-int4-gdn-spec-group-r228` (aaf920b0, `_xpu_ops.py`
c91d6b0d) -> `...:qwen38-int4-gdn-prefill-group-r236` (9488db61, `_xpu_ops.py` 015b4dce). Launch through the FP8 lane's
launchers with `MODEL_DIR=<relabel dir> MODEL_MANIFEST=<relabel manifest> QUANTIZATION=gptq VLLM_XPU_FP8_BLOCK_W8A16=0
VLLM_XPU_DRAFT_LM_HEAD_INT4=0 VLLM_XPU_W4A16_DETERMINISM_PAD=0 XPU_EXTENSION_SHA256_OVERRIDE=<_xpu_C sha256>
XPU_OPS_SHA256_OVERRIDE=<_xpu_ops.py sha256>` (see the R222-R239 wrappers under `experiments/qwen38-27b-b70/scripts/`).

**Single-request results, TP2 (R222/R227, R224 + batch-invariant):** MTP0 36.00 / 34.92 tok/s; MTP depth 1
50.83 / 51.12; depth 4 68.62 / 68.23; G1 (MTP0 pair), G2 (MTP pair) and G3 (each MTP server vs the MTP0 oracle) 12/12 at
every depth measured; G5 probe identical at 224/250/300 prompt tokens. Old ARK routing on the same stack: 32.8 tok/s and
no lossless speculation possible.

**Concurrency identity (c1-c64 ladder, 128 tokens per request, TP2):**

| configuration | MTP0 | MTP depth 4 |
|---|---|---|
| ARK routing (R216) | exact to c2 | exact to c2 |
| fixed-K kernel (R222) | exact to c32, c64 63/64 | c4 3/4, c8-c32 exact, c64 59/64 |
| + FP16 32-row pieces (R225) | c32 31/32, c64 64/64 | c8 7/8, c32 31/32, c64 60/64 |
| R226 (recorded as `VLLM_BATCH_INVARIANT=1`; the container actually ran with 0, so this is a repeat of the R224 configuration without the FA serial env) | **exact c1-c64** | exact c1-c16, c32 30/32, c64 59/64 |
| + Inductor `split_reductions=false` (R232) | exact c1-c64 | exact c1-c16, c32 31/32, c64 63/64 |

The remaining c32/c64 flips are a handful of near-tie prompts whose result depends on how the ladder's arrival timing
mixes prefills into the first steps (the MTP0 ladder with an unchanged configuration is all-exact in one run and one miss
at c32 or c64 in others). Every INT4 GEMM, FP16 linear and the attention decode are batch-invariant; the residual sits in
the GDN kernel's dependence on launch composition (grouping its launches, R229-R238, does not reproduce the
single-request arithmetic). Notes: `experiments/qwen38-27b-b70/notes/2026-09-05-qwen38-int4-{w4a16-fixed-k-two-tier-r220-r221,concurrency-identity-r222-r226}.md`;
data `experiments/qwen38-27b-b70/data/2026-09-05-qwen38-int4-*`.

**XPU graph capture (R246b/R247/R250, 2026-09-05 evening): the two-card headline.** With `VLLM_XPU_ENABLE_XPU_GRAPH=1`
and `cudagraph_mode FULL_DECODE_ONLY` (capture sizes 1-8) on the same final configuration, the captured verify step
no longer pays the per-op all-reduce host waits:

| depth | TP2 strict pair, graphs (tok/s) | TP2 eager (R239/R240) | gates |
|---|---|---|---|
| 4, + draft-only INT4 head (R257, R256 image) | **112.36 / 112.33** | 68.55 / 67.79 | G2 12/12, G3 12/12 x2 vs the eager MTP0 oracle; acceptance 3.51 |
| 5, + draft-only INT4 head (R258) | 109.97 / 110.07 | - | G2 12/12, G3 12/12 x2 |
| 6, + draft-only INT4 head (R258) | 108.34 / 108.37 | - | G2 12/12, G3 12/12 x2 |
| 0 (MTP0) | **49.83 / 49.89** | 34.21 / 35.64 | G1 12/12; 12/12 vs the eager R239 MTP0 oracle (R253) |
| 1 | 76.72 / 76.63 | 51.10 / 50.09 | G2 12/12, G3 12/12 x2, probe exact (R253) |
| 4 | **91.00 / 91.01** | 68.55 / 67.79 | G2 12/12, G3 12/12 x2 vs the eager MTP0 oracle |
| 5 | 88.84 / 89.12 | - | G2 12/12, G3 12/12 x2 |
| 6 | 84.06 / 83.98 | - | G2 12/12, G3 12/12 x2 |
| 4, TP1 | 56.91 / 56.90 | 56.29 / 56.25 | G2 12/12, G3 12/12 x2 |

Acceptance is unchanged by graph capture (3.00 per step at depth 4), so that gain is per-step time; on one card there
is no collective to remove. The draft-only INT4 lm_head (the FP8 lane's feature, unlocked for this lane by the R256
fallback, `docker/r256-draft-int4-head-fallback.py`) makes each draft pass a quarter of the bytes and, measured, raises
acceptance to 3.51: depth 4 goes from 91.0 to **112.4 tok/s**. The verifier head stays FP16; every gate stays 12/12. Graph replay is bit-identical to eager execution (every pair matched the eager oracle). The
launcher `scripts/run-fixed-k-mtp-server.sh` enables this by default (`XPU_GRAPH=0` restores eager). Two negatives from
the same evening: the community checkpoint's BF16 draft tensors (R244/R245) are lossless but accept fewer tokens than
the AutoRound INT4 draft (3.21 vs 3.52 per step on TP1; 52.0 vs 56.3 tok/s), and vLLM's generic `mtp` proposer is
identical in rate to `qwen3_next_mtp` (R248/R249). Concurrency identity under graph capture (R251, TP2): MTP0 exact at every level c1-c64 (aggregate 999.4 tok/s at c64);
depth 4 exact at c1, c2 and c16, benchmark-c003 (the composition-sensitive near-tie prompt) at c4 3/4 and c8 7/8,
c32 30/32, c64 58/64; depth 1 exact through c8 (358.5 tok/s), c16 12/16, c32 30/32, c64 61/64 (R253). Batches above
8 tokens run eagerly, so this is the same composition residual as the eager ladders, at a different mix. On the
headline configuration (graphs + draft-only INT4 head, R259): depth 4 exact through c8 (403 tok/s aggregate, 106.8 at
c1), c16 14/16, c32 30/32, c64 59/64; MTP0 exact except c32 31/32 (benchmark-c003; c64 64/64, 997.5 tok/s).

**Matrix R239, TP2 (2026-09-05, final configuration: R228 image + `split_reductions=false` (`VLLM_BATCH_INVARIANT=0`, see the correction above),
data `experiments/qwen38-27b-b70/data/2026-09-05-qwen38-int4-r239-matrix-result.json`):**

| depth | strict pair (tok/s) | gates | identity ladder c1 / c2 / c4 / c8 / c16 / c32 / c64 (aggregate tok/s at c16) |
|---|---|---|---|
| 0 (MTP0) | 34.21 / 35.64 | G1 12/12 | exact at every level in 3 of 4 ladders (c64 998.6 tok/s); one run c32 31/32, c64 63/64 |
| 1 | 51.10 / 50.09 | G2, G3 x2 12/12, probe exact | exact to c16 (650.1), c32 30/32, c64 58/64 |
| 2 | 61.14 / 61.54 | G2, G3 x2 12/12 | exact to c16 (522.3), c32 30/32, c64 60/64 |
| 3 | 67.61 / 67.83 | G2, G3 x2 12/12 | exact to c16 (599.4), c32 30/32, c64 59/64 |
| 4 | 68.55 / 67.79 (R240 re-run; first pass 68.22 / candidate b died at engine start) | G2, G3 x2 12/12 | c8 7/8, c16 exact (516.1), c32 30/32, c64 59/64 |

Depth 4 in R222/R227 on the same tensors and kernel: 68.62 / 68.23 with all gates 12/12. Aggregate ladder rates are
identity-qualified only where the level is exact.

**Matrix R239, TP1 (one card, `TENSOR_PARALLEL_SIZE=1 XPU_DEVICE_MASK=0 GPU_MEMORY_UTILIZATION=0.96`, same
configuration):**

| depth | strict pair (tok/s) | gates | identity ladder c1 / c2 / c4 / c8 / c16 / c32 / c64 (aggregate tok/s at c8) |
|---|---|---|---|
| 0 (MTP0) | 32.96 / 32.95 | G1 12/12 | exact at every level in 3 of 4 ladders (c64 447.6 tok/s); one run c64 63/64 |
| 1 | 49.64 / 49.47 | G2, G3 x2 12/12, probe exact | exact to c8 (261.6), c16 15/16, c32 30/32, c64 60/64 |
| 2 | 56.51 / 56.45 | G2, G3 x2 12/12 | exact to c8 (283.6), c16 15/16, c32 29/32, c64 62/64 |
| 3 | 58.47 / 58.47 | G2, G3 x2 12/12 | exact to c16 (270.3 at c8), c32 31/32, c64 61/64 |
| 4 | 56.29 / 56.25 | G2, G3 x2 12/12 | exact to c8 (216.3), c16 15/16, c32 31/32, c64 61/64 |

On one card the speculative turnover is at depth 3 (58.5 tok/s); on two cards depth 4 still gains (68.2). Single-card
MTP0 is within 5% of TP2 (33.0 vs 34.2-35.6), consistent with the launch-overhead-bound profile: the second card adds
almost nothing at one request and doubles aggregate throughput at concurrency. The single-card residual has the same
shape as TP2 without any collective, so the oneCCL all-reduce is excluded as a source.


New optimization lane, opened 2026-08-18, superseding the Qwen3.6 27B INT4
speculative lane. The two checkpoints have the same tensor architecture, so the
pinned Qwen3.6 source stack is mechanically compatible. New weights still
require independent numerical, quality, determinism, and performance gates.

> **Status correction, 2026-08-20:** this lane has no promoted record. The
> published `101.922` MTP5 and `100.497` MTP4 rows used an output-changing
> greedy margin and a baseline with the same setting; withdrawal is
> recommended. The honest margin-free working anchor is `101.170 tok/s`
> all-25, but its three arms agree on only 21–22/25 prompts. A fresh
> margin-free target-only oracle now exists, yet target A/B agreed on 24/25
> and a sealed-cache TP1 MTP5 pair agreed on only 2/4 diagnostic prompts. A
> later preregistered six-arm control produced two structured variants with
> the oneDNN INT4 prefill pad off and one shared variant in three pad-on arms.
> That passes the diagnostic criterion but does not establish lane-wide or
> full-25 TP2 determinism. Pad-on TP2 subsequently remained nondeterministic,
> including a recurring 512-zero final stream. A target/verifier post-forward
> sync arm observed a third endpoint family that still split at generated
> token 469; one treated arm cannot distinguish a sync effect from the lane's
> existing run-to-run variability. It is a negative boundary diagnostic, not a
> fix. A bounded prompt-24 replay-microscope arm then produced no trace because
> its anchored public request ID omitted vLLM's worker-side eight-hex suffix;
> prompt 6 also ended at 68 tokens, invalidating the strict metric window. That
> arm is an invalid false-null, not localization or speed evidence, and must not
> be retried. A later sealed graph-replay-bypass R1/R2 pair matched on 25/25
> token arrays, but its combined treatment also changed drafter geometry and
> startup allocation history; both arms remained only 18/25 exact versus target
> A and the pair central value was 56.363 tok/s. This is bounded diagnostic
> evidence, not a localized fix or performance candidate. Do not use the
> historical command below for a new promotion run. A later target-only split
> retained drafter PIECEWISE/M6 and both startup captures while bypassing only
> request-selected uniform target/verifier replay. Its two sealed arms matched
> only 24/25, splitting at prompt 24 token 469 between two sane historical
> families, and averaged only 60.938 tok/s preferred. That treatment is also
> terminal and insufficient; no retry or T3 is authorized.

**LocalMaxxing:** the depth-4 headline is approved as `cmtq3mx7e03gen701xxpwu67z` (112.344 tok/s, center of the R257 pair; ledger `results/localmaxxing-submissions.md`, attestation `experiments/qwen38-27b-b70/data/2026-09-06-qwen38-int4-r257-graph-drafthead-depth4-strict-result.json`).

## Optimization pass after the headline (2026-09-06 night, R261-R277)

Everything below was measured on the headline stack (R256/R228 images, TP2, XPU graph capture, draft-only INT4 head, MTP
depth 4, strict launcher env) unless stated; full detail in
[`notes/2026-09-06-qwen38-int4-optimization-night-r261-r268.md`](../../experiments/qwen38-27b-b70/notes/2026-09-06-qwen38-int4-optimization-night-r261-r268.md)
and the result files `2026-09-06-qwen38-int4-r262-headline-decode-profile-result.json` and the R261-R277 entries of
`2026-09-05-qwen38-int4-graph-capture-tp2-mtp4-r247-result.json`.

**Single user is at the practical floor of this stack.** The step is ~27 ms: ~12 ms of weight reads for the 5-token verify,
~6 ms for the four draft passes with the INT4 head, and ~8 ms of two-rank host round trip. Tested and equivalent or worse:

| lever | result | run |
|---|---|---|
| XPU graph capture sizes extended to 80 tokens | no change at any concurrency | R261 |
| vLLM V2 model runner (fused draft loop + prefill in FULL graphs), draft-only INT4 head ported to its loader (`docker/r266-v2-draft-int4-head.py`) | **112.70 / 112.96 tok/s, 12/12 vs the oracle**: lossless, speed-equal to V1 (112.36 / 112.33); identical text and acceptance on a shared prompt; no first-batch recompile stall at c2 | R263/R268/R269/R270 |
| `--async-scheduling` | already on by default for MTP in this vLLM (identical rates) | R267 |
| per-rank CPU binding (`--numa-bind --numa-bind-cpus`) | no change (each worker's main thread already spins on its own core) | R272 |
| GDN speculative group size 64 / 1 | 64: no change at c32; 1: uniformly slower | R273 |

**Multi-user findings.**

- Single-pass ladders under-report c2: the first two-sequence batch costs ~1.05 s of Dynamo recompilation (R264b). With a
  warm pass (R265b, every rung run twice) the steady curve is c2 147.6, c4 252.9, c8 405.4, c16 580.4 tok/s, identity-exact
  through c16 (16/16); c32 360.7 (31/32) and c64 445.3 (58/64) keep the near-tie residual. Warm-up the server with one
  two-user request after boot, or run ladders with `LADDER_REPEATS=2`.
- The c32/c64 rungs of the earlier ladders (~360 / ~445 tok/s) were an artifact, found by profiling (R278k): the launcher
  chain did not forward `VLLM_XPU_W4A16_DETERMINISM_PAD`, and the R213b overlay defaults to padding 128-512 verify rows to
  512 before the W4A16 GEMM (529 vs 180 us per launch; 163 vs 55 ms of GEMM per step at c32). Nothing below 128 rows
  (c1-c16, the single-user headline, every gate) was touched. Both launchers now forward the switch with 0 as the default;
  the corrected ladder through the runner (R281, warm pass) is c2 156, c4 268, c8 424, **c16 584 (16/16)**, c32 633
  (30/32), c64 604 (62/64, admission-limited at max-model-len 256). With the R276 image and capture sizes to 320 (R282):
  c2 191, c4 295, c8 423, c16 579, c32 635, c64 589, and its strict pair is lossless (R283: 112.90 / 113.00, 12/12), so
  R276 is the served image from here on.
- Below that, the device-side costs at c32 measured in R274/R274b: the W4A16 verify GEMMs are dequant-bound and
  scale ~linearly above ~32 rows (55 ms at M=160), and the GDN speculative recurrent kernel streams each sequence's fp32
  SSM state per layer (27 ms at 32 sequences, linear in sequences). The grouped GDN branch (group 16) added ~40 ms of host
  work per step and, because it syncs (`.tolist()`), cannot be graph-captured: no verify batch above 16 sequences was ever
  captured (R275 fails capture at 320 tokens). `docker/r276-gdn-spec-group-sync-free.py` computes the group boundaries
  arithmetically (uniform k+1 rows) so capture succeeds with the published group size.

### Many users: the c64 rung was never admission-limited (R284/R286/R287, 2026-09-06)

R284 re-ran the R282 ladders with max-num-seqs 128, max-model-len 512 and
max-num-batched-tokens 1024 (rungs 16 to 128; two passes; R276 image), and
R286/R287 repeated them at MTP depth 2 and depth 1. Warm-pass aggregate
tok/s with the exact-output count against each request's sequential oracle:

| users | no speculation | depth 1 | depth 2 | depth 4 |
| --- | --- | --- | --- | --- |
| 2 | 94.6 (2/2) | 147.3 (2/2) | - | 191.0 (2/2) |
| 4 | 178.5 (4/4) | 268.6 (4/4) | - | 294.7 (4/4) |
| 8 | 327.0 (8/8) | 456.0 (7/8) | - | 422.6 (7/8) |
| 16 | 533.8 (16/16) | 711.0 (16/16) | 591.0 (16/16) | 574.3 (16/16) |
| 32 | 815.0 (32/32) | 854.4 (32/32) | 723.4 (29/32) | 641.3 (32/32) |
| 64 | 991.4 (64/64) | 842.0 (61/64) | 815.2 (59/64) | 591.4 (58/64) |
| 96 | 915.4 (95/96) | 906.3 (92/96) | 777.8 (90/96) | 590.6 (89/96) |
| 128 | 1085.3 (128/128) | 894.8 (121/128) | 815.0 (121/128) | 584.7 (117/128) |

Every speculative depth plateaus once the verify batch exceeds about 32 rows,
and the non-exact requests above 32 users are the same prompts diverging at
the same token positions in every rung (for example `capacity-c006` at token
11), the signature of near-tie flips in the >32-row W4A16 GEMM tier rather
than anything in the harness. The 2/4/8-user rows for depth 1 and no speculation are R288 on the same
server shape; the depth-4 rows are R282 (max-num-seqs 64). Guidance: depth 4
for one to four users (112 tok/s single-user, 295 at four), depth 1 from
about 8 to 32 users (854 tok/s at 32, exact), and no speculation beyond that
(992 tok/s exact at 64). Entries
`R284_ladders_tp2_big_admission_mtp4_vs_mtp0`, `R286_…_mtp2` and
`R287_…_mtp1` in the graph-capture result JSON carry every rung and the
per-request divergence positions. A larger prefill budget does not help:
with max-num-batched-tokens 4096 and max-num-seqs 256 (R290) the same
near-tie prompts flip even without speculation from 64 users (57/64 in both
passes) and the aggregate is no better (1032.8 at 64, 966.8 at 128, 1021.3 at
256 users with 23 s first-token waits), so the published shape stays at 1024.

### One card, many users (R285, 2026-09-06)

The same large-admission ladders on one B70 (TP1): depth 4 saturates at about
245 tok/s from 16 users (c16 237.3 exact 16/16, c32 243.7 exact 32/32, then
62/64, 90/96, 123/128 with the same near-tie pattern), while MTP0 on the same
card is exact at every rung in both passes and peaks at 514.0 tok/s at 32
users (c64 418.5, c96 403.7, c128 443.8, all 128/128). One card therefore
wants depth 4 only up to about 8-16 users; two cards roughly double the
many-user ceiling (992 tok/s exact at 64 users) rather than the single-user
rate. Entry `R285_ladders_tp1_big_admission_mtp4_vs_mtp0` in the same JSON.

### The FP16 linear pieces were a throughput tax: R293, the class-consistent FP16 linear (R295-R298, 2026-09-11)

Every unquantized FP16 linear in this stack - the 2.5 GB `lm_head` (248320 x 5120 per card pair, a 1.27 GB shard per
card at TP2) and the MTP `fc` - ran in `<=32`-row pieces (R224) so that the oneDNN f16 GEMM stayed in its single-row
rounding class. Each piece re-reads the weight. The Qwen3.5 lanes found the size of that cost first (a 4B lost 41-54% of
depth-3 throughput above 32 rows; a 9B similar), measured the GEMM's row-count classes offline (rows 1-32, 33-128,
129-320, 321-512 at these shapes, each position-invariant, pad-invariant and deterministic), and built **R293**: the same
opaque op, with `VLLM_XPU_FP16_LINEAR_CLASSPAD=1`, measures the class map for each weight shape on first use, verifies
each class before trusting it, and pads or splits every call into one canonical class - one weight read per step, rows
bit-identical between one user and sixty-four. Mechanism and census:
[`experiments/qwen35-4b-b70/notes/2026-09-09-the-fp16-linear-chunk-is-a-throughput-tax.md`](../../experiments/qwen35-4b-b70/notes/2026-09-09-the-fp16-linear-chunk-is-a-throughput-tax.md)
and [`2026-09-11-r293-class-consistent-fp16-linear-on-the-server.md`](../../experiments/qwen35-4b-b70/notes/2026-09-11-r293-class-consistent-fp16-linear-on-the-server.md).
Image: R276 + `docker/r290..r293-*.py` (`neural-download/vllm-openai-xpu:qwen38-int4-fp16-linear-classpad-cheapest-r293`,
id `sha256:40d46730...`); with the switch off it runs the R276 code path unchanged.

On this lane the projection is a small share of a 19 GB model whose step is dominated by the INT4 GEMMs, so the gain is
real but modest. Measured on the served configuration (TP2, depth 4, sizes to 320, INT4 draft head), every server's log
carrying its census map (`R291 classpad census ... verdict=classpad`):

| | R276 (R282/R283) | R293 (R295-R297) |
| --- | ---: | ---: |
| strict pair, depth 4 (G2 12/12, G3 12/12 x2 vs its own fresh MTP0 pair) | 112.90 / 113.00 | 111.69 / 111.33 |
| strict pair, MTP0 (G1 12/12) | 49.83 / 49.89 (R253) | 49.39 / 49.39 |
| depth 4 ladder, warm pass: c8 / c16 / c32 / c64 | 422.6 / 578.9 / 634.5 / 589.0 | 436.1 / 611.1 / 682.6 / 627.8 |
| depth 4 identity: c16 / c32 / c64 | 16/16, 30/32, 60/64 | 16/16, 30/32, 59/64 |
| MTP0 ladder, warm pass: c32 / c64 | 815.8 / 989.8 | 815.4 / 1019.2 |
| MTP0 identity c1-c64 | exact | exact (128/128 at c64) |
| big admission (mns 256, mbt 4096), MTP0 c256 / depth 4 c128 | 1021.3 / 584.7 (R290 / R284) | 1079.5 / 647.7 |
| two cards, MTP0, c64, 5 ms admission stagger, ten passes | not measured | **640/640**, `output-identity-qualified`, 1014.4 |

Depth 4 gains 3-8% from c8 up and MTP0 3-6% from c64 up, at a 1% single-user cost; identity is unchanged at every rung.
The strict gate is regenerated on the same image (G1 on a fresh MTP0 pair, G3 against it): under R293 the projection
runs in a different rounding class from the eager R239 oracle the R283 gate used, so that stored oracle is not the
reference for this image. The staggered-admission recipe from the 4B lane (`--launch-stagger-ms 5`) was run here for the
first time: ten passes of 64 concurrent requests byte-identical to the sequential oracle, where the published claim had
been two passes.

`CLASSPAD=1` on the package launcher enables it; the default stays `0` (the published single-user configuration). One
card (R298) is in the replication matrix above.

### Kernel-library build reproducibility (clean-clone replay, 2026-09-06)

The published R220 + R221 build scripts were re-run from fresh clones of the
pinned commits (vllm-xpu-kernels `1e90ffa6`, oneDNN `0e2a5bfe`, sycl-tla
`cd763790`) with the in-repository patches. The rebuilt `_xpu_C` matched the
shipped extension bit-for-bit (`271db0d4…`) and both chain images reproduced
their ids (R220 `36360702…`, R221 `699e2699…`), so the kernel library in every
image above is reproducible from source, not only from the GHCR digest. Two
script defects found by the replay (a stale hash pin on the builder helper and
a `grep -q` under `pipefail`) were fixed in place; see the manifest's
`known_corrections` and
`experiments/qwen38-27b-b70/data/replay-w4a16-r220-r221-clean-clone-20260906T192527Z/`.

### Real-content 2K-32K depth ladder on the headline configuration (R260b, 2026-09-06)

Same harness as the FP8 lane's R189 (`bench-w8a16-real-content-depth.sh`, fixture `qwen38-bce40ca-mixed-content-depth-v1`): one slot,
`--max-model-len 33024 --max-num-seqs 1 --max-num-batched-tokens 4096`, exact active context of unrepeated technical prose, Python
code and structured documents, 128 returned token IDs, cache zero, canaries before and after each arm. R256 image, TP2, XPU graph
capture, draft-only INT4 head, strict launcher env. The MTP0 arm is the oracle; every MTP depth-4 array equalled it (18/18).
Per-depth values are medians of the three classes; the MTP4 class spread is wide because Python code accepts more draft tokens.

| active context | MTP0 tok/s | MTP0 TTFT s | MTP4 tok/s | MTP4 TTFT s | MTP4 prose / Python / docs | MTP4 vs MTP0 oracle |
|---|---|---|---|---|---|---|
| 2,048 | 50.02 | 0.62 | 120.56 | 0.62 | 110.9 / 183.0 / 120.6 | 3/3 exact |
| 4,096 | 48.78 | 1.19 | 125.72 | 1.22 | 87.3 / 125.7 / 153.3 | 3/3 exact |
| 8,192 | 47.46 | 2.42 | 124.35 | 2.48 | 124.4 / 169.3 / 111.5 | 3/3 exact |
| 16,384 | 45.92 | 5.02 | 90.18 | 5.17 | 85.3 / 156.2 / 90.2 | 3/3 exact |
| 24,576 | 44.17 | 7.84 | 84.63 | 8.09 | 82.7 / 151.7 / 84.6 | 3/3 exact |
| 32,768 | 42.83 | 10.84 | 100.27 | 11.19 | 100.3 / 134.2 / 94.3 | 3/3 exact |

Evidence: [`2026-09-06-qwen38-int4-r256-real-content-depth-r260b-result.json`](../../experiments/qwen38-27b-b70/data/2026-09-06-qwen38-int4-r256-real-content-depth-r260b-result.json); wrapper
`experiments/qwen38-27b-b70/scripts/run-20260906-qwen38-int4-r256-real-content-depth-r260.sh`. The first attempt (R260) aborted at the
MTP4 arm because the public launcher then exported `VLLM_BATCH_INVARIANT=1`, which vLLM rejects for the GDN backend; see the correction above.

### Standalone `docker run` (no lab scripts)

The launcher above is the contract-checked path (it routes through the FP8 lane's strict launchers, which pin the process-level determinism env below). Served image since 2026-09-06 afternoon: **R276** (`@sha256:521eb277…`, R256 plus the sync-free grouped GDN branch, `docker/r276-gdn-spec-group-sync-free.py`), which lets XPU graph capture cover verify batches above 16 sequences (`cudagraph_capture_sizes` to 320): lossless (R283 strict pair 112.90 / 113.00, 12/12 vs the oracle) and +22% / +10% at two / four users (R282). The R228 / R256 images run the same code for the single-user headline but must keep capture sizes 1-8 (their grouped branch syncs during capture; `XPU_GRAPH_SIZES=8` in the launcher). This is the same container invocation written out, for anyone replicating the
headline profile (TP2, MTP depth 4, XPU graphs, draft-only INT4 head) without cloning the lab. `MODEL_DIR` is the gptq-relabelled
directory produced by `scripts/make-gptq-relabel.py` (hard links to the AutoRound checkpoint plus a rewritten `config.json`;
manifest `manifests/model-gptq-relabel-r212.json`). Change `--tensor-parallel-size`, `ZE_AFFINITY_MASK`/`ONEAPI_DEVICE_SELECTOR`
(`0` for one card), `num_speculative_tokens` (drop `--speculative-config` for MTP0) and `VLLM_XPU_ENABLE_XPU_GRAPH`/the
`cudagraph_mode` block (`"PIECEWISE"`, sizes `[1]`, for graphs off) to reach the other rows of the replication matrix.

> The `gptq-relabel` directory below is the path this block used when the numbers were
> taken; it is not present on the lab host today. The live local copy of the same pinned
> revision is `/mnt/usb-models/llm-models/qwen3.8-27b-int4-autoround-devan` (see the
> model-identity section), so set `MODEL_DIR` to that unless you have reconstructed the
> relabelled copy yourself. Checked 2026-09-08.

```bash
MODEL_DIR=/mnt/fast-ai/llm-models/qwen3.8-27b-int4-autoround-gptq-relabel
CACHE_DIR=$(mktemp -d)   # empty compile cache; every published number was taken from an empty cache
docker run --rm --name qwen38-int4-fixed-k-mtp4 --ulimit core=0 \
  --memory 12g --memory-swap 16g --device /dev/dri:/dev/dri --group-add render \
  --cap-add SYS_PTRACE --security-opt label=disable --ipc=host --shm-size=8g \
  --publish 127.0.0.1:18134:8000 --volume "${MODEL_DIR}:/model:ro" --volume "${CACHE_DIR}:/root/.cache/vllm" \
  --env ZE_AFFINITY_MASK=0,1 --env ONEAPI_DEVICE_SELECTOR=level_zero:0,1 \
  --env VLLM_TARGET_DEVICE=xpu --env VLLM_WORKER_MULTIPROC_METHOD=spawn \
  --env VLLM_XPU_ENABLE_XPU_GRAPH=1 --env VLLM_XPU_ALLREDUCE_HOST_WAIT=1 \
  --env VLLM_XPU_FP8_BLOCK_W8A16=1 --env VLLM_XPU_W4A16_DETERMINISM_PAD=0 \
  --env VLLM_XPU_GEMMA_RMSNORM_TRITON=0 --env VLLM_XPU_RMSNORM_TRITON=0 \
  --env VLLM_XPU_GDN_SPLIT_MIXED=1 --env VLLM_XPU_GDN_SPEC_GROUP=16 --env VLLM_XPU_GDN_PREFILL_GROUP=1 \
  --env VLLM_XPU_FP16_LINEAR_ROWCHUNK=32 --env VLLM_BATCH_INVARIANT=0 --env VLLM_XPU_GDN_SPEC_PERSISTENT_SCRATCH=1 \
  --env VLLM_XPU_QWEN_GEMMA_RMSNORM_BATCH_INVARIANT=0 --env VLLM_XPU_QWEN_GEMMA_RMSNORM_PACKED_SERIAL_EXACT=1 \
  --env VLLM_XPU_DRAFT_LM_HEAD_INT4=1 --env VLLM_XPU_DRAFT_LM_HEAD_INT4_GROUP_SIZE=128 --env VLLM_XPU_DRAFT_LM_HEAD_INT4_SCALE_DTYPE=bf16 \
  --env VLLM_XPU_DRAFT_LM_HEAD_INT4_CHUNK_ROWS=2048 --env VLLM_XPU_GDN_NATIVE_FALLBACK=1 \
  --env TORCHINDUCTOR_DETERMINISTIC=1 --env VLLM_ENABLE_INDUCTOR_MAX_AUTOTUNE=0 --env VLLM_ENABLE_INDUCTOR_COORDINATE_DESCENT_TUNING=0 --env PYTHONHASHSEED=0 \
  --env PYTORCH_ALLOC_CONF=expandable_segments:True \
  --env CCL_ATL_TRANSPORT=ofi --env FI_PROVIDER=tcp --env FI_TCP_IFACE=lo --env CCL_ZE_IPC_EXCHANGE=pidfd \
  --env CCL_SEND=direct --env CCL_RECV=direct --env CCL_TOPO_P2P_ACCESS=1 \
  --env CCL_SYCL_ALLREDUCE_SIMPLE_THRESHOLD=4294967296 --env CCL_SYCL_ALLGATHERV_SIMPLE_THRESHOLD=4294967296 \
  --env CCL_SYCL_REDUCE_SCATTER_SIMPLE_THRESHOLD=4294967296 \
  ghcr.io/steveseguin/vllm-openai-xpu-qwen38-int4@sha256:521eb277c0733f8c2ce47aea1bb98ed576c6f1ad63bf5baf22d38fc07abf54ad \
  --model /model --served-model-name qwen38-int4-fixed-k-mtp4 --host 0.0.0.0 --port 8000 \
  --tensor-parallel-size 2 --dtype float16 --quantization gptq --kv-cache-dtype auto \
  --gpu-memory-utilization 0.95 --max-model-len 1024 --block-size 64 \
  --max-num-seqs 1 --max-num-batched-tokens 1024 \
  --no-enable-prefix-caching --enable-prompt-tokens-details --language-model-only \
  --speculative-config '{"method":"qwen3_next_mtp","num_speculative_tokens":4}' \
  --compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY","cudagraph_capture_sizes":[1,2,3,4,5,6,8,10,15,16,20,25,30,32,40,50,60,64,80,100,120,160,200,240,320],"max_cudagraph_capture_size":320,"splitting_ops":[],"inductor_compile_config":{"combo_kernels":false,"benchmark_combo_kernel":false,"deterministic":true,"benchmark_epilogue_fusion":false,"split_reductions":false,"triton.autotune_pointwise":false}}'
```

Then `curl -s localhost:18134/v1/completions -d '{"model":"qwen38-int4-fixed-k-mtp4","prompt":"...","max_tokens":512,"temperature":0}'`.
The strict suite (`repro/qwen36-27b-autoround-int4-b70/realistic-suite-v1.json`, 12 prompts, six classes, 512-token cap) and
the concurrency ladder client (`scripts/bench-openai-concurrency-oracle.py`) are the only clients used for the published numbers;
`max-model-len 256 / max-num-seqs 64 / max-num-batched-tokens 512` is the ladder server shape. The lab's launcher additionally
verifies the image id and the digests of `_xpu_C.abi3.so`, `_xpu_ops.py` and `layernorm.py` inside the container
(`verify-image-contract.sh`); the registry digest above pins the same bytes.

## Model

`devan-carlin/Qwen3.8-27B-int4-AutoRound`, base `Qwen/Qwen3.8-27B`, Apache-2.0.
Local copy: `/mnt/usb-models/llm-models/qwen3.8-27b-int4-autoround-devan`
(19.02 GB, 8 safetensors shards + 11 small files, all verified).

Content manifest: [`manifests/model.json`](manifests/model.json). The upstream
revision is pinned at
[`bce40cacab0a4535b92fb3d57615c2bea9adf3d1`](https://huggingface.co/devan-carlin/Qwen3.8-27B-int4-AutoRound/tree/bce40cacab0a4535b92fb3d57615c2bea9adf3d1).
An independent download at
`/mnt/fast-ai/llm-models/qwen3.8-27b-int4-autoround` matched all recorded
file identities and the exact `19,016,930,167`-byte payload.

## Why the existing stack transfers

| Property | Qwen3.6 (old lane) | Qwen3.8 (this lane) |
| --- | --- | --- |
| architecture class | `Qwen3_5ForConditionalGeneration` | same |
| text `model_type` | `qwen3_5_text` | same |
| layers / hidden | 64 / 5120 | same |
| **vocab size** | 248320 | **same** |
| attention heads / KV heads | 24 / 4 | same |
| linear key / value heads | 16 / 48 | same |
| `full_attention_interval` | 4 | same |
| `mtp_num_hidden_layers` | 1 | same |
| `quant_method` | `auto-round` | same |
| bits / group size / packing | 4 / 128 / `auto_round:auto_gptq` | same |

Consequences:

- vLLM routes `auto-round` to `INCConfig`
  (`vllm/model_executor/layers/quantization/__init__.py:164`), so the whole INT4
  W4A16 path applies: `int4_gemm_w4a16`, the oneDNN completion barriers and
  input-dependency controls, and the INT8 LM head.
- `Qwen3_5ForConditionalGeneration` and `Qwen3_5MTP` are both registered
  (`registry.py:564`, `:634`), and the checkpoint ships **29 MTP tensors**
  (`mtp.fc.weight`, `mtp.layers.0.*`, `mtp.norm.weight`), so MTP speculative
  decoding is available.
- The vocabulary is byte-for-byte the same size, so the masked-max greedy
  sampler fix carries its full benefit.
- The upstream README warns that mixed symmetric/asymmetric INT4 checkpoints
  need devan-carlin's empty/shape-compatible qzeros guard on the newer XPU/ARK
  path ([vLLM PR #52428](https://github.com/vllm-project/vllm/pull/52428)).
  The pinned older vLLM tree has only a present/non-null check at
  `inc.py:822-825`, not that newer guard. This exact checkpoint nevertheless
  loaded successfully in the recorded baseline. Do not generalize that success
  to a different AutoRound export or claim PR #52428 is already present.

## Reference point

The model author measured **47.8 tok/s** on 4x B70 at TP=4, no speculation,
`max_tokens=16384` (versus 30.2 for BF16). That is not comparable to this lane's
TP=2 + MTP3 configuration, and it is far below what the identical Qwen3.6
architecture reaches here (~95 tok/s), so it should be treated as a floor, not a
target.

## Running an arm

Run the read-only [`scripts/preflight.sh`](scripts/preflight.sh) first. It
checks the pinned sources, Python package family, retained graph-safe
FlashAttention and oneCCL identities, complete model manifest, host memory, and
two-card inventory without importing torch or opening a GPU. The outstanding
reference-host portability and low-RAM evidence is tracked in
[`REFERENCE-HOST-HANDOFF.md`](REFERENCE-HOST-HANDOFF.md).

The Qwen3.6 harness is reused directly. Two environment variables retarget it:

```bash
MODEL_DIR=/mnt/usb-models/llm-models/qwen3.8-27b-int4-autoround-devan \
VALIDATION_MODEL_MANIFEST=$repo/repro/qwen38-27b-autoround-int4-b70/manifests/model.json \
VALIDATION_EXPECT_XPU_COUNT=2 \
VALIDATION_EXPECT_VLLM_VERSION=0.21.1rc1.dev289+g44fc8fde0 \
VALIDATION_HF_HOME=/mnt/fast-ai/llm-cache/hf \
VALIDATION_XPU_RUNTIME_MANIFEST=/path/to/xpu-runtime.sha256 \
VALIDATION_ONECCL_MANIFEST=/path/to/oneccl-runtime.sha256 \
VALIDATION_GRAPH_STAGE_MANIFEST=/path/to/graph-stage.sha256 \
  ...run-arm.sh spec-native-partition-exact-native 0,1 "$root" "$baseline"
```

`VALIDATION_MODEL_MANIFEST` was added for this lane and defaults to the Qwen3.6
manifest, so every existing Qwen3.6 arm is unaffected. The device-count
override keeps the recorded four-B70 default fail-closed while permitting this
explicitly identified two-B70 host; the runner still requires the requested
TP2 pair and logs the expected count in the run identity.
The version override is the metadata emitted by a fresh editable build of
source `44fc8fde09` on 2026-08-18. The historical environment retained an older
distribution label even after its editable source moved to that commit; source
head and diff remain the authoritative code identity in either case.
This two-card host's pinned low-memory XPU rebuild, package manifest, peak-RSS
warning, and import-path check are recorded in
[`RUNTIME-BUILD-20260818.md`](RUNTIME-BUILD-20260818.md).
The matching public oneCCL build, checksum manifest, and passing two-rank graph
oracles are recorded in
[`ONECCL-BUILD-20260818.md`](ONECCL-BUILD-20260818.md).
The matching model-specific graph-safe attention stage passed 12,000 replay
oracles across both B70s. Its complete loadable-package manifest is
[`manifests/graph-stage-qwen38-head256-oneapi2025.3.3-20260818.sha256`](manifests/graph-stage-qwen38-head256-oneapi2025.3.3-20260818.sha256),
and the replay summary is
[`evidence/graph-stage-oracles-20260818.json`](evidence/graph-stage-oracles-20260818.json).

The three optional runtime manifests make a rebuilt host identity explicit
without weakening the retained reference-host defaults. Entries are
`SHA256 relative/path` pairs rooted at, respectively,
`BASE_STAGE/vllm_xpu_kernels`, `ONECCL_INSTALL_DIR`, and `STAGE`. Absolute or
parent-traversing paths, missing files, malformed hashes, and empty manifests
all fail closed. Omit these variables only when reproducing the historical
binary hashes embedded in the validator.
`VALIDATION_HF_HOME` prevents the historical `/mnt/usb-models` cache path from
being recreated on hosts that keep their model and transient Hugging Face
metadata elsewhere; it does not change model identity verification.
The runner also places the verified `BASE_STAGE` first on `PYTHONPATH`; this is
required to keep a rebuilt `_xpu_C` paired with its matching device libraries
instead of silently importing `_xpu_C` from an installed wheel.

The known-good deterministic configuration and its flag set are documented in
[`../qwen36-27b-autoround-int4-b70-determinism-20260818/README.md`](../qwen36-27b-autoround-int4-b70-determinism-20260818/README.md)
section 7a; start from that rather than re-deriving it.

The 2026-08-18 transition pull was independently audited on the two-card,
15 GiB host. Its Qwen3.6 closeout bundle and flat patch are complete and
reconstruct the recorded vLLM tree, while the Qwen3.8 runtime and raw-evidence
handoff remains incomplete. See the
[transition audit](../../experiments/qwen38-27b-b70/notes/2026-08-18-autoround-int4-transition-handoff-audit.md).

## Open items

- A valid fresh margin-free target-only quality oracle now exists at
  `qwen38-marginfree-targetoracle-25-a-20260820`. Its A/B throughput was
  `49.759` / `50.016 tok/s`, but the pair agreed on only 24/25 prompts: long
  rollover diverged at token 469. Use A for the semantic baseline while
  retaining that target-only determinism caveat.
- The vision tower (333 tensors) is unused for text benchmarking; the config
  carries `language_model_only`.
- The current margin-free MTP5 anchor is `101.170 tok/s` all-25 and `92.851`
  selection-12, the median of three arms. Pairwise token parity is 21/25,
  21/25, and 22/25, so it is a research baseline rather than a result.
- Post-recovery dual-view-verified MTP5 arms reached `102.132` and
  `102.176 tok/s`, but agreed on 21/25 and each matched target oracle A on only
  15/25. A byte-identical sealed-cache TP1 pair agreed on only 2/4. This proves
  runtime nondeterminism without TP2 collectives. A preregistered six-arm TP1
  control then produced structured variants `G/F2/G` with the global oneDNN
  INT4 prefill pad off and `G/G/G` with it on, under the same binary and sealed
  cache. The follow-up pad-on composite TP2 full-25 pair passed the new
  fail-closed engagement/direct-load/cache/freshness/quality gates but agreed
  on only 22/25 token arrays. A2's long-rollover response was all-zero from the
  first token, while B2 was sane. The exact C1 recurrence arm repeated A2's
  512-zero stream and formed third SQL/factual output families under the same
  sealed identity. Its preferred median was `101.059 tok/s`; all three arms are
  nonpromotable. A sealed graph-replay-bypass R1/R2 pair subsequently matched
  all 25 arrays and emitted the sane S1/target-A prompt-24 family, but only
  under a combined target-verifier replay, drafter graph/geometry, and startup
  allocation treatment. Both arms were 18/25 exact versus target A and their
  preferred central value was `56.363 tok/s`, 44.263% below B2. Preserve the
  pair as bounded diagnostic evidence, run no further arm under its
  preregistration, and do not promote it. The draft-fallback-margin path is
  also closed: its real 598-call TP2 qualification exceeded the required error
  bound on every call and left 9 repaired argmax mismatches versus full FP16.
  Do not retry it or run a full-25 throughput A/B; see the
  [terminal result](../../experiments/qwen38-27b-b70/notes/2026-08-20-draft-margin-tp2-qualification-result.md).
- Do not use stock `intel/llm-scaler-vllm:0.21.0-b3.1` as a substitute for the
  pinned source stack on a 16 GB host. An independent eager TP2 smoke first hit
  its FP8-only GDN output-projection probe on an INT4 `qweight`; disabling that
  optional probe allowed all weights to load, but a 9 GiB cgroup then killed a
  worker during warmup and triggered one BCS reset. See the
  [safety note](../../experiments/qwen38-27b-b70/notes/2026-08-18-autoround-int4-stock-image-lowram-unsafe.md).

## Invalidated historical measurement — 101.922 tok/s, MTP5 (2026-08-18)

LocalMaxxing `cmszbkxco0e11ms01l2rixxbt`; withdrawal recommended. Median of three cold arms
(`100.896` / `102.042` / `101.922`) on the 25-prompt suite; all three pairwise
comparisons were 25/25 only because a `0.03125` greedy margin masked runtime
flips and changed output on 18/25 prompts. Its quality baseline used the same
margin, so the quality pass is invalid. **Selection-12 was `95.167`** — lower than the MTP4 row's `96.627`,
because depth helps the newer holdout prompts and hurts the historical ones.

Identical to the MTP4 command below except:

```bash
VALIDATION_NUM_SPECULATIVE_TOKENS=5 \
VALIDATION_COMPILATION_CONFIG_OVERRIDE='{"use_inductor_graph_partition":true,"pass_config":{"fuse_rope_kvcache_cat_mla":false},"cudagraph_mode":"PIECEWISE","cudagraph_capture_sizes":[6],"max_cudagraph_capture_size":6}' \
```

The historical depth sweep measured MTP3 `96.616`, MTP4 `100.497`, MTP5
`101.922`, and MTP6 `99.464`, all under the invalid margin-assisted identity.
It does not establish a valid optimum for the margin-free lane.

### What a third party cannot reproduce exactly

Two things are honestly out of reach without host access:

1. **The staged graph-safe FlashAttention binaries.** These are a 3.1 GB AOT
   SYCL package that is not published and cannot be rebuilt bit-identically
   across toolchains. Build your own from
   `experiments/qwen27_graphsafe_flash_attention/` (`build.sh` plus the four
   patches) and run with:

   ```bash
   VALIDATION_ALLOW_UNPINNED_BINARIES=1 ...run-arm.sh ...
   ```

   That downgrades every binary-hash mismatch to a loud warning and writes the
   actual hashes to `binary-identity.txt` in the arm root, so your run stays
   auditable against what was really loaded. **It relaxes no correctness gate** —
   freshness, cache-zero, determinism and quality all still apply. Runs made
   this way are reproductions, not record-identity runs, and should not be used
   to promote a submission.
2. **The torch.compile cache.** The historical margin-on artifact was evaluated
   against a pinned compile cache, but that does not establish determinism for
   the current margin-free lane. Even three arms sharing one cache agreed on
   only 21/25, 21/25, and 22/25 prompts. The later TP1 F2/G pair began and
   ended on a byte-identical cache tree but still agreed on only 2/4 prompts,
   so exact token repeatability is a confirmed runtime problem rather than an
   untested cache hypothesis.

Everything else — model manifest, both source trees on the public forks and
the harness — is published. The historical quality baseline is published for
audit only and is not a valid margin-free oracle.

## Invalidated historical measurement — 100.497 tok/s (2026-08-18)

This three-arm MTP4 measurement is invalid for the same reason as MTP5: the
greedy margin changed emitted tokens and the quality baseline shared it. The
historical numbers remain below for audit, not promotion. Full original analysis:
[`../../notes/2026-08-18-qwen38-int4-100tps-uninitialized-gdn-scratch.md`](../../notes/2026-08-18-qwen38-int4-100tps-uninitialized-gdn-scratch.md).

| Arm | all-25 | selection-12 |
| --- | ---: | ---: |
| A | `101.653` | `96.499` |
| B | `100.497` | `96.627` |
| C | `99.905` | `96.895` |
| **median** | **`100.497`** | **`96.627`** |

Carry these caveats with the number: arm C is below 100 so the arms are not
unanimously over the line, the median is; **selection-12 at `96.627` has not
crossed 100**, and that is the subset any record comparison rests on; and this
is the pinned-compile-cache gate, with a fresh-compile arm still outstanding.

### Historical command identity — do not use for a new measurement

This command intentionally preserves the invalid margin-assisted identity so
the old artifact can be audited. It must not be copied into a new run. A new
run requires margin `0`, persistent scratch `1`, and a fresh target-only
quality oracle; those gates are not yet represented by a promoted command.

```bash
repo=$(git -C . rev-parse --show-toplevel)
LABEL=qwen38-mtp4-noscratch-repro-a
root=${BENCH_ROOT:-/mnt/usb-models/bench-results/qwen38-27b-autoround-int4-b70}/$LABEL
cache=${CACHE_ROOT:-/mnt/usb-models/llm-runtime/vllm-cache}/qwen38-mtp4-noscratch
qbase="$repo/data/qwen38-27b-autoround-int4-b70-baselines/quality-qwen38-int4-mtp3-fast-20260818.json"

VALIDATION_GDN_SPEC_PERSISTENT_SCRATCH=0 \
MODEL_DIR=/mnt/usb-models/llm-models/qwen3.8-27b-int4-autoround-devan \
VALIDATION_MODEL_MANIFEST="$repo/repro/qwen38-27b-autoround-int4-b70/manifests/model.json" \
VALIDATION_VLLM_CACHE_ROOT="$cache" \
VALIDATION_RUN_SMOKE=1 VALIDATION_RUN_BENCH=1 VALIDATION_RUN_QUALITY=1 \
VALIDATION_BENCH_MAX_TOKENS=512 VALIDATION_BENCH_METRIC_TOKENS=100 \
VALIDATION_NUM_SPECULATIVE_TOKENS=4 \
VALIDATION_COMPILATION_CONFIG_OVERRIDE='{"use_inductor_graph_partition":true,"pass_config":{"fuse_rope_kvcache_cat_mla":false},"cudagraph_mode":"PIECEWISE","cudagraph_capture_sizes":[5],"max_cudagraph_capture_size":5}' \
VALIDATION_ENABLE_XPU_GRAPH=1 \
VALIDATION_GDN_NATIVE_SPEC_RECURRENT_SERIAL_EXACT=0 \
VALIDATION_GDN_CAPTURE_NATIVE_SPEC=1 VALIDATION_GDN_NATIVE_SPEC_COMPLETION_BARRIER=0 \
VALIDATION_ONEDNN_INT4_COMPLETION_BARRIER=1 VALIDATION_ONEDNN_INT4_INPUT_DEPENDENCY=1 \
VALIDATION_ONEDNN_INT4_INPUT_DEPENDENCY_SCOPE=all_target \
VALIDATION_ONEDNN_INT8_COMPLETION_BARRIER=1 VALIDATION_ONEDNN_INT8_INPUT_DEPENDENCY=1 \
VALIDATION_LM_HEAD_INT8=1 VALIDATION_DETERMINISTIC_GREEDY_MARGIN=0.03125 \
VALIDATION_VLLM_EXTRA_ARGS='--dtype float16' \
LABEL=$LABEL \
"$repo/experiments/qwen36-27b-autoround-int4-b70/validation-20260815/run-arm.sh" \
spec-native-partition-exact-native 0,1 "$root" "$qbase"
```

Historically, all three arms shared one `VALIDATION_VLLM_CACHE_ROOT`. Fresh
compilations produce different-but-internally-deterministic code, so a
fresh-cache rerun will not reproduce token-for-token. The compile cache is part
of the run identity.

The published command's `VALIDATION_GDN_SPEC_PERSISTENT_SCRATCH=0` attribution
was also wrong: the old harness scrubbed that value and hard-exported `1`, as
proved by 96 scratch-allocation messages in each record arm. The harness now
propagates the validation variable and records the effective flag.
