Qwen3.8 Flash-Next · 125B-A6B

Qwen's experimental 125B-A6B hybrid-attention MoE. The official FP8 export serves on four B70s with selective host placement, and all 25 practical TP4 eager-text MTP/context cells through 8K are classified. The retained current runtime now independently qualifies TP4 eager MTP0 at short context and exact active 4K: 5.224 tok/s on the established short screen and 4.758 tok/s conventional at exact 4K, with 6/7 semantic quality, 16/16 repeats, an exact cache-zero 4K needle, and card-clean teardown. At 16K, current-source MTP0 completed one exact generic-depth request, then the semantic program produced one correct fresh response, a corrupted same-server repeat, and a separate fresh-server worker timeout at 1,600 computed prompt tokens. This is Grade-D quarantined capability and nondeterministic runtime-stability evidence; it does not authorize 24K/32K. MTP2's bounded 16K treatment tranche is exhausted. MTP3 remains the preferred exact-4K recipe at 15.502 tok/s decode. The target-only official-thinking MTP0 profile passed 25/25, and graph, vision, other topologies, clean-host replay, and deployment qualification remain open.

39.3tok/s
measured
≈ 29 words a secondqwen38-flash-next · qwen38-flash-next-fp8-bcd9f01 · official FP8 block-128 · vLLM XPU 6d872457 (fused-QSA overlay on the Triton-HC MTP1 head) + staged kernels 2f829747 · TP4 · MTP1 · auto KV · max context 4,352 · active context 4,096 · New authority whose difference is at depth; quality profile byte-identical to the certified battery; lossless MTP1 within the lineage; LocalMaxxing 37.83 tok/s approved
Qwen3.8 Flash-Next · base post-trained weights
1 exact artifacts · 1 quantizations
  • qwen38-flash-next-fp8-bcd9f01 · FP8 block-128 · Qwen/Qwen3.8-Flash-Next-FP8@bcd9f01ddc9cff2316eb84281bebcd5b058bddce
Qwen3.8 Flash-Next FP8 · TP4 MTP0, exact-4K, never-routed experts host-placed, both reference Triton kernels restored33.44tok/sqwen38-flash-next · qwen38-flash-next-fp8-bcd9f01 · official FP8 block-128 · vLLM XPU 2a372e86 (fused-QSA overlay on the Triton-HC MTP0 head) + staged kernels 2f829747 · TP4 · MTP0 · auto KV · max context 4,352 · active context 4,096New authority whose difference is at depth; quality profile byte-identical to the certified battery; LocalMaxxing 33.80 tok/s approvedQwen3.8 Flash-Next FP8 · TP4 lossless MTP1, exact-4K, never-routed experts host-placed, Triton HC glue36.37tok/sqwen38-flash-next · qwen38-flash-next-fp8-bcd9f01 · official FP8 block-128 · vLLM XPU 62219122 (Triton-HC overlay on the lossless MTP1 placement head 005dc578) + staged kernels 2f829747 · TP4 · MTP1 · auto KV · max context 4,352 · active context 4,096New deterministic authority reproduced on five servers; quality profile equal to the certified rows; lossless MTP1 within the lineage; LocalMaxxing 37.05 tok/s approvedQwen3.8 Flash-Next FP8 · TP4 MTP0, exact-4K, never-routed experts host-placed, Triton HC glue32.6tok/sqwen38-flash-next · qwen38-flash-next-fp8-bcd9f01 · official FP8 block-128 · vLLM XPU 8d7d6fd8 (Triton-HC overlay on the MTP0 placement head cb59004b) + staged kernels 2f829747 · TP4 · MTP0 · auto KV · max context 4,352 · active context 4,096New deterministic authority reproduced on three servers; quality profile equal to the certified rows; LocalMaxxing 32.90 tok/s approvedQwen3.8 Flash-Next FP8 · TP4 deterministic full-decode graph, exact-4K, never-routed experts host-placed (no speculation)27.4tok/sqwen38-flash-next · qwen38-flash-next-fp8-bcd9f01 · official FP8 block-128 · vLLM XPU cb59004b (placement overlay on 2169dbfe) + staged kernels 2f829747 · TP4 · MTP0 · auto KV · max context 4,352 · active context 4,096Certified line: outputs bit-identical to the promoted authorities on two servers; 6/7 semantic with the inherited miss, 16/16 repeat, exact needle; LocalMaxxing 27.64 tok/s approvedQwen3.8 Flash-Next FP8 · TP4 lossless MTP1, exact-4K, never-routed experts host-placed32.49tok/sqwen38-flash-next · qwen38-flash-next-fp8-bcd9f01 · official FP8 block-128 · vLLM XPU 005dc578 (placement overlay on the lossless MTP1 head 1b2a17c1) + staged kernels 2f829747 · TP4 · MTP1 · auto KV · max context 4,352 · active context 4,096Every output pin equal to the MTP0 line at exact 2K and 4K; certified battery; LocalMaxxing 31.93 tok/s approved
Read the lab reportAll packets and recipesNo step-by-step install guide is published for this model yet.
B70 fit
Four card screened
Quality evidence
Official target quality pass research deployment
Interest
2.22K
Measured results
13

Packets and recipes

The deployment variants of this family, at every maturity.

qwen38-flash-nextB70-verified originating-host replay

Qwen3.8 Flash-Next FP8 · TP4+EP4 graph + lossless MTP1 + never-routed experts host-placed

Reproduce 31.93 tok/s on 4× B70 · candidate package

Median of prompt-class medians over 99 inter-token intervals after TTFT on the fixed cold 12-prompt realistic suite, sent once (A226, 2026-09-06).
Capability B of A–DEvidence B of A–D
decodeexactnessqualityruntime provenancerecipe
qwen38-flash-nextB70-verified originating-host replay

Qwen3.8 Flash-Next FP8 · TP4+EP4 graph + lossless MTP1 + never-routed experts host-placed + Triton HC glue

Reproduce 37.05 tok/s on 4× B70 · candidate package

Median of prompt-class medians over 99 inter-token intervals after TTFT on the fixed cold 12-prompt realistic suite, sent once (A272, 2026-09-07).
Capability B of A–DEvidence B of A–D
decodeexactnessqualityruntime provenancerecipe
qwen38-flash-nextB70-verified originating-host replay

Qwen3.8 Flash-Next FP8 · TP4+EP4 graph + lossless MTP1 + never-routed experts host-placed + both reference Triton kernels

Reproduce 37.83 tok/s on 4× B70 · candidate package

Median of prompt-class medians over 99 inter-token intervals after TTFT on the fixed cold 12-prompt realistic suite, sent once (A306, 2026-09-07).
Capability B of A–DEvidence B of A–D
decodeexactnessqualityruntime provenancerecipe
qwen38-flash-nextB70-verified originating-host replay; clean-host replay pending

Qwen3.8 Flash-Next FP8 ? TP4+EP4 graph, no speculation, W13-N64

Reproduce 34.5 tok/s on 4× B70 · candidate package

median of prompt-class medians, 99 inter-token intervals after TTFT, fixed realistic suite run once cold; A326 fresh-server repeat, 2026-09-08.
Capability B of A–DEvidence B of A–D
decodeexactnessruntime provenancerecipe
qwen38-flash-nextB70-verified originating-host replay

Qwen3.8 Flash-Next FP8 · TP4+EP4 graph + lossless MTP1

Reproduce 27.05 tok/s on 4× B70 · candidate package

Median of prompt-class medians over 99 inter-token intervals after TTFT on the fixed cold 12-prompt realistic suite, sent once (A189, 2026-09-05).
Capability B of A–DEvidence B of A–D
decodeexactnessqualityruntime provenancerecipe
qwen38-flash-nextB70-screened

Qwen3.8 Flash-Next FP8 · TP4 research server

Measured evidence: 5.22 tok/s on 4× B70 · research-only-quality-caveat

One instrumentation-free server; p146/o256/c1; three sequential repetitive-prompt requests; full-output rate after first text, not the conventional 99-interval final gate lab report — documents the result; not a step-by-step install guide
Capability C of A–DEvidence C of A–D
TP4/EP4MTP0eagertext-onlyconfigured max 512/1536/3072/4352/8448short quality plus screened 1K/2K/4K and formal 8K context
qwen38-flash-nextB70-screened

Qwen3.8 Flash-Next FP8 · TP4 MTP0 current-runtime research quick check

Measured evidence: 4.76 tok/s on 4× B70 · research-only-quality-caveat

Two exact p4096/o128 requests; conventional decode across the 99 inter-token intervals between generated-token events 1 and 100 after TTFT lab report — documents the result; not a step-by-step install guide
Capability C of A–DEvidence C of A–D
TP4/EP4MTP0eagertext-onlyconfigured max 4352short screen plus repeated exact active-4K screensame-boot quality and card-clean teardown
qwen38-flash-nextB70-screened

Qwen3.8 Flash-Next FP8 · TP4 MTP1 research quick check

Measured evidence: 9.37 tok/s on 4× B70 · research-only-quality-caveat

One instrumentation-free server; p146/o256/c1; one untimed warmup then three sequential measured requests; full-output rate after first text, not the conventional 99-interval final gate lab report — documents the result; not a step-by-step install guide
Capability D of A–DEvidence C of A–D
TP4/EP4MTP1eagertext-onlyconfigured max 512 and 4352/headroom32configured max 1536 and 3072 retained as first-request bounded negativesconfigured max 8448 retained as an active-8K cross-runtime parity quarantinematched MTP0 quality plus 16/16 repeat and exact-4K needlethree exact-4K deployment-shaped rows plus a separate formal row
qwen38-flash-nextB70-screened

Qwen3.8 Flash-Next FP8 · TP4 MTP3 research quick check

Measured evidence: 15.5 tok/s on 4× B70 · research-only-quality-caveat

Three separately salted exact p4096/o256/c1 requests with no harness-added warmups; full-output rate after first text; formal p4096/o128 99-interval row retained separately lab report — documents the result; not a step-by-step install guide
Capability C of A–DEvidence C of A–D
TP4/EP4MTP3eagertext-onlyconfigured max 512 and 4352configured max 1536 retained as an active-1K external-signal quarantineconfigured max 3072 retained as an active-2K target-parity quarantineconfigured max 8448 retained as an active-8K bounded no-receipt quarantinematched MTP0 quality plus 16/16 repeat and exact-4K needlevariable 512-token screen plus exact-4K service and formal rows
qwen38-flash-nextB70-screened

Qwen3.8 Flash-Next FP8 · TP4 MTP2 research quick check

Measured evidence: 11.9 tok/s on 4× B70 · research-only-quality-caveat

One instrumentation-free server; p146/o256/c1; one untimed warmup per invocation then three sequential measured requests; full-output rate after first text, not the conventional 99-interval final gate lab report — documents the result; not a step-by-step install guide
Capability D of A–DEvidence C of A–D
TP4/EP4MTP2eagertext-onlyconfigured max 512 and 4352/headroom32configured max 1536 retained as an active-1K clean-host quarantineconfigured max 3072 and 8448 retained as active-2K and active-8K target-parity quarantinesmatched MTP0 quality plus 16/16 repeat and exact-4K needleshort and exact-4K three-row speed screens plus separate formal rows
qwen38-flash-nextB70-screened

Qwen3.8 Flash-Next FP8 · TP4 MTP4 research quick check

Measured evidence: 20.73 tok/s on 4× B70 · research-only-quality-caveat

One instrumentation-free server; p146/o256/c1; one untimed warmup per invocation then three sequential measured requests; full-output rate after first text, not the conventional 99-interval final gate lab report — documents the result; not a step-by-step install guide
Capability D of A–DEvidence C of A–D
TP4/EP4MTP4eagertext-onlyconfigured max 512configured max 1536 retained as an active-1K teardown quarantineconfigured max 3072 retained as an active-2K no-output and reset quarantineconfigured max 4352 retained as an exact-4K runtime quarantineconfigured max 8448 retained as an exact active-8K cross-runtime parity quarantinematched MTP0 quality plus 16/16 repeat and small needlethree-row speed screen

What has been classified

Dense scorecards summarize every declared combination; measured slices retain exact evidence links.

Practical TP4 eager text coverage

Official FP8 child artifact on TP4/EP4, eager text, and automatic KV. MTP0 uses vLLM 658965050 plus kernels 2f829747; MTP1-4 use vLLM 1372c62d plus staged kernels 2f829747. This is a coverage map across exact runtime identities, not a controlled depth curve. Fixed: revision=qwen38-flash-next · artifact_id=qwen38-flash-next-fp8-bcd9f01 · runtime_family=vLLM XPU · tp=4 · ep=4 · parallel_profile=TP=EP · graph_mode=off · kv=auto · modality=text.

  • ◇ Speed check only0, no speculative decoding 0·MTP05.22 tok/sreport · evidence
  • ◇ Speed check only1K, no speculative decoding 1K·MTP05.13 tok/sreport · evidence
  • ◇ Speed check only2K, no speculative decoding 2K·MTP05.23 tok/sreport · evidence
  • ◇ Speed check only4K, no speculative decoding 4K·MTP04.46 tok/sreport · evidence
  • ◇ Speed check only8K, no speculative decoding 8K·MTP03.98 tok/sreport · evidence
  • ◇ Speed check only0, + speculative decoding (depth 1) 0·MTP19.37 tok/sreport · evidence
  • ◇ Speed check only4K, + speculative decoding (depth 1) 4K·MTP18.9 tok/sreport · evidence
  • ◇ Speed check only0, + speculative decoding (depth 2) 0·MTP211.9 tok/sreport · evidence
  • ◇ Speed check only4K, + speculative decoding (depth 2) 4K·MTP29.89 tok/sreport · evidence
  • ◇ Speed check only0, + speculative decoding (depth 3) 0·MTP314.89 tok/sreport · evidence
  • ◇ Speed check only4K, + speculative decoding (depth 3) 4K·MTP315.5 tok/sreport · evidence
  • ◇ Speed check only0, + speculative decoding (depth 4) 0·MTP420.73 tok/sreport · evidence
  • ⚠ Quarantined1K, + speculative decoding (depth 1) 1K·MTP1 — The exact 32-block configured-1536 identity passed source, runtime, four-rank, placement, cache, capacity, and health gates, then reached the unchanged worker-response deadline during its first 1K request. No output was returned; request two was not sent and the separate 2K boot did not run. Observed: 768 computed prompt tokens; no output.evidence
  • ⚠ Quarantined2K, + speculative decoding (depth 1) 2K·MTP1 — The exact 32-block configured-3072 identity passed all startup gates and exposed 7,561 cache tokens. Its first exact-2K exchange had a zero-byte completion body and no output token recorded when the fixed 360-second client bound expired; the subsequent engine diagnostic showed 448 computed prompt tokens and zero output, while vLLM completed-request, token, and MTP counters remained zero. The engine independently reported its own sampling timeout. Request two was not sent. The post-failure teardown window recorded compute- and copy-class resets on all four cards before all four were rediscovered; no post-reset collective was run. Existing MTP1 configured-512 and exact-4K passes and every captured speed remain unchanged. Observed: 360-second client timeout; 448 computed tokens; no output or speed.evidence
  • ⚠ Quarantined8K, + speculative decoding (depth 1) 8K·MTP1 — The exact 32-block current-source identity exposed 13,516 cache tokens and completed one p8192/o128 request with exact usage, zero cache reuse, all 25 generic depth gates, and positive MTP1 counters at position zero. Its token array first diverged from the frozen cross-runtime/cache MTP0 authority at zero-based generated-token index 72. The 4.151 tok/s rate and 953.3-second TTFT receive no speed or quality credit. Controlled teardown passed its cleanup gates and returned all four cards idle; the intentional stop retained the known shutdown-time output-handler notice and one shared-memory cleanup warning. The host window contained corrected local-NVMe events but no B70 event. Observed: exact 8K completed; parity mismatch at token 73; 4.151 tok/s diagnostic only.evidence
  • ⚠ Quarantined1K, + speculative decoding (depth 2) 1K·MTP2 — Both exact active-1K requests returned the frozen MTP0 text hash with zero cache reuse, identical text, and perfect MTP2 acceptance at positions zero and one. Request one observed 10.683 tok/s after first text and the repeat sentinel 12.642 tok/s. Eleven corrected local-NVMe events after the preregistered journal cutoff failed the strict clean-host gate, so neither rate receives speed, quality, or deployment credit. No event named a B70 address; existing MTP2 configured-512, active-2K, and exact-4K results remain unchanged. Observed: exact parity twice; corrected local-NVMe events; diagnostic only.evidence
  • ⚠ Quarantined2K, + speculative decoding (depth 2) 2K·MTP2 — The exact 32-block configured-3072 identity passed all startup gates. Its first exact-2K request returned 128 tokens with zero cache reuse and active MTP2 counters, but the token array diverged from the frozen MTP0 authority at zero-based generated-token index 12. Request two was not sent. The observed 4.527 tok/s is diagnostic only and does not alter the passing MTP2 512 or exact-4K results. Observed: completed; target-parity mismatch at token 13; no speed credit.evidence
  • ⚠ Quarantined8K, + speculative decoding (depth 2) 8K·MTP2 — The exact 32-block current-source identity exposed 11,264 cache tokens and completed one p8192/o128 request with exact usage, zero cache reuse, all generic depth gates, and positive MTP2 counters at both positions. Its token array first diverged from the frozen cross-runtime/cache MTP0 authority at zero-based generated-token index 26. The 6.235 tok/s rate and 649.7-second TTFT receive no speed or quality credit. The bounded host window contained corrected storage/root-port events but no B70 event; teardown was clean and all four cards returned idle. Observed: exact 8K completed; parity mismatch at token 27; 6.235 tok/s diagnostic only.evidence
  • ⚠ Quarantined1K, + speculative decoding (depth 3) 1K·MTP3 — The exact 25-block local-NVMe identity passed source, runtime, fresh four-rank, placement, cache, capacity, served-identity, and health gates. Request one began under the frozen protocol, then the server received an external SIGTERM before the response completed. No request JSON, usage, output hash, or performance result exists. Partial server metrics showed six drafted and six accepted tokens with 1.000 acceptance at all three MTP3 positions, but those counters receive no parity, speed, quality, or deployment credit. Request two was not sent. Existing MTP3 configured-512, active-2K, and exact-4K results remain unchanged. Observed: external SIGTERM during request one; six accepted draft tokens; no completed response or speed.evidence
  • ⚠ Quarantined2K, + speculative decoding (depth 3) 2K·MTP3 — The exact 25-block configured-3072 identity passed source, runtime, four-rank, placement, cache, capacity, and health gates. Its first exact-2K request returned 128 tokens with zero cache reuse and active MTP3 counters, but the token array diverged from the frozen MTP0 authority at zero-based generated-token index 4. Request two was not sent. The observed 5.932 tok/s is diagnostic only and does not alter the passing MTP3 512 or exact-4K results. Observed: completed; target-parity mismatch at token five; no speed credit.evidence
  • ⚠ Quarantined8K, + speculative decoding (depth 3) 8K·MTP3 — The exact current-source TP4/EP4 eager MTP3 boot passed source/runtime, fresh four-rank, placement, served-identity, health, and 32-block capacity gates, reporting 9,654 cache tokens. Its sole p8192/o128 request reached the fixed 900-second client bound without a completed response receipt or any durably recorded output token. No usage, cache-zero, MTP-counter, parity, quality, TTFT, or speed result is claimed. The failed-request supervisor path cleanly removed the listener, process group, compile/RPC paths and returned all four cards idle. The host window contained corrected local-NVMe events but no B70-addressed event. Existing MTP3 configured-512, active-2K, exact-4K, and every captured speed remain unchanged. Observed: 900-second client timeout; no completed receipt or speed.evidence
  • ⚠ Quarantined1K, + speculative decoding (depth 4) 1K·MTP4 — Both exact active-1K requests returned the frozen MTP0 text hash with zero cache reuse, identical text, and perfect MTP4 acceptance at positions zero through three. Request one observed 13.326 tok/s after first text and the repeat sentinel 17.291 tok/s. The exact stop sentinel terminated the timeout supervisor without reaching the detached server group; direct recovery produced an orderly shutdown, but the frozen teardown rule quarantines the cell. Seven corrected-only local-NVMe records separately block clean-host qualification. Existing MTP4 configured-512 and exact-4K results remain unchanged. Observed: exact parity twice; detached teardown gate failed; diagnostic only.evidence
  • ⚠ Quarantined2K, + speculative decoding (depth 4) 2K·MTP4 — The exact 29-block configured-3072 identity passed source, runtime, fresh four-rank, placement, cache, capacity, served-identity, and health gates. Request one reached the fixed 360-second client bound without a response receipt; about five seconds later the engine independently reported its own sampling timeout at 384 computed prompt tokens and zero output. No HTTP status, usage, parity hash, MTP counter delta, or speed exists, and request two was not sent. The corrected supervisor returned zero and left no listener, recorded process, compile path, or RPC path, but the teardown window recorded compute- and copy-class resets on all four B70s. All four cards were rediscovered at low memory use; no post-reset collective or known-good generation canary was run. Existing MTP4 configured-512, active-1K, exact-4K, and all captured speeds remain unchanged. Observed: 360-second client timeout; 384 computed tokens; no output; four-card teardown resets.evidence
  • ⚠ Quarantined4K, + speculative decoding (depth 4) 4K·MTP4 — The exact 29-block configuration admitted 4,352 tokens but stopped during the 4K quality request; no durable quality or timing result was produced. Observed: worker timeout at 3,904 computed tokens; no speed.evidence
  • ⚠ Quarantined8K, + speculative decoding (depth 4) 8K·MTP4 — The exact 36-block current-source identity exposed 9,504 cache tokens and completed one p8192/o128 request with exact usage, zero cache reuse, all 25 generic depth gates, and positive MTP4 counters at all four positions. Its token array first diverged from the frozen cross-runtime/cache MTP0 authority at zero-based generated-token index 26. The 4.026 tok/s rate and 918.4-second TTFT receive no speed or quality credit. Controlled teardown passed its cleanup gates and returned all four cards idle; the five-second grace expired before the remaining EngineCore and workers were stopped, and the intentional stop retained the known output-handler and shared-memory cleanup notices. Six corrected local-NVMe endpoint records block clean-host wording, but no B70 event appears. The superseded attempt-1 command-identity stop sent no request and grants no matrix credit. Observed: exact 8K completed; parity mismatch at token 27; 4.026 tok/s diagnostic only.evidence

Archived MTP0 context anchors and estimates

Legacy vLLM 658965050 TP4/EP4 eager-text MTP0 only. The 4K/8K cells are formal exact-depth measurements; 24K/32K are deterministic Grade-D extrapolations with 50%-150% bands. They do not transfer to the latest runtime or qualify boot, fit, quality, deployment, records, or promotion. Fixed: revision=qwen38-flash-next · artifact_id=qwen38-flash-next-fp8-bcd9f01 · runtime=vLLM XPU 658965050 + kernels 2f829747 · runtime_family=vLLM XPU · tp=4 · ep=4 · parallel_profile=TP=EP · graph_mode=off · kv=auto · modality=text.

  • ◇ Speed check only4K, no speculative decoding 4K·MTP04.46 tok/sreport · evidence
  • ◇ Speed check only8K, no speculative decoding 8K·MTP03.98 tok/sreport · evidence
  • ≈ Estimate24K, no speculative decoding 24K·MTP0 — qwen38-flash-next-mtp0-context-estimator 1.0.0≈ 3.33 tok/s (1.66–4.99)evidence
  • ≈ Estimate32K, no speculative decoding 32K·MTP0 — qwen38-flash-next-mtp0-context-estimator 1.0.0≈ 3.17 tok/s (1.59–4.76)evidence

TP4 eager text deeper-context coverage

Official FP8 child artifact on current-source vLLM 1372c62d plus staged kernels 2f829747, TP4/EP4 eager text, and automatic KV. Single generic requests and bounded negatives are Grade-D matrix evidence until semantic and repeat qualification exists. Fixed: revision=qwen38-flash-next · artifact_id=qwen38-flash-next-fp8-bcd9f01 · runtime=vLLM XPU 1372c62d + staged kernels 2f829747 · runtime_family=vLLM XPU · tp=4 · ep=4 · parallel_profile=TP=EP · graph_mode=off · kv=auto · modality=text.

  • ⚠ Quarantined16K, no speculative decoding 16K·MTP0 — The current-source 33-block identity first completed one generic p16384/o128 request at a diagnostic-only 5.219 tok/s. The semantic program then returned one correct 16,213-token fresh response, corrupted the identical same-server repeat into repeated text despite zero reported cache use, and stopped a separate fresh-server attempt at 1,600 computed prompt tokens with no first output after an RPC sampling timeout. Cleanup of the latter recorded eight B70 engine resets and 61 unsuccessful responses before all four cards re-enumerated idle. The cell is nondeterministically unstable and receives no speed, curve, quality, deployment, or headline credit; unchanged retries and 24K/32K serving remain blocked pending a material runtime treatment. Observed: 16K sometimes completes; same-server corruption and fresh-server timeout; diagnostic only.evidence
  • ⚠ Quarantined16K, + speculative decoding (depth 2) 16K·MTP2 — The current-source MTP2 identity admitted exactly 40 blocks / 20,014 cache tokens in both bounded arms. Scheduler-64 stopped at 3,200 computed prompt tokens; the scheduler-32 arm was observed 2,240 tokens / 70% farther at 5,440, but the sole p16384/o128 treatment request still returned no output before the unchanged 300-second runtime response deadline. This is treatment evidence, not repeat-confirmed causality. Its shutdown was followed by eight card reset records and 58 unsuccessful card responses, so the frozen postflight rule fails and the treatment tranche is exhausted. No capability, speed, quality, parity, deployment, or headline credit is granted; current checks found no listener or owned residue and all cards are below 43 MiB. Observed: exact 16K stopped at 5,440 computed prompt tokens after scheduler-32 treatment; no output; four-card teardown events; diagnostic only.evidence
13 untested combinations
24K·MTP032K·MTP016K·MTP124K·MTP132K·MTP124K·MTP232K·MTP216K·MTP324K·MTP332K·MTP316K·MTP424K·MTP432K·MTP4
Full matrix and exact selectors

Card-fit summary

Official FP8 child artifact, vLLM 658965050 plus kernels 2f829747, eager MTP0 text, zero prior context, automatic KV, and TP=EP. TP1 and TP2 need separate fit/offload designs; they are not claimed unsupported. Fixed: revision=qwen38-flash-next · artifact_id=qwen38-flash-next-fp8-bcd9f01 · runtime=vLLM XPU 658965050 + kernels 2f829747 · runtime_family=vLLM XPU · parallel_profile=TP=EP · active_context_tokens=0 · graph_mode=off · kv=auto · modality=text.

  • ◇ Speed check only4 cards, no speculative decoding TP4·MTP05.22 tok/sreport · evidence
2 untested combinations
TP1·MTP0TP2·MTP0
Full matrix and exact selectors

Graph and modality summary

Official FP8 child artifact, vLLM 1372c62d plus staged kernels 2f829747, TP4/EP4, MTP0, zero prior context, and automatic KV. Eager text is screened; PIECEWISE text has bounded compile-resource evidence but no API, quality, or speed result; vision remains unmeasured. Fixed: revision=qwen38-flash-next · artifact_id=qwen38-flash-next-fp8-bcd9f01 · runtime=vLLM XPU 1372c62d + staged kernels 2f829747 · runtime_family=vLLM XPU · tp=4 · ep=4 · parallel_profile=TP=EP · mtp=0 · active_context_tokens=0 · kv=auto.

  • ◇ Speed check onlytext, eager text·eager5.22 tok/sreport · evidence
  • ⚠ Quarantinedtext, graph text·graph — Attempt 7 loaded the complete checkpoint on all four ranks and entered PIECEWISE torch.compile with one compile thread. The phase-aware guard latched below its 30-GiB host-memory floor before the later TTM/global-OOM teardown window. Graph capture did not complete, the API never became healthy, and no client, output, quality, replay, or speed row exists. This is Grade-D bounded resource evidence only. Observed: all 131 shards loaded on four ranks; graph compile began; 30-GiB host-memory guard tripped; no API, quality, or speed.evidence
2 untested combinations
vision·eagervision·graph
Full matrix and exact selectors
Full 270-cell coverage contracts · 26 classified
Coverage · 2 matrices26/270 classified
16 quarantined244 missing9 screened1 closed

Text serving coverage

Exact text-serving combinations across topology, MTP depth, active context, and graph mode on the latest retained runtime. Older-runtime evidence is preserved in a separate archival contract and never silently transferred into these cells.

26/240
240 exact cells26 classified0 measured214 gaps240 retry-tagged
1 ■ closed9 ◇ screened214 · untested16 ⚠ quarantined

Fixed: runtime=vLLM XPU 1372c62d + staged kernels 2f829747 · runtime_family=vLLM XPU · parallel_profile=TP=EP

Break down by axis
Weight revisionqwen38-flash-next 26/240
Artifactqwen38-flash-next-fp8-bcd9f01 26/240
Tensor parallel1 1/802 0/804 25/80
MTP depth0 5/481 5/482 6/483 5/484 5/48
Active context0 7/301024 4/302048 4/304096 5/308192 4/3016384 2/3024576 0/3032768 0/30
Graph modeoff 25/120PIECEWISE 1/120
KV cacheauto 26/240
Modalitytext 26/240

Fixed-vision serving coverage

Vision capability and quality combinations under one versioned fixed-image fixture. Context depth is intentionally excluded until a vision anchor establishes meaningful image-plus-text token accounting.

0/30
30 exact cells0 classified0 measured30 gaps30 retry-tagged
30 · untested

Fixed: runtime=vLLM XPU 1372c62d + staged kernels 2f829747 · runtime_family=vLLM XPU · parallel_profile=TP=EP · active_context_tokens=0 · modality=vision · workload_profile=fixed-vision-fixture-v1-required

Break down by axis
Weight revisionqwen38-flash-next 0/30
Artifactqwen38-flash-next-fp8-bcd9f01 0/30
Tensor parallel1 0/102 0/104 0/10
MTP depth0 0/61 0/62 0/63 0/64 0/6
Graph modeoff 0/15PIECEWISE 0/15
KV cacheauto 0/30

Measured results

Every number links to its proof.

family data

Research speed screen by TP ◇ screened, experimental

Only TP4 has bounded screens; each MTP cell has three same-server samples under the declared workload, so no TP mini-curve is implied

05.811.617.423.2124tensor parallel cards · tok/sofficial FP8 · eager MTP0 · Decode (decode_tok_s) · tensor parallel cards=4.0 · value=5.221849709057954 tok/sofficial FP8 · eager MTP1 · Decode (decode_tok_s) · tensor parallel cards=4.0 · value=9.37225436776222 tok/sofficial FP8 · eager MTP2 (variable screen) · Decode (decode_tok_s) · tensor parallel cards=4.0 · value=11.895061402541456 tok/sofficial FP8 · eager MTP3 (variable screen) · Decode (decode_tok_s) · tensor parallel cards=4.0 · value=14.88878979448863 tok/sofficial FP8 · eager MTP4 · Decode (decode_tok_s) · tensor parallel cards=4.0 · value=20.72717637199404 tok/sgapgap
official FP8 · eager MTP0official FP8 · eager MTP1official FP8 · eager MTP2 (variable screen)official FP8 · eager MTP3 (variable screen)official FP8 · eager MTP4missing
official FP8 · eager MTP0 4: 5.22
official FP8 · eager MTP1 4: 9.37
official FP8 · eager MTP2 (variable screen) 4: 11.9
official FP8 · eager MTP3 (variable screen) 4: 14.89
official FP8 · eager MTP4 4: 20.73
evidence 1 · evidence 2 · evidence 3 · evidence 4 · evidence 5

Legacy-comparable after-first-text context screen ◇ screened, experimental

The plotted 5.22/5.13/5.23/5.23 tok/s values use the legacy after-first-text comparison and stop at 4K. Separate formal cache-zero rates are 3.865 at 2K, 4.456 at 4K, and 3.980 tok/s at 8K; 8K is not plotted here because its third comparison row did not complete. All cells retain the short-quality caveat.

01.52.94.45.901K2K4Kactive context tokens · tok/sofficial FP8 · TP4 eager MTP0 · Decode (decode_tok_s) · active context tokens=0.0 · value=5.221849709057954 tok/sofficial FP8 · TP4 eager MTP0 · Decode (decode_tok_s) · active context tokens=1024.0 · value=5.13358756138473 tok/sofficial FP8 · TP4 eager MTP0 · Decode (decode_tok_s) · active context tokens=2048.0 · value=5.228429046201661 tok/sofficial FP8 · TP4 eager MTP0 · Decode (decode_tok_s) · active context tokens=4096.0 · value=5.233664731906276 tok/s
official FP8 · TP4 eager MTP0
official FP8 · TP4 eager MTP0 0: 5.22 · 1K: 5.13 · 2K: 5.23 · 4K: 5.23
evidence 1

Exact-4K service tradeoff by MTP depth ◇ screened, experimental

Workload-aligned p4096/o256/c1 service screens with no harness-added warmups, but from different vLLM source revisions and cache allocations. MTP1 and MTP2 use 32-block headroom recipes; MTP3 remains faster with lower TTFT and higher wall output. Separate single-point series avoid implying a causal MTP-depth or cache curve; only MTP4 remains quarantined.

04.38.71317.401234MTP depth · tok/sofficial FP8 · TP4 eager MTP0 · Decode (decode_tok_s) · MTP depth=0.0 · value=5.233664731906276 tok/sofficial FP8 · TP4 eager MTP1 · headroom32 · Decode (decode_tok_s) · MTP depth=1.0 · value=8.904420575355882 tok/sofficial FP8 · TP4 eager MTP2 · headroom32 · Decode (decode_tok_s) · MTP depth=2.0 · value=9.89315479235244 tok/sofficial FP8 · TP4 eager MTP3 · Decode (decode_tok_s) · MTP depth=3.0 · value=15.50156510641242 tok/sgap
official FP8 · TP4 eager MTP0official FP8 · TP4 eager MTP1 · headroom32official FP8 · TP4 eager MTP2 · headroom32official FP8 · TP4 eager MTP3missing
official FP8 · TP4 eager MTP0 0: 5.23
official FP8 · TP4 eager MTP1 · headroom32 1: 8.9
official FP8 · TP4 eager MTP2 · headroom32 2: 9.89
official FP8 · TP4 eager MTP3 3: 15.5
evidence 1 · evidence 2 · evidence 3 · evidence 4
Fine print

Qwen4Exp hybrid Gated DeltaNet/QSA sparse MoE, 48 layers, 2560 hidden. The FP8 repository is a quantized child of Qwen3.8 Flash-Next, not a separate model. Measurements remain artifact- and runtime-specific. Measurements, artifact hashes, outputs, quality decisions, and speed stay pinned to their exact recorded identity.

Many people at once Not measured

Multi-user (aggregate) throughput has not been measured for this family. The lab’s one measured sweep so far is in the multi-user report; the projection block below includes a projected users curve where a projection exists.

How much faster could this get? No projection

No projection is shown: this model is not in the ML Bottleneck catalog yet. The measured numbers above stand on their own.