Packets and recipes
The deployment variants of this family, at every maturity.
Qwen3.8 Flash-Next FP8 · TP4+EP4 graph + lossless MTP1 + never-routed experts host-placed
Reproduce 31.93 tok/s on 4× B70 · candidate package
Median of prompt-class medians over 99 inter-token intervals after TTFT on the fixed cold 12-prompt realistic suite, sent once (A226, 2026-09-06).Qwen3.8 Flash-Next FP8 · TP4+EP4 graph + lossless MTP1 + never-routed experts host-placed + Triton HC glue
Reproduce 37.05 tok/s on 4× B70 · candidate package
Median of prompt-class medians over 99 inter-token intervals after TTFT on the fixed cold 12-prompt realistic suite, sent once (A272, 2026-09-07).Qwen3.8 Flash-Next FP8 · TP4+EP4 graph + lossless MTP1 + never-routed experts host-placed + both reference Triton kernels
Reproduce 37.83 tok/s on 4× B70 · candidate package
Median of prompt-class medians over 99 inter-token intervals after TTFT on the fixed cold 12-prompt realistic suite, sent once (A306, 2026-09-07).Qwen3.8 Flash-Next FP8 ? TP4+EP4 graph, no speculation, W13-N64
Reproduce 34.5 tok/s on 4× B70 · candidate package
median of prompt-class medians, 99 inter-token intervals after TTFT, fixed realistic suite run once cold; A326 fresh-server repeat, 2026-09-08.Qwen3.8 Flash-Next FP8 · TP4+EP4 graph + lossless MTP1
Reproduce 27.05 tok/s on 4× B70 · candidate package
Median of prompt-class medians over 99 inter-token intervals after TTFT on the fixed cold 12-prompt realistic suite, sent once (A189, 2026-09-05).Qwen3.8 Flash-Next FP8 · TP4 research server
Measured evidence: 5.22 tok/s on 4× B70 · research-only-quality-caveat
One instrumentation-free server; p146/o256/c1; three sequential repetitive-prompt requests; full-output rate after first text, not the conventional 99-interval final gate lab report — documents the result; not a step-by-step install guideQwen3.8 Flash-Next FP8 · TP4 MTP0 current-runtime research quick check
Measured evidence: 4.76 tok/s on 4× B70 · research-only-quality-caveat
Two exact p4096/o128 requests; conventional decode across the 99 inter-token intervals between generated-token events 1 and 100 after TTFT lab report — documents the result; not a step-by-step install guideQwen3.8 Flash-Next FP8 · TP4 MTP1 research quick check
Measured evidence: 9.37 tok/s on 4× B70 · research-only-quality-caveat
One instrumentation-free server; p146/o256/c1; one untimed warmup then three sequential measured requests; full-output rate after first text, not the conventional 99-interval final gate lab report — documents the result; not a step-by-step install guideQwen3.8 Flash-Next FP8 · TP4 MTP3 research quick check
Measured evidence: 15.5 tok/s on 4× B70 · research-only-quality-caveat
Three separately salted exact p4096/o256/c1 requests with no harness-added warmups; full-output rate after first text; formal p4096/o128 99-interval row retained separately lab report — documents the result; not a step-by-step install guideQwen3.8 Flash-Next FP8 · TP4 MTP2 research quick check
Measured evidence: 11.9 tok/s on 4× B70 · research-only-quality-caveat
One instrumentation-free server; p146/o256/c1; one untimed warmup per invocation then three sequential measured requests; full-output rate after first text, not the conventional 99-interval final gate lab report — documents the result; not a step-by-step install guideQwen3.8 Flash-Next FP8 · TP4 MTP4 research quick check
Measured evidence: 20.73 tok/s on 4× B70 · research-only-quality-caveat
One instrumentation-free server; p146/o256/c1; one untimed warmup per invocation then three sequential measured requests; full-output rate after first text, not the conventional 99-interval final gate lab report — documents the result; not a step-by-step install guideWhat has been classified
Dense scorecards summarize every declared combination; measured slices retain exact evidence links.
Practical TP4 eager text coverage
Official FP8 child artifact on TP4/EP4, eager text, and automatic KV. MTP0 uses vLLM 658965050 plus kernels 2f829747; MTP1-4 use vLLM 1372c62d plus staged kernels 2f829747. This is a coverage map across exact runtime identities, not a controlled depth curve. Fixed: revision=qwen38-flash-next · artifact_id=qwen38-flash-next-fp8-bcd9f01 · runtime_family=vLLM XPU · tp=4 · ep=4 · parallel_profile=TP=EP · graph_mode=off · kv=auto · modality=text.
- ◇ Speed check only0, no speculative decoding
0·MTP05.22 tok/sreport · evidence - ◇ Speed check only1K, no speculative decoding
1K·MTP05.13 tok/sreport · evidence - ◇ Speed check only2K, no speculative decoding
2K·MTP05.23 tok/sreport · evidence - ◇ Speed check only4K, no speculative decoding
4K·MTP04.46 tok/sreport · evidence - ◇ Speed check only8K, no speculative decoding
8K·MTP03.98 tok/sreport · evidence - ◇ Speed check only0, + speculative decoding (depth 1)
0·MTP19.37 tok/sreport · evidence - ◇ Speed check only4K, + speculative decoding (depth 1)
4K·MTP18.9 tok/sreport · evidence - ◇ Speed check only0, + speculative decoding (depth 2)
0·MTP211.9 tok/sreport · evidence - ◇ Speed check only4K, + speculative decoding (depth 2)
4K·MTP29.89 tok/sreport · evidence - ◇ Speed check only0, + speculative decoding (depth 3)
0·MTP314.89 tok/sreport · evidence - ◇ Speed check only4K, + speculative decoding (depth 3)
4K·MTP315.5 tok/sreport · evidence - ◇ Speed check only0, + speculative decoding (depth 4)
0·MTP420.73 tok/sreport · evidence - ⚠ Quarantined1K, + speculative decoding (depth 1)
1K·MTP1— The exact 32-block configured-1536 identity passed source, runtime, four-rank, placement, cache, capacity, and health gates, then reached the unchanged worker-response deadline during its first 1K request. No output was returned; request two was not sent and the separate 2K boot did not run. Observed: 768 computed prompt tokens; no output.evidence - ⚠ Quarantined2K, + speculative decoding (depth 1)
2K·MTP1— The exact 32-block configured-3072 identity passed all startup gates and exposed 7,561 cache tokens. Its first exact-2K exchange had a zero-byte completion body and no output token recorded when the fixed 360-second client bound expired; the subsequent engine diagnostic showed 448 computed prompt tokens and zero output, while vLLM completed-request, token, and MTP counters remained zero. The engine independently reported its own sampling timeout. Request two was not sent. The post-failure teardown window recorded compute- and copy-class resets on all four cards before all four were rediscovered; no post-reset collective was run. Existing MTP1 configured-512 and exact-4K passes and every captured speed remain unchanged. Observed: 360-second client timeout; 448 computed tokens; no output or speed.evidence - ⚠ Quarantined8K, + speculative decoding (depth 1)
8K·MTP1— The exact 32-block current-source identity exposed 13,516 cache tokens and completed one p8192/o128 request with exact usage, zero cache reuse, all 25 generic depth gates, and positive MTP1 counters at position zero. Its token array first diverged from the frozen cross-runtime/cache MTP0 authority at zero-based generated-token index 72. The 4.151 tok/s rate and 953.3-second TTFT receive no speed or quality credit. Controlled teardown passed its cleanup gates and returned all four cards idle; the intentional stop retained the known shutdown-time output-handler notice and one shared-memory cleanup warning. The host window contained corrected local-NVMe events but no B70 event. Observed: exact 8K completed; parity mismatch at token 73; 4.151 tok/s diagnostic only.evidence - ⚠ Quarantined1K, + speculative decoding (depth 2)
1K·MTP2— Both exact active-1K requests returned the frozen MTP0 text hash with zero cache reuse, identical text, and perfect MTP2 acceptance at positions zero and one. Request one observed 10.683 tok/s after first text and the repeat sentinel 12.642 tok/s. Eleven corrected local-NVMe events after the preregistered journal cutoff failed the strict clean-host gate, so neither rate receives speed, quality, or deployment credit. No event named a B70 address; existing MTP2 configured-512, active-2K, and exact-4K results remain unchanged. Observed: exact parity twice; corrected local-NVMe events; diagnostic only.evidence - ⚠ Quarantined2K, + speculative decoding (depth 2)
2K·MTP2— The exact 32-block configured-3072 identity passed all startup gates. Its first exact-2K request returned 128 tokens with zero cache reuse and active MTP2 counters, but the token array diverged from the frozen MTP0 authority at zero-based generated-token index 12. Request two was not sent. The observed 4.527 tok/s is diagnostic only and does not alter the passing MTP2 512 or exact-4K results. Observed: completed; target-parity mismatch at token 13; no speed credit.evidence - ⚠ Quarantined8K, + speculative decoding (depth 2)
8K·MTP2— The exact 32-block current-source identity exposed 11,264 cache tokens and completed one p8192/o128 request with exact usage, zero cache reuse, all generic depth gates, and positive MTP2 counters at both positions. Its token array first diverged from the frozen cross-runtime/cache MTP0 authority at zero-based generated-token index 26. The 6.235 tok/s rate and 649.7-second TTFT receive no speed or quality credit. The bounded host window contained corrected storage/root-port events but no B70 event; teardown was clean and all four cards returned idle. Observed: exact 8K completed; parity mismatch at token 27; 6.235 tok/s diagnostic only.evidence - ⚠ Quarantined1K, + speculative decoding (depth 3)
1K·MTP3— The exact 25-block local-NVMe identity passed source, runtime, fresh four-rank, placement, cache, capacity, served-identity, and health gates. Request one began under the frozen protocol, then the server received an external SIGTERM before the response completed. No request JSON, usage, output hash, or performance result exists. Partial server metrics showed six drafted and six accepted tokens with 1.000 acceptance at all three MTP3 positions, but those counters receive no parity, speed, quality, or deployment credit. Request two was not sent. Existing MTP3 configured-512, active-2K, and exact-4K results remain unchanged. Observed: external SIGTERM during request one; six accepted draft tokens; no completed response or speed.evidence - ⚠ Quarantined2K, + speculative decoding (depth 3)
2K·MTP3— The exact 25-block configured-3072 identity passed source, runtime, four-rank, placement, cache, capacity, and health gates. Its first exact-2K request returned 128 tokens with zero cache reuse and active MTP3 counters, but the token array diverged from the frozen MTP0 authority at zero-based generated-token index 4. Request two was not sent. The observed 5.932 tok/s is diagnostic only and does not alter the passing MTP3 512 or exact-4K results. Observed: completed; target-parity mismatch at token five; no speed credit.evidence - ⚠ Quarantined8K, + speculative decoding (depth 3)
8K·MTP3— The exact current-source TP4/EP4 eager MTP3 boot passed source/runtime, fresh four-rank, placement, served-identity, health, and 32-block capacity gates, reporting 9,654 cache tokens. Its sole p8192/o128 request reached the fixed 900-second client bound without a completed response receipt or any durably recorded output token. No usage, cache-zero, MTP-counter, parity, quality, TTFT, or speed result is claimed. The failed-request supervisor path cleanly removed the listener, process group, compile/RPC paths and returned all four cards idle. The host window contained corrected local-NVMe events but no B70-addressed event. Existing MTP3 configured-512, active-2K, exact-4K, and every captured speed remain unchanged. Observed: 900-second client timeout; no completed receipt or speed.evidence - ⚠ Quarantined1K, + speculative decoding (depth 4)
1K·MTP4— Both exact active-1K requests returned the frozen MTP0 text hash with zero cache reuse, identical text, and perfect MTP4 acceptance at positions zero through three. Request one observed 13.326 tok/s after first text and the repeat sentinel 17.291 tok/s. The exact stop sentinel terminated the timeout supervisor without reaching the detached server group; direct recovery produced an orderly shutdown, but the frozen teardown rule quarantines the cell. Seven corrected-only local-NVMe records separately block clean-host qualification. Existing MTP4 configured-512 and exact-4K results remain unchanged. Observed: exact parity twice; detached teardown gate failed; diagnostic only.evidence - ⚠ Quarantined2K, + speculative decoding (depth 4)
2K·MTP4— The exact 29-block configured-3072 identity passed source, runtime, fresh four-rank, placement, cache, capacity, served-identity, and health gates. Request one reached the fixed 360-second client bound without a response receipt; about five seconds later the engine independently reported its own sampling timeout at 384 computed prompt tokens and zero output. No HTTP status, usage, parity hash, MTP counter delta, or speed exists, and request two was not sent. The corrected supervisor returned zero and left no listener, recorded process, compile path, or RPC path, but the teardown window recorded compute- and copy-class resets on all four B70s. All four cards were rediscovered at low memory use; no post-reset collective or known-good generation canary was run. Existing MTP4 configured-512, active-1K, exact-4K, and all captured speeds remain unchanged. Observed: 360-second client timeout; 384 computed tokens; no output; four-card teardown resets.evidence - ⚠ Quarantined4K, + speculative decoding (depth 4)
4K·MTP4— The exact 29-block configuration admitted 4,352 tokens but stopped during the 4K quality request; no durable quality or timing result was produced. Observed: worker timeout at 3,904 computed tokens; no speed.evidence - ⚠ Quarantined8K, + speculative decoding (depth 4)
8K·MTP4— The exact 36-block current-source identity exposed 9,504 cache tokens and completed one p8192/o128 request with exact usage, zero cache reuse, all 25 generic depth gates, and positive MTP4 counters at all four positions. Its token array first diverged from the frozen cross-runtime/cache MTP0 authority at zero-based generated-token index 26. The 4.026 tok/s rate and 918.4-second TTFT receive no speed or quality credit. Controlled teardown passed its cleanup gates and returned all four cards idle; the five-second grace expired before the remaining EngineCore and workers were stopped, and the intentional stop retained the known output-handler and shared-memory cleanup notices. Six corrected local-NVMe endpoint records block clean-host wording, but no B70 event appears. The superseded attempt-1 command-identity stop sent no request and grants no matrix credit. Observed: exact 8K completed; parity mismatch at token 27; 4.026 tok/s diagnostic only.evidence
Archived MTP0 context anchors and estimates
Legacy vLLM 658965050 TP4/EP4 eager-text MTP0 only. The 4K/8K cells are formal exact-depth measurements; 24K/32K are deterministic Grade-D extrapolations with 50%-150% bands. They do not transfer to the latest runtime or qualify boot, fit, quality, deployment, records, or promotion. Fixed: revision=qwen38-flash-next · artifact_id=qwen38-flash-next-fp8-bcd9f01 · runtime=vLLM XPU 658965050 + kernels 2f829747 · runtime_family=vLLM XPU · tp=4 · ep=4 · parallel_profile=TP=EP · graph_mode=off · kv=auto · modality=text.
- ◇ Speed check only4K, no speculative decoding
4K·MTP04.46 tok/sreport · evidence - ◇ Speed check only8K, no speculative decoding
8K·MTP03.98 tok/sreport · evidence - ≈ Estimate24K, no speculative decoding
24K·MTP0— qwen38-flash-next-mtp0-context-estimator 1.0.0≈ 3.33 tok/s (1.66–4.99)evidence - ≈ Estimate32K, no speculative decoding
32K·MTP0— qwen38-flash-next-mtp0-context-estimator 1.0.0≈ 3.17 tok/s (1.59–4.76)evidence
TP4 eager text deeper-context coverage
Official FP8 child artifact on current-source vLLM 1372c62d plus staged kernels 2f829747, TP4/EP4 eager text, and automatic KV. Single generic requests and bounded negatives are Grade-D matrix evidence until semantic and repeat qualification exists. Fixed: revision=qwen38-flash-next · artifact_id=qwen38-flash-next-fp8-bcd9f01 · runtime=vLLM XPU 1372c62d + staged kernels 2f829747 · runtime_family=vLLM XPU · tp=4 · ep=4 · parallel_profile=TP=EP · graph_mode=off · kv=auto · modality=text.
- ⚠ Quarantined16K, no speculative decoding
16K·MTP0— The current-source 33-block identity first completed one generic p16384/o128 request at a diagnostic-only 5.219 tok/s. The semantic program then returned one correct 16,213-token fresh response, corrupted the identical same-server repeat into repeated text despite zero reported cache use, and stopped a separate fresh-server attempt at 1,600 computed prompt tokens with no first output after an RPC sampling timeout. Cleanup of the latter recorded eight B70 engine resets and 61 unsuccessful responses before all four cards re-enumerated idle. The cell is nondeterministically unstable and receives no speed, curve, quality, deployment, or headline credit; unchanged retries and 24K/32K serving remain blocked pending a material runtime treatment. Observed: 16K sometimes completes; same-server corruption and fresh-server timeout; diagnostic only.evidence - ⚠ Quarantined16K, + speculative decoding (depth 2)
16K·MTP2— The current-source MTP2 identity admitted exactly 40 blocks / 20,014 cache tokens in both bounded arms. Scheduler-64 stopped at 3,200 computed prompt tokens; the scheduler-32 arm was observed 2,240 tokens / 70% farther at 5,440, but the sole p16384/o128 treatment request still returned no output before the unchanged 300-second runtime response deadline. This is treatment evidence, not repeat-confirmed causality. Its shutdown was followed by eight card reset records and 58 unsuccessful card responses, so the frozen postflight rule fails and the treatment tranche is exhausted. No capability, speed, quality, parity, deployment, or headline credit is granted; current checks found no listener or owned residue and all cards are below 43 MiB. Observed: exact 16K stopped at 5,440 computed prompt tokens after scheduler-32 treatment; no output; four-card teardown events; diagnostic only.evidence
13 untested combinations
24K·MTP032K·MTP016K·MTP124K·MTP132K·MTP124K·MTP232K·MTP216K·MTP324K·MTP332K·MTP316K·MTP424K·MTP432K·MTP4Card-fit summary
Official FP8 child artifact, vLLM 658965050 plus kernels 2f829747, eager MTP0 text, zero prior context, automatic KV, and TP=EP. TP1 and TP2 need separate fit/offload designs; they are not claimed unsupported. Fixed: revision=qwen38-flash-next · artifact_id=qwen38-flash-next-fp8-bcd9f01 · runtime=vLLM XPU 658965050 + kernels 2f829747 · runtime_family=vLLM XPU · parallel_profile=TP=EP · active_context_tokens=0 · graph_mode=off · kv=auto · modality=text.
2 untested combinations
TP1·MTP0TP2·MTP0Graph and modality summary
Official FP8 child artifact, vLLM 1372c62d plus staged kernels 2f829747, TP4/EP4, MTP0, zero prior context, and automatic KV. Eager text is screened; PIECEWISE text has bounded compile-resource evidence but no API, quality, or speed result; vision remains unmeasured. Fixed: revision=qwen38-flash-next · artifact_id=qwen38-flash-next-fp8-bcd9f01 · runtime=vLLM XPU 1372c62d + staged kernels 2f829747 · runtime_family=vLLM XPU · tp=4 · ep=4 · parallel_profile=TP=EP · mtp=0 · active_context_tokens=0 · kv=auto.
- ◇ Speed check onlytext, eager
text·eager5.22 tok/sreport · evidence - ⚠ Quarantinedtext, graph
text·graph— Attempt 7 loaded the complete checkpoint on all four ranks and entered PIECEWISE torch.compile with one compile thread. The phase-aware guard latched below its 30-GiB host-memory floor before the later TTM/global-OOM teardown window. Graph capture did not complete, the API never became healthy, and no client, output, quality, replay, or speed row exists. This is Grade-D bounded resource evidence only. Observed: all 131 shards loaded on four ranks; graph compile began; 30-GiB host-memory guard tripped; no API, quality, or speed.evidence
2 untested combinations
vision·eagervision·graphFull 270-cell coverage contracts · 26 classified
Text serving coverage
Exact text-serving combinations across topology, MTP depth, active context, and graph mode on the latest retained runtime. Older-runtime evidence is preserved in a separate archival contract and never silently transferred into these cells.
Fixed: runtime=vLLM XPU 1372c62d + staged kernels 2f829747 · runtime_family=vLLM XPU · parallel_profile=TP=EP
Break down by axis
Fixed-vision serving coverage
Vision capability and quality combinations under one versioned fixed-image fixture. Context depth is intentionally excluded until a vision anchor establishes meaningful image-plus-text token accounting.
Fixed: runtime=vLLM XPU 1372c62d + staged kernels 2f829747 · runtime_family=vLLM XPU · parallel_profile=TP=EP · active_context_tokens=0 · modality=vision · workload_profile=fixed-vision-fixture-v1-required
Break down by axis
Measured results
Every number links to its proof.
Research speed screen by TP ◇ screened, experimental
Only TP4 has bounded screens; each MTP cell has three same-server samples under the declared workload, so no TP mini-curve is implied
official FP8 · eager MTP1 4: 9.37
official FP8 · eager MTP2 (variable screen) 4: 11.9
official FP8 · eager MTP3 (variable screen) 4: 14.89
official FP8 · eager MTP4 4: 20.73
official FP8 · eager MTP1 4: 9,349
official FP8 · eager MTP2 (variable screen) 4: 11,278
official FP8 · eager MTP3 (variable screen) 4: 11,818
official FP8 · eager MTP4 4: 10,023
Legacy-comparable after-first-text context screen ◇ screened, experimental
The plotted 5.22/5.13/5.23/5.23 tok/s values use the legacy after-first-text comparison and stop at 4K. Separate formal cache-zero rates are 3.865 at 2K, 4.456 at 4K, and 3.980 tok/s at 8K; 8K is not plotted here because its third comparison row did not complete. All cells retain the short-quality caveat.
Exact-4K service tradeoff by MTP depth ◇ screened, experimental
Workload-aligned p4096/o256/c1 service screens with no harness-added warmups, but from different vLLM source revisions and cache allocations. MTP1 and MTP2 use 32-block headroom recipes; MTP3 remains faster with lower TTFT and higher wall output. Separate single-point series avoid implying a causal MTP-depth or cache curve; only MTP4 remains quarantined.
official FP8 · TP4 eager MTP1 · headroom32 1: 8.9
official FP8 · TP4 eager MTP2 · headroom32 2: 9.89
official FP8 · TP4 eager MTP3 3: 15.5
official FP8 · TP4 eager MTP1 · headroom32 1: 0.98
official FP8 · TP4 eager MTP2 · headroom32 2: 0.89
official FP8 · TP4 eager MTP3 3: 1.25
official FP8 · TP4 eager MTP1 · headroom32 1: 232,079
official FP8 · TP4 eager MTP2 · headroom32 2: 263,279
official FP8 · TP4 eager MTP3 3: 187,899
Fine print
Qwen4Exp hybrid Gated DeltaNet/QSA sparse MoE, 48 layers, 2560 hidden. The FP8 repository is a quantized child of Qwen3.8 Flash-Next, not a separate model. Measurements remain artifact- and runtime-specific. Measurements, artifact hashes, outputs, quality decisions, and speed stay pinned to their exact recorded identity.
Many people at once Not measured
Multi-user (aggregate) throughput has not been measured for this family. The lab’s one measured sweep so far is in the multi-user report; the projection block below includes a projected users curve where a projection exists.
How much faster could this get? No projection
No projection is shown: this model is not in the ML Bottleneck catalog yet. The measured numbers above stand on their own.