Packets and recipes
The deployment variants of this family, at every maturity.
Laguna S 2.1 INT4 · TP4+EP4 + DFlash11
Reproduce 125.46 tok/s on 4× B70 · candidate package
Conventional 99-interval median across the sealed 13-prompt, one-start, cache-zero record gate.Laguna S 2.1 INT4 · historical TP4+EP4 + DFlash11
Evidence packet · superseded-approved-record
Laguna S 2.1 INT4 · TP4+EP4 long-context quick check
Evidence packet · bounded-research-screen-not-q1-exact
lab report — documents the result; not a step-by-step install guideWhat has been classified
Dense scorecards summarize every declared combination; measured slices retain exact evidence links.
long context · screened
Pinned INT4 target and DFlash draft · BF16 KV · q12 verifier · one request · cache zero. TP4 points are bounded retrieval-screen medians, not exact-output measurements; TP1, TP2, and TP3 were not run. Fixed: revision=laguna-s-2.1-int4-4bbfc28 · variant=INT4 target + INT4 DFlash draft / BF16 KV · runtime=vLLM XPU 1a7f61feffbc61b21b73f812d231c7426386ccdc + XPU kernels 99886d783372e621941228250091dc8ebdc1595d · profile_id=laguna-q12-long-context-gpu080-composed · ep=4 · speculative_method=DFlash · draft_depth=11 · verifier_width=12 · kv=bfloat16 · configured_max_context_tokens=32768 · output_tokens=128.
- ◇ Speed check only4 cards, 1K
TP4·1K— Two-row median after excluding the first-live capture/JIT row; retrieval and q1 output equality passed; DFlash acceptance 23.56%.D153.6 · P5,172 · T201.77 · AR0.24evidence - ◇ Speed check only4 cards, 4K
TP4·4K— Three-row median; retrieval passed but q1 output equality failed; DFlash acceptance 11.36%.D80.24 · P7,332 · T564.3 · AR0.11evidence - ◇ Speed check only4 cards, 8K
TP4·8K— Three-row median; retrieval passed but q1 output equality failed; DFlash acceptance 2.34%.D46.81 · P4,130 · T1,993 · AR0.02evidence - ◇ Speed check only4 cards, 16K
TP4·16K— Three-row median; retrieval passed but q1 output equality failed; DFlash acceptance 0.86%.D40.41 · P5,111 · T3,228 · AR0.01evidence - ◇ Speed check only4 cards, 24K
TP4·24K— Three-row median; retrieval passed but q1 output equality failed; DFlash acceptance 0.85%.D39.78 · P5,053 · T4,883 · AR0.01evidence - ◇ Speed check only4 cards, 32,640
TP4·32,640— Three-row median; retrieval passed but q1 output equality failed; DFlash acceptance 0.47%.D39.59 · P7,345 · T4,478 · AR0evidence
Codes in the rows: D = decode tok/s · P = prefill tok/s · T = ms to first token · AR = share of drafted tokens accepted · A = combined tok/s.
18 untested combinations
TP1·1KTP2·1KTP3·1KTP1·4KTP2·4KTP3·4KTP1·8KTP2·8KTP3·8KTP1·16KTP2·16KTP3·16KTP1·24KTP2·24KTP3·24KTP1·32,640TP2·32,640TP3·32,640Measured results
Every number links to its proof.
How the record improved
Same pinned INT4 target and DFlash draft, TP4+EP4, DFlash11, width-12 verifier, BF16 KV; runtime revisions differ and values use conventional 99-interval accounting
Draft-model efficiency (earlier runs)
The sealed 2026-07-26 packet is the only stored family measurement with aggregate acceptance and emitted tokens per draft cycle
Fine print
single pinned target/draft pair with revision-specific runtime and native artifacts. Measurements, artifact hashes, outputs, quality decisions, and speed stay pinned to their exact recorded identity.
4 combinations the lab has stopped pursuing: No qualified performance or quality profile is stored for this exact target/draft pair outside the four-card TP4+EP4 topology. evidence A bounded TP4 retrieval screen covers 1K through 32,640, but every 4K-and-longer output diverged from the target-only q1 oracle; no promotion-grade exact context curve is stored. evidence The current record remains tied to exact originating-host runtime and native artifacts; a portable runtime rebuild is not stored. evidence No non-originating-host replay has completed the full record gate. evidence Exact selectors are in the family data.
Many people at once Not measured
Multi-user (aggregate) throughput has not been measured for this family. The lab’s one measured sweep so far is in the multi-user report; the projection block below includes a projected users curve where a projection exists.
How much faster could this get? Projected — not measured
The ML Bottleneck physics engine projects a tuned-run target and the physical ceiling for the headline setup (laguna-s-2.1-int4- · INT4 target + INT4 DFlash draft / BF16 KV · vLLM XPU · 4× B70). The grade is optimization headroom against the tuned-run target, not model quality.
Loading projections from mlbottleneck.com…