B70 fit High
Quality evidence Strong deployment evidence
Measured results 3
Packets and recipes The deployment variants of this family, at every maturity.
What has been classified Dense scorecards summarize every declared combination; measured slices retain exact evidence links.
MTP depth × TP Gemma 4 26B A4B · UD-Q8_K_XL target + locally derived Q4_0 MTP draft · llama.cpp SYCL · f16 KV · 32K configured ceiling. Fixed: revision=gemma4-26b-a4b-it-gguf-3bb10d5 · variant=UD-Q8_K_XL target + locally derived Q4_0 MTP draft · runtime=llama.cpp SYCL c926ad09857517978575d6a74d225b463f7417a0 · kv=f16 · configured_max_context_tokens=32768.
✓ Measured 1 card, + speculative decoding (depth 3) TP1·MTP3 122.16 tok/s evidence 14 untested combinations TP1·MTP0TP2·MTP0TP4·MTP0TP1·MTP1TP2·MTP1TP4·MTP1TP1·MTP2TP2·MTP2TP4·MTP2TP2·MTP3TP4·MTP3TP1·MTP4TP2·MTP4TP4·MTP4
Full matrix and exact selectors
Measured results Every number links to its proof.
family data
Speed as the conversation grows Gemma 4 UD-Q8_K_XL target + Q4_0 MTP draft · TP1 llama.cpp · four-lane arithmetic means; prefill is prompt tokens / TTFT
Decode Prefill TTFT
0 45.3 90.7 136 181.3 741 2.8K 5.6K 11K 16.2K 22.7K 30.4K 32.6K active context tokens · tok/s measured service profile · Decode (decode_tok_s) · active context tokens=741.0 · value=161.89843663684636 tok/s measured service profile · Decode (decode_tok_s) · active context tokens=2806.0 · value=144.8662058096366 tok/s measured service profile · Decode (decode_tok_s) · active context tokens=5643.0 · value=141.11391660033576 tok/s measured service profile · Decode (decode_tok_s) · active context tokens=10976.0 · value=135.93492558486764 tok/s measured service profile · Decode (decode_tok_s) · active context tokens=16213.0 · value=127.63665710590647 tok/s measured service profile · Decode (decode_tok_s) · active context tokens=22730.0 · value=120.48241930243164 tok/s measured service profile · Decode (decode_tok_s) · active context tokens=30400.0 · value=114.00353788462567 tok/s measured service profile · Decode (decode_tok_s) · active context tokens=32571.0 · value=114.8486529751413 tok/s 0 411.4 822.9 1,234 1,646 741 2.8K 5.6K 11K 16.2K 22.7K 30.4K 32.6K active context tokens · tok/s measured service profile · Prefill (prefill_tok_s) · active context tokens=741.0 · value=847.9619132833765 tok/s measured service profile · Prefill (prefill_tok_s) · active context tokens=2806.0 · value=1401.0150752344389 tok/s measured service profile · Prefill (prefill_tok_s) · active context tokens=5643.0 · value=1469.4385849395155 tok/s measured service profile · Prefill (prefill_tok_s) · active context tokens=10976.0 · value=1330.3881630879202 tok/s measured service profile · Prefill (prefill_tok_s) · active context tokens=16213.0 · value=1256.9527613956354 tok/s measured service profile · Prefill (prefill_tok_s) · active context tokens=22730.0 · value=1128.9763236685953 tok/s measured service profile · Prefill (prefill_tok_s) · active context tokens=30400.0 · value=1028.1957534609096 tok/s measured service profile · Prefill (prefill_tok_s) · active context tokens=32571.0 · value=1001.687171893108 tok/s 0 9,105 18,210 27,315 36,419 741 2.8K 5.6K 11K 16.2K 22.7K 30.4K 32.6K active context tokens · ms measured service profile · TTFT (ttft_ms) · active context tokens=741.0 · value=873.9445414976217 ms measured service profile · TTFT (ttft_ms) · active context tokens=2806.0 · value=2002.938948746305 ms measured service profile · TTFT (ttft_ms) · active context tokens=5643.0 · value=3840.5072987661697 ms measured service profile · TTFT (ttft_ms) · active context tokens=10976.0 · value=8250.644568965072 ms measured service profile · TTFT (ttft_ms) · active context tokens=16213.0 · value=12899.11724528065 ms measured service profile · TTFT (ttft_ms) · active context tokens=22730.0 · value=20134.24033729825 ms measured service profile · TTFT (ttft_ms) · active context tokens=30400.0 · value=29567.641104455106 ms measured service profile · TTFT (ttft_ms) · active context tokens=32571.0 · value=32517.333841737127 ms
measured service profile
measured service profile 741: 161.9 · 2.8K: 144.87 · 5.6K: 141.11 · 11K: 135.93 · 16.2K: 127.64 · 22.7K: 120.48 · 30.4K: 114 · 32.6K: 114.85
measured service profile 741: 847.96 · 2.8K: 1,401 · 5.6K: 1,469 · 11K: 1,330 · 16.2K: 1,257 · 22.7K: 1,129 · 30.4K: 1,028 · 32.6K: 1,002
measured service profile 741: 873.94 · 2.8K: 2,003 · 5.6K: 3,841 · 11K: 8,251 · 16.2K: 12,899 · 22.7K: 20,134 · 30.4K: 29,568 · 32.6K: 32,517
All measured metric summaries Decode measured service profile 741: 161.9 · 2.8K: 144.87 · 5.6K: 141.11 · 11K: 135.93 · 16.2K: 127.64 · 22.7K: 120.48 · 30.4K: 114 · 32.6K: 114.85Prefill measured service profile 741: 847.96 · 2.8K: 1,401 · 5.6K: 1,469 · 11K: 1,330 · 16.2K: 1,257 · 22.7K: 1,129 · 30.4K: 1,028 · 32.6K: 1,002TTFT measured service profile 741: 873.94 · 2.8K: 2,003 · 5.6K: 3,841 · 11K: 8,251 · 16.2K: 12,899 · 22.7K: 20,134 · 30.4K: 29,568 · 32.6K: 32,517
evidence 1
Record history Distinct MTP depths and context contracts; values are shown as historical operating points, not a controlled depth comparison
qualified records · draft tokens 3 122.16 tok/s Decode lab measurement; quality scope in evidence · evidence qualified records · draft tokens 7 95.26 tok/s Decode lab measurement; quality scope in evidence · evidence
qualified records
evidence 1 · evidence 2
Fine print Gemma 4 26B A4B MoE, 30 layers. single pinned target revision and reconstructed auxiliary draft path. Measurements, artifact hashes, outputs, quality decisions, and speed stay pinned to their exact recorded identity.
4 combinations the lab has stopped pursuing: No qualified multi-card performance profile is stored for this exact target, draft, runtime, and gate identity. evidence The architecture's trained context capacity is not a measurement; the stored service sweep ends at 32,571 active tokens. evidence The packaged and measured deployment is text-only; no multimodal package result is stored. evidence The current package has not completed a clean-host rebuild and endpoint replay, and the historical local Q4_0 draft identity was not retained. evidence Exact selectors are in the family data .
Many people at once Not measured
Multi-user (aggregate) throughput has not been measured for this family. The lab’s one measured sweep so far is in the multi-user report ; the projection block below includes a projected users curve where a projection exists.
How much faster could this get? Projected — not measured The ML Bottleneck physics engine projects a tuned-run target and the physical ceiling for the headline setup (UD-Q8_K_XL target + Q4_0 MTP draft · llama.cpp SYCL · 1× B70). The grade is optimization headroom against the tuned-run target, not model quality.
Loading projections from mlbottleneck.com…
Projections need JavaScript; the measured numbers above are static.