Gemma 4 · 26B A4B

Google's Gemma 4, the 26B mixture-of-experts with about 4B active per word. A strong all-rounder for chat, writing, and homework that runs on a single Arc Pro B70 and can also read images.

122.16tok/s
measured
≈ 92 words a secondgemma4-26b-a4b-it-gguf-3bb10d5 · UD-Q8_K_XL target + Q4_0 MTP draft · llama.cpp SYCL c926ad09857517978575d6a74d225b463f7417a0 · 1× B70 · B70-verified source reconstruction
Gemma 4 26B A4B IT · UD-Q8_K_XL target · target/verifier checkpoint
B70 fit
High
Quality evidence
Strong deployment evidence
Measured results
3

Packets and recipes

The deployment variants of this family, at every maturity.

What has been classified

Dense scorecards summarize every declared combination; measured slices retain exact evidence links.

MTP depth × TP

Gemma 4 26B A4B · UD-Q8_K_XL target + locally derived Q4_0 MTP draft · llama.cpp SYCL · f16 KV · 32K configured ceiling. Fixed: revision=gemma4-26b-a4b-it-gguf-3bb10d5 · variant=UD-Q8_K_XL target + locally derived Q4_0 MTP draft · runtime=llama.cpp SYCL c926ad09857517978575d6a74d225b463f7417a0 · kv=f16 · configured_max_context_tokens=32768.

  • ✓ Measured1 card, + speculative decoding (depth 3) TP1·MTP3122.16 tok/sevidence
14 untested combinations
TP1·MTP0TP2·MTP0TP4·MTP0TP1·MTP1TP2·MTP1TP4·MTP1TP1·MTP2TP2·MTP2TP4·MTP2TP2·MTP3TP4·MTP3TP1·MTP4TP2·MTP4TP4·MTP4
Full matrix and exact selectors

Measured results

Every number links to its proof.

family data

Speed as the conversation grows

Gemma 4 UD-Q8_K_XL target + Q4_0 MTP draft · TP1 llama.cpp · four-lane arithmetic means; prefill is prompt tokens / TTFT

045.390.7136181.37412.8K5.6K11K16.2K22.7K30.4K32.6Kactive context tokens · tok/smeasured service profile · Decode (decode_tok_s) · active context tokens=741.0 · value=161.89843663684636 tok/smeasured service profile · Decode (decode_tok_s) · active context tokens=2806.0 · value=144.8662058096366 tok/smeasured service profile · Decode (decode_tok_s) · active context tokens=5643.0 · value=141.11391660033576 tok/smeasured service profile · Decode (decode_tok_s) · active context tokens=10976.0 · value=135.93492558486764 tok/smeasured service profile · Decode (decode_tok_s) · active context tokens=16213.0 · value=127.63665710590647 tok/smeasured service profile · Decode (decode_tok_s) · active context tokens=22730.0 · value=120.48241930243164 tok/smeasured service profile · Decode (decode_tok_s) · active context tokens=30400.0 · value=114.00353788462567 tok/smeasured service profile · Decode (decode_tok_s) · active context tokens=32571.0 · value=114.8486529751413 tok/s
measured service profile
measured service profile 741: 161.9 · 2.8K: 144.87 · 5.6K: 141.11 · 11K: 135.93 · 16.2K: 127.64 · 22.7K: 120.48 · 30.4K: 114 · 32.6K: 114.85
evidence 1

Record history

Distinct MTP depths and context contracts; values are shown as historical operating points, not a controlled depth comparison

qualified records · draft tokens 3122.16tok/sDecodelab measurement; quality scope in evidence · evidence
qualified records · draft tokens 795.26tok/sDecodelab measurement; quality scope in evidence · evidence
qualified records
evidence 1 · evidence 2
Fine print

Gemma 4 26B A4B MoE, 30 layers. single pinned target revision and reconstructed auxiliary draft path. Measurements, artifact hashes, outputs, quality decisions, and speed stay pinned to their exact recorded identity.

4 combinations the lab has stopped pursuing: No qualified multi-card performance profile is stored for this exact target, draft, runtime, and gate identity. evidence The architecture's trained context capacity is not a measurement; the stored service sweep ends at 32,571 active tokens. evidence The packaged and measured deployment is text-only; no multimodal package result is stored. evidence The current package has not completed a clean-host rebuild and endpoint replay, and the historical local Q4_0 draft identity was not retained. evidence Exact selectors are in the family data.

Many people at once Not measured

Multi-user (aggregate) throughput has not been measured for this family. The lab’s one measured sweep so far is in the multi-user report; the projection block below includes a projected users curve where a projection exists.

How much faster could this get? Projected — not measured

The ML Bottleneck physics engine projects a tuned-run target and the physical ceiling for the headline setup (UD-Q8_K_XL target + Q4_0 MTP draft · llama.cpp SYCL · 1× B70). The grade is optimization headroom against the tuned-run target, not model quality.