Candidate package · llama.cpp SYCL · 1× Intel Arc Pro B70

Gemma 4 26B A4B UD-Q8_K_XL on one Intel Arc Pro B70

Google's Gemma 4 26B mixture-of-experts on one Arc Pro B70 in 8-bit form, with its built-in draft head speeding up generation. A strong all-rounder that also reads images.

What we measured Lab-measured

122tok/s decode (MTP-assisted)

Class-balanced median of per-input-class medians using conventional 99-interval rates on the fixed cold suite with the Q4_0 MTP draft assisting. The all-prompt median is 123.727369 tok/s; the historical 100-event compatibility figure is 124.977141 tok/s. proof file

99-interval median = the middle rate across 99 measured stretches of a long answer · MTP = multi-token prediction, a small draft the main model verifies

Model
Gemma 4 26B A4B
Publisher
Google
Checkpoint
unsloth/gemma-4-26B-A4B-it-GGUF
Compression
UD-Q8_K_XL + Q4_0 MTP
Software
llama.cpp SYCL
Cards
1× Intel Arc Pro B70 32 GiB
Model weight bytes
27.6 GB
Operating systems
Linux
Delivery
native
Good for
general coding long context
Published
2026-08-22
Clean-host replay
not yet

Still missing before this becomes an install guide

  • historical llama-server binary SHA-256
  • historical locally quantized Q4_0 MTP draft SHA-256 and byte size (the reconstructed draft is repeat-stable at 1f6706e4… across three builds)
  • beginner recovery flow exercised end-to-end on a fresh host (documented; the pinned oneAPI 2026.0 container path and two clean-rebuild gate replays were verified on the lab host on 2026-09-07)

Measured performance profiles Lab-measured

Decode over active context Lab-measured

1621451411361281201141157412.7K5.5K10.7K15.8K22.2K29.7K31.8KActual prompt / active context tokens · y: tok/s
Mean after-TTFT decode rate across four independent one-B70 lanes; unique cache-zero prompts and exact JSON gates. evidence

Prefill over prompt length Lab-measured

8481,4011,4691,3301,2571,1291,0281,0027412.7K5.5K10.7K15.8K22.2K29.7K31.8KActual prompt tokens · y: tok/s
Approximate prompt tokens divided by TTFT, averaged across four independent one-B70 lanes. evidence

Time to first token Lab-measured

8742,0033,8418,25112,89920,13429,56832,5177412.7K5.5K10.7K15.8K22.2K29.7K31.8KActual prompt tokens · y: ms
Mean TTFT across four independent one-B70 lanes under the same cache-zero service sweep. evidence

Many people at once Not published

No qualified multi-user aggregate profile is published for this exact package. Diagnostic or unsupported boundaries may still appear under “What to know” or in the full guide. Nothing is interpolated or promoted from a different model, quantization, runtime, or card count; the projection below remains clearly labeled as projected.

How much faster could this get? Projected — not measured

The ML Bottleneck physics engine projects what stock software, a tuned run, and the physical ceiling look like for this exact model, compression, card count, and software. The grade is optimization headroom against the tuned-run target, not model quality.