I have one B70
Start with Qwen
The Qwen3.8 Q4 package is our most complete one-card candidate. Gemma is faster in this lab, but its self-contained package is still being finished.
Open one-card packageUnofficial Intel XPU Optimization Lab · community-run, not affiliated with Intel
Find an AI model for writing, coding, and everyday questions on your own PC. Explore tested setups for Intel Arc Pro B70 graphics cards, with measured speeds and guides to help you get started.
Speeds are in tok/s: a token is about ¾ of a word, so 100 tok/s is roughly 75 words a second. Compare the hardware and speeds below, then open a model’s details for setup instructions.
| Model and deployment | Best for | 1 user | 32K input | Many users | Actions |
|---|---|---|---|---|---|
| LFM2.5 2.6B 1× B70 · Q8_0 · no MTP · Liquid AI | Lightweight | 132.14 | 89.94† | —not measured | |
| Ornith 1.5 35B-A3B 1× B70 · Q4_K_M · no MTP · Ornith AI | General | —strict headline pending | 99.61† | 216.5† | |
| Laguna-S-2.1 4× B70 · INT4 · DFlash depth 11 · Poolside | Coding | 125.46 | —withheld after quality failure | —not measured | |
| Gemma-4-26B 1× B70 · Q8 · MTP depth 3 · Google | General | 122.16 | 114.85† | —not measured | |
| Qwen3.5-9B 1× B70 · FP8 · MTP depth 3 · Alibaba | General | 98.14 | 86.8† | 564.7 | |
| Qwen3.5-9B 1× B70 · INT4 W4A16 · MTP depth 3 · Alibaba | General | 113.27 | 89.5† | 1268.4 | |
| Qwen3.5-4B 1× B70 · INT4 W4A16 · MTP depth 3 · Alibaba | General | 177.29 | 149.9† | 1593.9 | |
| Muse-Glimmer-30B 4× B70 · Q8/WOQ · DFlash draft n=15 · Meta | Vision | 100.37 | —not measured | —not measured | |
| Qwen3.6-35B-A3B Alibaba · uses ~3B of 35B per word Long context | |||||
| 4× B70 · Quark W8A8 INT8 · no MTP | 93.55 | —not measured | —not measured | ||
| 1× B70 · AutoRound INT4 · no MTP · experimental | 90.91 | —not measured | 1,039 | ||
| MiniMax-M2.7 4× B70 · AutoRound INT4 · no MTP · MiniMax | Long context | 89.31 | 63.91† | —not measured | |
| DeepSeek-V4-Flash-180B 4× B70 · FP8 + FP4 experts · DSpark depth 7 · community trim · experimental | Research | 80.82 | —not measured | —not measured | |
| Nemotron 3.5 Lightning 30B-A3B 1× B70 · UD-Q4_K_M · no MTP · NVIDIA | General | —strict headline pending | 64.62† | —not measured | |
| Ornith 1.5 9B 1× B70 · Q8_0 · no MTP · Ornith AI | Beginner | —strict headline withheld after output-gate failure | 39.84† | —not measured | |
| Qwen3.8-27B Alibaba · newest all-rounder · reads pictures General | |||||
| 1× B70 · Q4_K_M · no MTP | 27.83 | 24.49† | 83.8† | ||
| 1× B70 · Q4_K_M + Q4_0 draft · MTP depth 2 | Fast interactive | 42.64 | 36.51† | 68.3† | |
| 1× B70 · Q8_0 · no MTP | Highest quality | 19.62 | 18.02† | 68.6† | |
| 1× B70 · Q8_0 + Q4_0 draft · MTP depth 2 | Fast high-quality | 37.06 | —not measured | —not measured | |
| 2× B70 · Q4_K_M · no MTP | 49.72 | 44.44† | 192.3† | ||
| 2× B70 · Q4_K_M + Q4_0 draft · MTP depth 2 | Fast interactive | 64.24 | —not measured | —not measured | |
| 2× B70 · Q8_0 · no MTP | Highest quality | 36.73 | 33.85† | 163.6† | |
| 2× B70 · official FP8 + lab W8A16 · no MTP | 33.31 | 29.78† | 931.4 | ||
| 2× B70 · official FP8 + lab W8A16 · MTP depth 1 | Fast interactive FP8 | 54.60 | 46.64† | 474.3 | |
| 2× B70 · AutoRound INT4 fixed-K (lab oneDNN W4A16) · no MTP · XPU graphs | 49.86 | 42.83† | 1000.2 | ||
| 2× B70 · AutoRound INT4 fixed-K · MTP depth 4 · XPU graphs + draft-only INT4 head | Fastest lossless single user | 112.36 | 100.27† | 641.3 | |
| 2× B70 · INT4 · MTP depth 5 · experimental | 101.17 | —not measured | —not measured on this host | ||
| Qwen3.8 Flash-Next 125B-A6B Alibaba · four-card setup · experienced users | |||||
| 4× B70 · FP8 · no MTP | 34.50 | —not measured | —not measured | ||
| 4× B70 · FP8 · MTP depth 1 | 37.83 | —not measured | —not measured | ||
† Different test conditions; see the model details. — No qualified result yet. Experimental setups still have quality or repeatability checks open. Model names and speed labels explained →
A B70 has 32 GB of graphics memory. Some deployments fit on one card; larger or faster research configurations use two to four.
I have one B70
The Qwen3.8 Q4 package is our most complete one-card candidate. Gemma is faster in this lab, but its self-contained package is still being finished.
Open one-card packageI have two to four B70s
Explore the Qwen3.8 FP8 two-card setup. The guide covers requirements, measured performance, and setup limitations.
Open two-card packageNew to local AI?
Find out how much memory you need, what the model names mean, and how to choose a setup for your PC.
Explore the beginner guidesThese setups are still being tested for installation on a fresh PC. Check each guide’s requirements before starting. Browse all recipes · See recent improvements
Measured on the lab's own Arc Pro B70 machines and re-checked; every number links to its proof. One row per model at its fewest cards — hover a column heading for what it means.
| Model | Software | Weights | Generation | Cards | Speed (tok/s) | Proof |
|---|---|---|---|---|---|---|
| Lab-verified: Laguna-S-2.1 coding assistant · Jul 2026 | vLLM XPU | INT4 W4A16 | DFlash · depth 11 | 4×fewer untested | 125.46Fastest promoted row in this table. | test report |
| Lab-verified: Gemma-4-26B-A4B all-rounder · uses ~4B of 26B per word · Apr 2026 | llama.cpp SYCL | UD-Q8_K_XL | MTP · depth 3 | 1× | 122.16 | test report |
| Lab-verified: Qwen3.5-9B compact all-rounder · thinks before answering · 2026 | vLLM XPU | FP8-dynamic | MTP · depth 3 | 1× | 98.14 | test report |
| Lab-verified: Qwen3.5-9B compact all-rounder · thinks before answering · 2026 | vLLM XPU | INT4 W4A16 | MTP · depth 3 | 1× | 113.27 | test report |
| Lab-verified: Muse-Glimmer-30B Meta · reads pictures · Aug 2026 | llama.cpp SYCL | UD-Q8_K_XL | DFlash draft · n=15 | 4×needs ≥2 | 100.37 | full report |
| Lab-verified: Qwen3.6-27B general assistant · uses all 27B per word · Apr 2026 | vLLM XPU | AutoRound INT4 | MTP · depth 3 | 2×fits on 1 · TP2 | 98.77 | full rankings |
| Lab-verified: Qwen3.6-35B-A3B long documents · uses ~3B of 35B per word · Apr 2026 | vLLM XPU | Quark W8A8 INT8 | no MTP | 4×needs ≥2 | 93.55 | test report |
| Lab-verified: MiniMax-M2.7 remembers long conversations (32K) · Mar 2026 | vLLM XPU | AutoRound INT4 W4A16 | no MTP | 4×needs 4 | 89.31 | full report |
| Lab-verified: DeepSeek-V4-Flash-180B experimental community-trimmed copy · research only | vLLM XPU | FP8 + FP4 experts | DSpark · depth 7 | 4×needs 4 | 80.82 | full rankings |
| Lab-tested: Qwen3.8-27B newest all-rounder · reads pictures · Aug 2026 | llama.cpp SYCL | Q4_K_M | no MTP | 1× | 27.83 | candidate guide |
R = research lane, run-to-run determinism still open · hover a column heading for what it means · every number links to its proof.
Estimates from ML Bottleneck (same author) show how much faster each setup might run with further optimization. These are projections, not measured speeds; grades describe speed potential, not answer quality.
Loading projections from mlbottleneck.com…
Not on the table yet? Pick a model, card, compression, and software to get a projected decode and prompt-processing rate for Intel Arc Pro cards, plus whether it fits. The full planner on mlbottleneck.com adds memory maps, scaling charts, and every other GPU.
Setups shared by the community and tested on our machines. Meet the contributors and see the tests.
| Model | Software | Weights | Generation | Cards | Speed (tok/s) | Proof |
|---|---|---|---|---|---|---|
| B70-tested: Qwen3.6-27B community setup we re-ran and confirmed | Docker / vLLM | FP8 | no MTP | 2× | 30.17 | lab check |
Community optimization
Share a speed improvement, setup guide, or correction. We test what we can and credit your contribution. Experiments that didn’t work are useful too.