Unofficial Intel XPU Optimization Lab · community-run, not affiliated with Intel

Run AI at home.
We make it faster.

Find an AI model for writing, coding, and everyday questions on your own PC. Explore tested setups for Intel Arc Pro B70 graphics cards, with measured speeds and guides to help you get started.

Which model should I try?

Speeds are in tok/s: a token is about ¾ of a word, so 100 tok/s is roughly 75 words a second. Compare the hardware and speeds below, then open a model’s details for setup instructions.

Model and deploymentBest for1 user32K inputMany usersActions
LFM2.5 2.6B 1× B70 · Q8_0 · no MTP · Liquid AI Lightweight 132.14 89.94† not measured
Details pageGitHub guide
Ornith 1.5 35B-A3B 1× B70 · Q4_K_M · no MTP · Ornith AI General strict headline pending 99.61† 216.5†
Details pageGitHub guide
Laguna-S-2.1 4× B70 · INT4 · DFlash depth 11 · Poolside Coding 125.46 withheld after quality failure not measured
Details pageGitHub guide
Gemma-4-26B 1× B70 · Q8 · MTP depth 3 · Google General 122.16 114.85† not measured
Details pageGitHub guide
Qwen3.5-9B 1× B70 · FP8 · MTP depth 3 · Alibaba General 98.14 86.8† 564.7
Details pageGitHub guide
Qwen3.5-9B 1× B70 · INT4 W4A16 · MTP depth 3 · Alibaba General 113.27 89.5† 1268.4
Details pageGitHub guide
Qwen3.5-4B 1× B70 · INT4 W4A16 · MTP depth 3 · Alibaba General 177.29 149.9† 1593.9
Details pageGitHub guide
Muse-Glimmer-30B 4× B70 · Q8/WOQ · DFlash draft n=15 · Meta Vision 100.37 not measured not measured
Details pageGitHub guide
Qwen3.6-35B-A3B Alibaba · uses ~3B of 35B per word Long context
4× B70 · Quark W8A8 INT8 · no MTP 93.55 not measured not measured
Details pageGitHub guide
1× B70 · AutoRound INT4 · no MTP · experimental 90.91 not measured 1,039
MiniMax-M2.7 4× B70 · AutoRound INT4 · no MTP · MiniMax Long context 89.31 63.91† not measured
Details pageGitHub guide
DeepSeek-V4-Flash-180B 4× B70 · FP8 + FP4 experts · DSpark depth 7 · community trim · experimental Research 80.82 not measured not measured
Nemotron 3.5 Lightning 30B-A3B 1× B70 · UD-Q4_K_M · no MTP · NVIDIA General strict headline pending 64.62† not measured
Details pageGitHub guide
Ornith 1.5 9B 1× B70 · Q8_0 · no MTP · Ornith AI Beginner strict headline withheld after output-gate failure 39.84† not measured
Details pageGitHub guide
Qwen3.8-27B Alibaba · newest all-rounder · reads pictures General
1× B70 · Q4_K_M · no MTP 27.83 24.49† 83.8†
Details pageGitHub guide
1× B70 · Q4_K_M + Q4_0 draft · MTP depth 2 Fast interactive 42.64 36.51† 68.3†
Details pageGitHub guide
1× B70 · Q8_0 · no MTP Highest quality 19.62 18.02† 68.6†
Details pageGitHub guide
1× B70 · Q8_0 + Q4_0 draft · MTP depth 2 Fast high-quality 37.06 not measured not measured
Details pageGitHub guide
2× B70 · Q4_K_M · no MTP 49.72 44.44† 192.3†
Details pageGitHub guide
2× B70 · Q4_K_M + Q4_0 draft · MTP depth 2 Fast interactive 64.24 not measured not measured
Details pageGitHub guide
2× B70 · Q8_0 · no MTP Highest quality 36.73 33.85† 163.6†
Details pageGitHub guide
2× B70 · official FP8 + lab W8A16 · no MTP 33.31 29.78† 931.4
Details pageGitHub guide
2× B70 · official FP8 + lab W8A16 · MTP depth 1 Fast interactive FP8 54.60 46.64† 474.3
Details pageGitHub guide
2× B70 · AutoRound INT4 fixed-K (lab oneDNN W4A16) · no MTP · XPU graphs 49.86 42.83† 1000.2
Details pageGitHub guide
2× B70 · AutoRound INT4 fixed-K · MTP depth 4 · XPU graphs + draft-only INT4 head Fastest lossless single user 112.36 100.27† 641.3
Details pageGitHub guide
2× B70 · INT4 · MTP depth 5 · experimental 101.17 not measured not measured on this host
Open details
Qwen3.8 Flash-Next 125B-A6B Alibaba · four-card setup · experienced users
4× B70 · FP8 · no MTP 34.50 not measured not measured
4× B70 · FP8 · MTP depth 1 37.83 not measured not measured

† Different test conditions; see the model details. — No qualified result yet. Experimental setups still have quality or repeatability checks open. Model names and speed labels explained →

Start here

A B70 has 32 GB of graphics memory. Some deployments fit on one card; larger or faster research configurations use two to four.

I have one B70

Start with Qwen

The Qwen3.8 Q4 package is our most complete one-card candidate. Gemma is faster in this lab, but its self-contained package is still being finished.

Open one-card package

I have two to four B70s

Put your extra cards to work

Explore the Qwen3.8 FP8 two-card setup. The guide covers requirements, measured performance, and setup limitations.

Open two-card package

New to local AI?

Learn the basics

Find out how much memory you need, what the model names mean, and how to choose a setup for your PC.

Explore the beginner guides

Lab-tested speeds checked on our own machines

Measured on the lab's own Arc Pro B70 machines and re-checked; every number links to its proof. One row per model at its fewest cards — hover a column heading for what it means.

Model Software Weights Generation Cards Speed (tok/s) More cards How measured Proof
Lab-verified: Laguna-S-2.1 coding assistant · Jul 2026 vLLM XPUINT4 W4A16DFlash · depth 11fewer untested 125.46Fastest promoted row in this table. conventional median test report
Lab-verified: Gemma-4-26B-A4B all-rounder · uses ~4B of 26B per word · Apr 2026 llama.cpp SYCLUD-Q8_K_XLMTP · depth 3 122.16 typical · remembers ~24K words test report
Lab-verified: Qwen3.5-9B compact all-rounder · thinks before answering · 2026 vLLM XPUFP8-dynamicMTP · depth 3 98.14 2× 147.8 typical · remembers ~24K words (32K measured) test report
Lab-verified: Qwen3.5-9B compact all-rounder · thinks before answering · 2026 vLLM XPUINT4 W4A16MTP · depth 3 113.27 2× 147.8 8K served · lossless to 64 users test report
Lab-verified: Muse-Glimmer-30B Meta · reads pictures · Aug 2026 llama.cpp SYCLUD-Q8_K_XLDFlash draft · n=15needs ≥2 100.37 average across repeat runs full report
Lab-verified: Qwen3.6-27B general assistant · uses all 27B per word · Apr 2026 vLLM XPUAutoRound INT4MTP · depth 3fits on 1 · TP2 98.77 average across 4 setups full rankings
Lab-verified: Qwen3.6-35B-A3B long documents · uses ~3B of 35B per word · Apr 2026 vLLM XPUQuark W8A8 INT8no MTPneeds ≥2 93.55 ×1 INT4 90.91R typical, strictest quality gate test report
Lab-verified: MiniMax-M2.7 remembers long conversations (32K) · Mar 2026 vLLM XPUAutoRound INT4 W4A16no MTPneeds 4 89.31 average of 4 clean runs full report
Lab-verified: DeepSeek-V4-Flash-180B experimental community-trimmed copy · research only vLLM XPUFP8 + FP4 expertsDSpark · depth 7needs 4 80.82 record-suite high · 78.29 three-suite center · all 24 checks passed full rankings
Lab-tested: Qwen3.8-27B newest all-rounder · reads pictures · Aug 2026 llama.cpp SYCLQ4_K_Mno MTP 27.83 ×2 FP8+W8A16+MTP pending · ×2 INT4+MTP5 101.17R typical of the promoted 1-card package candidate guide

R = research lane, run-to-run determinism still open · hover a column heading for what it means · every number links to its proof.

measured · experimental 1,039 tok/s combined — one Arc Pro B70 serving 64 people at once with Qwen3.6-35B-A3B. Read the multi-user report

Community work we validated B70-tested in this lab

Setups shared by the community and tested on our machines. Meet the contributors and see the tests.

Model Software Weights Generation Cards Speed (tok/s) Proof
B70-tested: Qwen3.6-27B community setup we re-ran and confirmed Docker / vLLMFP8no MTP 30.17 lab check

Community optimization

Made a model faster?

Share a speed improvement, setup guide, or correction. We test what we can and credit your contribution. Experiments that didn’t work are useful too.

  1. Share what changedTell us what you tried, which hardware you used, and how it performed.
  2. We test itWe check speed and answer quality before recommending it.
  3. You get the creditYour name and contribution stay linked from the guides that use your work.