Families first · quantizations are deployment variants

Models

Start with the model family, then choose weights, quantization, cards, context, runtime, and speed-up. Mini graphs show decode, prefill, TTFT, quality, and gaps without turning every permutation into another model.

Model families

Shape-compatible weight updates share implementation work, while every measurement stays pinned to its exact checkpoint and runtime.

Model family · 1 revision · 1 packet

DeepSeek Coder V2 · Lite Instruct

DeepSeek's coding model in its small mixture-of-experts form: 16B parameters total, about 2.4B active per word, built for code completion and programming help. Fits one Arc Pro B70 with room to spare.

TP 1 / 2 / 4 · target only · context, KV, graph, quant, prefill, TTFT, quality

Open the model page →
Model family · 1 revision · 1 packet

DeepSeek V4 · Flash 180B

DeepSeek's large mixture-of-experts reasoning model in a community-trimmed 180B form (about 13B active per word). A research-grade deployment: it needs all four B70 cards and an experimental checkpoint, so treat it as a look at what the biggest open models can do on this hardware, not a daily driver.

TP 1 / 2 / 4 · target / MTP0–4 / DSpark5–8 · context, KV, graph, quant, prefill, TTFT, quality

Open the model page →
Model family · 1 revision · 2 packets

Gemma 4 · 26B A4B

Google's Gemma 4, the 26B mixture-of-experts with about 4B active per word. A strong all-rounder for chat, writing, and homework that runs on a single Arc Pro B70 and can also read images.

TP 1 / 2 / 4 · target-verified MTP 0–7 · context, KV, graph, quant, prefill, TTFT, quality

Open the model page →
Model family · 1 revision · 1 packet

GLM 4.7 · Flash 30B-A3B

Zhipu's GLM-4.7 Flash, a 30B mixture-of-experts with about 3B active per word. A fast general assistant that fits one Arc Pro B70 in 4-bit form.

TP 1 / 2 / 4 · target only · context, KV, graph, quant, prefill, TTFT, quality

Open the model page →
Model family · 1 revision · 3 packets

Laguna S · 2.1

Poolside's Laguna S 2.1, a coding assistant model. Runs across four Arc Pro B70 cards with a draft model speeding up generation; measured as a coding-first deployment.

TP 4 · DFlash 11 · context, KV, graph, quant, prefill, TTFT, quality

Open the model page →
Model family · 1 revision · 1 packet

LFM2.5 2.6B

Liquid AI's LFM2.5, a compact 2.6B model with a hybrid convolution-and-attention design. Very fast on one Arc Pro B70 and a good fit for lightweight assistants and edge-style workloads.

TP 1 / 2 / 4 · target only · context, KV, graph, quant, prefill, TTFT, quality

Open the model page →
Model family · 1 revision · 4 packets

MiniMax M2.7

MiniMax M2.7, a 229B mixture-of-experts (about 10B active per word) built for long conversations and agent-style work. Needs all four Arc Pro B70 cards in 4-bit form.

TP 1 / 2 / 4 · target only · context, KV, graph, quant, prefill, TTFT, quality

Open the model page →
Model family · 1 revision · 1 packet

Mistral Small 3.2 · 24B Instruct 2506

Mistral Small 3.2, a dense 24B instruct model from Mistral AI. A solid general assistant that fits one Arc Pro B70 in 4-bit form.

TP 1 / 2 / 4 · target only · context, KV, graph, quant, prefill, TTFT, quality

Open the model page →
Model family · 1 revision · 3 packets

Muse-Glimmer 30B

Meta's Muse-Glimmer 30B, a dense multimodal model that reads images as well as text. Runs across four Arc Pro B70 cards with a draft model speeding up generation.

TP 1 / 2 / 4 · none / DFlash 0–15 · context, KV, graph, quant, prefill, TTFT, quality

Open the model page →
Model family · 1 revision · 1 packet

Nemotron 3.5 Lightning 30B-A3B

NVIDIA's Nemotron 3.5 Lightning, a 30B hybrid Mamba mixture-of-experts with about 3B active per word. A fast general assistant that runs on one Arc Pro B70 with stock software.

Cards 1 / 2 · target only · context, KV, graph, quant, prefill, TTFT, quality

Open the model page →
Model family · 1 revision · 1 packet

Nemotron Cascade 2 · 30B-A3B

NVIDIA's Nemotron Cascade 2, a 30B hybrid attention/Mamba mixture-of-experts with about 3B active per word. Runs on one Arc Pro B70 in 4-bit form.

TP 1 / 2 / 4 · target only · context, KV, graph, quant, prefill, TTFT, quality

Open the model page →
Model family · 1 revision · 2 packets

Ornith 1.5 · dense 9B + MoE 35B-A3B

Ornith AI's Ornith 1.5 line: a dense 9B and a 35B mixture-of-experts (about 3B active per word), both Qwen-derived. The 35B is the fastest single-card model this lab has measured; the 9B is a smaller stock deployment.

Cards 1 / 2 · target only · context, KV, graph, quant, prefill, TTFT, quality

Open the model page →
Model family · 1 revision · 1 packet

Phi-4 Mini · Instruct

Microsoft's Phi-4 Mini Instruct, a 3.8B model built to punch above its size on reasoning and instruction following. Very fast on one Arc Pro B70; the full-size Phi-4 is not covered yet.

TP 1 / 2 / 4 · target only · context, KV, graph, quant, prefill, TTFT, quality

Open the model page →
Model family · 2 revisions · 72 packets

Qwen 27B · 3.6 → 3.8

Alibaba's dense 27B Qwen line (3.6 and the newer 3.8), the lab's most-tested general assistant: it reads images, handles long context, and has been measured on one, two, and four Arc Pro B70 cards with both llama.cpp and vLLM.

TP 1 / 2 / 3 / 4 · MTP 0–9 · context, KV, graph, quant, prefill, TTFT, quality

Open the model page →
Model family · 1 revision · 1 packet

Qwen3 14B · Instruct

Alibaba's Qwen3 14B Instruct, a dense mid-size general assistant. Fits one Arc Pro B70 in 4-bit form.

TP 1 / 2 / 4 · target only · context, KV, graph, quant, prefill, TTFT, quality

Open the model page →
Model family · 1 revision · 2 packets

Qwen3 30B-A3B · General + Coder

Alibaba's Qwen3 30B mixture-of-experts (about 3B active per word) in its General Instruct and Coder forms. Fast on one Arc Pro B70; the two siblings share a design but not their weights.

TP 1 / 2 / 4 · target only · context, KV, graph, quant, prefill, TTFT, quality

Open the model page →
Model family · 2 revisions · 1 packet

Qwen3.5 4B

Alibaba's smallest dense-hybrid Qwen3.5 with a shipped MTP head, served on one Arc Pro B70. Only the INT4 W4A16 build is packaged: the FP8 build of the same model is not repeat-exact on this stack, while the INT4 route passes every identity gate.

TP 1 / 2 · MTP 0–3 · context, KV, graph, quant, prefill, TTFT, quality

Open the model page →
Model family · 2 revisions · 2 packets

Qwen3.5 9B

Alibaba's dense-hybrid Qwen3.5 9B (linear-attention layers with a full-attention layer every fourth, and a shipped MTP head), served on one Arc Pro B70 as RedHatAI's FP8-dynamic checkpoint through the lab's deterministic vLLM XPU stack with lossless speculative decoding. Two routes are measured: RedHatAI's FP8-dynamic and its W4A16 INT4, the latter faster and byte-exact at every concurrency through 64 users.

TP 1 / 2 · MTP 0–6 · context, KV, graph, quant, prefill, TTFT, quality

Open the model page →
Model family · 1 revision · 2 packets

Qwen3.6 35B A3B

Alibaba's Qwen3.6 35B mixture-of-experts (about 3B active per word), built for long documents. Measured on four cards under the lab's strictest quality gate, on a single card experimentally, and serving 64 people at once from one card.

TP 1 / 2 / 4 · MTP 0–4 · context, KV, graph, quant, prefill, TTFT, quality

Open the model page →
Model family · 1 revision · 11 packets

Qwen3.8 Flash-Next · 125B-A6B

Qwen's experimental 125B-A6B hybrid-attention MoE. The official FP8 export serves on four B70s with selective host placement, and all 25 practical TP4 eager-text MTP/context cells through 8K are classified. The retained current runtime now independently qualifies TP4 eager MTP0 at short context and exact active 4K: 5.224 tok/s on the established short screen and 4.758 tok/s conventional at exact 4K, with 6/7 semantic quality, 16/16 repeats, an exact cache-zero 4K needle, and card-clean teardown. At 16K, current-source MTP0 completed one exact generic-depth request, then the semantic program produced one correct fresh response, a corrupted same-server repeat, and a separate fresh-server worker timeout at 1,600 computed prompt tokens. This is Grade-D quarantined capability and nondeterministic runtime-stability evidence; it does not authorize 24K/32K. MTP2's bounded 16K treatment tranche is exhausted. MTP3 remains the preferred exact-4K recipe at 15.502 tok/s decode. The target-only official-thinking MTP0 profile passed 25/25, and graph, vision, other topologies, clean-host replay, and deployment qualification remain open.

TP 1 / 2 / 4 · MTP 0–4 · context, KV, graph, quant, prefill, TTFT, quality

Open the model page →

Deployment packets

Measured quantized variants, recipes, and research packets at their honest maturity. Speed is one field, not the sorting rule.

Candidate package · llama.cpp SYCL

Gemma 4 26B A4B UD-Q8_K_XL on one Intel Arc Pro B70

Google's Gemma 4 26B mixture-of-experts on one Arc Pro B70 in 8-bit form, with its built-in draft head speeding up generation. A strong all-rounder that also reads images.

122 tok/s measured · 1× B70 · UD-Q8_K_XL + Q4_0 MTP

Candidate package · vLLM XPU

Laguna S 2.1 INT4 with DFlash on four Intel Arc Pro B70 cards

Poolside's Laguna S 2.1 coding assistant on four Arc Pro B70 cards in 4-bit form, with a draft model speeding up generation. A replay of the lab's record run, quality-checked.

125 tok/s measured · 4× B70 · INT4 / BF16 KV

Candidate package · llama.cpp SYCL

LFM2.5 2.6B Q8_0 on one Intel Arc Pro B70

Liquid AI's compact LFM2.5 2.6B on one Arc Pro B70 in 8-bit form with stock, unpatched llama.cpp - the simplest single-command recipe on the site.

132 tok/s measured · 1× B70 · Q8_0

Candidate package · vLLM XPU + llm-scaler

MiniMax M2.7 AutoRound INT4 on four Intel Arc Pro B70 cards

MiniMax M2.7, a 229B mixture-of-experts for long conversations, on four Arc Pro B70 cards in 4-bit form with vLLM. Quality-checked token by token.

89.3 tok/s measured · 4× B70 · AutoRound W4A16 INT4

Candidate package · llama.cpp SYCL

Muse-Glimmer 30B Q8/WOQ with DFlash on four Intel Arc Pro B70 cards

Meta's Muse-Glimmer 30B, a model that reads images as well as text, on four Arc Pro B70 cards in 8-bit form with a draft model speeding up generation.

100 tok/s measured · 4× B70 · UD-Q8_K_XL / BF16 draft

Candidate package · llama.cpp SYCL

Nemotron 3.5 Lightning 30B-A3B UD-Q4_K_M on one Intel Arc Pro B70

NVIDIA's Nemotron 3.5 Lightning, a 30B hybrid Mamba mixture-of-experts, on one Arc Pro B70 in 4-bit form with stock llama.cpp. Its speed barely drops as the conversation grows.

strict headline pending · 1× B70 · UD-Q4_K_M

Candidate package · llama.cpp SYCL

Ornith 1.5 35B-A3B Q4_K_M on one Intel Arc Pro B70

Ornith AI's 35B mixture-of-experts (about 3B active per word) on one Arc Pro B70 in 4-bit form - the fastest single-card result this lab has measured, using the lab's tuned kernel stack.

strict headline pending · 1× B70 · Q4_K_M

Candidate package · llama.cpp SYCL

Ornith 1.5 9B Q8_0 on one Intel Arc Pro B70

Ornith AI's dense 9B on one Arc Pro B70 in 8-bit form with stock llama.cpp. A small, official, beginner-friendly deployment.

strict headline withheld · 1× B70 · Q8_0

Candidate package · vLLM XPU

Qwen3.5 4B W4A16 with its own MTP head on one or two Intel Arc Pro B70

RedHatAI's W4A16 Qwen3.5-4B served by vLLM XPU on one B70 through the lab's R276 image, with the publisher's MTP head as a lossless speculative draft. The FP8-dynamic build of the same model cannot pass the base identity gate on this stack (three fresh-server pairs scored 11/12, 9/12, 11/12 with three tie-prone prompts); on the row-invariant W4A16 INT4 kernel the same gate passes 12/12, which is why this is the packaged route. The kernel is not the only reduction on that path: the RMSNorm this route runs is itself row-count dependent (measured 2026-09-08), so the route is exact in the regimes measured rather than exact by construction.

177 tok/s measured · 1× B70 · W4A16 (compressed-tensors INT4 weights, FP16 activations)

Candidate package · vLLM XPU

Qwen3.5 9B FP8-dynamic with its own MTP head on one Intel Arc Pro B70

RedHatAI's FP8-dynamic Qwen3.5-9B served by vLLM XPU on one B70 through the lab's published R276 image and the Qwen3.8 FP8 recipe's strict launchers, with the model's own MTP head as a lossless speculative draft (qwen3_5_mtp) and full decode-only XPU graph capture. On this hardware the W4A16 build of the same model (packages/qwen35-9b-w4a16-b70) is the better route: faster at every depth and byte-exact at every concurrency through 64 users, where this FP8 route flips a near-tie token from 16 users up.

98.1 tok/s measured · 1× B70 · FP8-dynamic (compressed-tensors, per-channel FP8 weights, dynamic activations)

Candidate package · vLLM XPU

Qwen3.5 9B W4A16 with its own MTP head on one or two Intel Arc Pro B70

RedHatAI's W4A16 Qwen3.5-9B served by vLLM XPU on one B70 through the lab's R276 image, with the publisher's MTP head as a lossless speculative draft. vLLM's compressed-tensors path selects the lab's wNa16 INT4 kernel, whose fixed-K two-tier strategy is row-count invariant: without speculation this route is byte-exact against a single request at every concurrency through 64 users, where the FP8 route on the same model flips a near-tie token from 16 users up. It is also 28% faster without speculation and 15% faster with it. The kernel is not the only reduction on that path: the RMSNorm this route runs is itself row-count dependent (measured 2026-09-08), so the route is exact in the regimes measured rather than exact by construction.

113 tok/s measured · 1× B70 · W4A16 (compressed-tensors INT4 weights, FP16 activations)

Candidate package · llama.cpp SYCL

Qwen3.8 27B Q4_K_M on one Intel Arc Pro B70

Alibaba's Qwen3.8 27B on one Arc Pro B70 in 4-bit form with llama.cpp, no draft model. The lab's promoted one-card package, rebuilt from pinned source.

27.8 tok/s measured · 1× B70 · Q4_K_M

Candidate package · llama.cpp SYCL

Qwen3.8 27B Q4_K_M + MTP2 on one Intel Arc Pro B70

The one-card Qwen3.8 Q4_K_M target with its separately downloaded Q4_0 MTP draft at depth 2. Two fresh servers measured a 55.75% strict decode gain with target-exact output.

42.6 tok/s measured · 1× B70 · Q4_K_M target + Q4_0 MTP draft

Candidate package · llama.cpp SYCL

Qwen3.8 27B Q8_0 on one Intel Arc Pro B70

A quality-conservative one-card Qwen3.8 lane using Q8_0 weights, F16 KV, llama.cpp/SYCL, and no speculative decoding.

19.6 tok/s measured · 1× B70 · Q8_0

Candidate package · llama.cpp SYCL

Qwen3.8 27B Q8_0 + MTP2 on one Intel Arc Pro B70

The quality-conservative Q8_0 target with a separately downloaded Q4_0 MTP draft at depth 2. Two fresh servers measured 37.062 tok/s with all 24 speculative outputs exact to a matched MTP0 control.

37.1 tok/s measured · 1× B70 · Q8_0 target + Q4_0 MTP draft

Candidate package · vLLM XPU

Qwen3.8 27B AutoRound INT4, fixed-K batch-invariant profile, on two Intel Arc Pro B70 cards

devan-carlin's AutoRound INT4 tensors served through vLLM's plain-GPTQ oneDNN W4A16 path with a rebuilt kernel library that pins a two-tier fixed-K GEMM strategy, FP16 linears in 32-row pieces, single-split attention and size-independent Inductor reductions, on the FP8 lane's whole-graph deterministic stack. Single-request output is repeat-exact and speculative decoding is lossless against the MTP0 oracle at every depth measured; MTP0 output is byte-identical to a single request through 64 concurrent users (a near-tie prompt can differ in some runs), speculative depths through 16.

strict headline pending · 2× B70 · AutoRound INT4 W4A16 (group 128, symmetric)

Candidate package · llama.cpp SYCL

Qwen3.8 27B 256K + vision + MTP draft on one Intel Arc Pro B70

Alibaba's Qwen3.8 27B on one Arc Pro B70 with its full 262K-token context, image input, and draft head all resident at once. A capability-first fit, not a speed record.

26.7 tok/s measured · 1× B70 · UD-Q5_K_S (shipped) / UD-Q4_K_XL (alternative)

Candidate package · vLLM XPU

Qwen3.8 27B official FP8 on two Intel Arc Pro B70 cards

Alibaba's Qwen3.8 27B official block-FP8 weights on two Arc Pro B70 cards. The R187 profile (the R156 row-invariant W8A16 kernel and mixed-step GDN split, served with one whole-graph torch.compile instead of vLLM's piecewise split) is clean-boot-qualified at 54.935 tok/s MTP1 (FP16 target verifier, draft-only INT4 head), 70.142 tok/s MTP depth 2, 79.183 tok/s MTP depth 3, 82.396 tok/s MTP depth 4, 86.182 tok/s MTP depth 5, and 33.097 tok/s MTP0; every pair 12/12 against a same-configuration MTP0 oracle, repeat-exact at 224-300-token prompts, MTP0 output-identical to a single request through 64 concurrent users, MTP1 and depths 3-5 through 16, depth 2 through 4. On the piecewise compile MTP depth 2 emitted a phantom first token on one request in 64; on the whole-graph compile no pass has shown it (R182-R193, 2026-09-03). The cause is an unfixed upstream vLLM defect that also occurs on the unmodified image (R192/R194), so this is a configuration that avoids it on our deterministic build, not a fix; no patch, no image rebuild. A prebuilt copy of the exact image is on GitHub Container Registry (ghcr.io/steveseguin/vllm-openai-xpu-qwen38-fp8, digest sha256:173660ec…, equal to the image id the launchers verify); the source build in the guide remains the authoritative route.

86.2 tok/s measured · 2× B70 · FP8

Candidate package · llama.cpp SYCL

Qwen3.8 27B Q4_K_M + MTP2 on two Intel Arc Pro B70 cards

Two-B70 Qwen3.8 27B Q4_K_M with target-exact Q4_0 MTP2.

64.2 tok/s measured · 2× B70 · Q4_K_M target + Q4_0 MTP draft / F16 KV

Candidate package · llama.cpp SYCL

Qwen3.8 27B Q4_K_M on two Intel Arc Pro B70 cards

Alibaba's Qwen3.8 27B on two Arc Pro B70 cards in 4-bit form with llama.cpp, no draft model. Roughly 1.8x the one-card speed using the lab's two-card kernel stack.

49.7 tok/s measured · 2× B70 · Q4_K_M / F16 KV

Candidate package · llama.cpp SYCL

Qwen3.8 27B Q8_0 on two Intel Arc Pro B70 cards

Alibaba's Qwen3.8 27B on two Arc Pro B70 cards in 8-bit form with llama.cpp - the quality-conservative choice when 4-bit is not enough.

36.7 tok/s measured · 2× B70 · Q8_0 / F16 KV

Candidate package · vLLM XPU

Qwen3.8 Flash-Next FP8 with lossless MTP1 on four Intel Arc Pro B70 cards

Qwen's 125B-A6B hybrid-attention MoE, served from its official FP8 weights across four Arc Pro B70 cards with the expert and n-gram tables that do not fit placed in host memory. One speculative token per step with every output identical to the no-speculation line. A replay of the lab's certified run.

27.0 tok/s measured · 4× B70 · FP8 block-128 weights / BF16 KV

Candidate package · vLLM XPU

Qwen3.8 Flash-Next FP8 with lossless MTP1 and never-routed experts host-placed on four Intel Arc Pro B70 cards

Qwen's 125B-A6B hybrid-attention MoE, served from its official FP8 weights across four Arc Pro B70 cards. The n-gram table and embeddings live in host memory, every hot expert stays on the cards, and the experts a routing census never selects are parked in host memory behind a per-expert table in the MoE kernel. One speculative token per step with every output identical to the no-speculation line. A replay of the lab's certified run.

31.9 tok/s measured · 4× B70 · FP8 block-128 weights / BF16 KV

Candidate package · vLLM XPU

Qwen3.8 Flash-Next FP8 with lossless MTP1, never-routed experts host-placed and both reference Triton kernels restored, on four Intel Arc Pro B70 cards

Qwen's 125B-A6B hybrid-attention MoE, served from its official FP8 weights across four Arc Pro B70 cards. The n-gram table and embeddings live in host memory, every hot expert stays on the cards, the experts a routing census never selects are parked in host memory behind a per-expert table in the MoE kernel, and the two Triton kernels the XPU port had replaced, the hyper-connection glue and the QSA pre-indexer, are restored to the model's own reference implementations. One speculative token per step, lossless within the lineage. The outputs are a new authority whose difference is at depth: exact-2K coincides with the certified stream and exact-4K does not, while the quality profile is preserved byte for byte. A replay of the lab's certified run.

37.8 tok/s measured · 4× B70 · FP8 block-128 weights / BF16 KV

Candidate package · vLLM XPU

Qwen3.8 Flash-Next FP8 with lossless MTP1, never-routed experts host-placed and the Triton hyper-connection glue on four Intel Arc Pro B70 cards

Qwen's 125B-A6B hybrid-attention MoE, served from its official FP8 weights across four Arc Pro B70 cards. The n-gram table and embeddings live in host memory, every hot expert stays on the cards, the experts a routing census never selects are parked in host memory behind a per-expert table in the MoE kernel, and the model's own Triton hyper-connection glue kernels run on XPU instead of torch fallbacks. One speculative token per step with every output identical to the no-speculation line. The outputs are a new deterministic authority (the Triton glue rounds at the last bf16 bit), reproduced across five servers with the quality profile of the certified rows. A replay of the lab's certified run.

37.0 tok/s measured · 4× B70 · FP8 block-128 weights / BF16 KV

Candidate package · vLLM XPU

Qwen3.8 Flash-Next FP8 with no speculation, never-routed experts host-placed, both reference Triton kernels restored and the W13 MoE tile at the base width, on four Intel Arc Pro B70 cards

Qwen's 125B-A6B hybrid-attention MoE, served from its official FP8 weights across four Arc Pro B70 cards, with no speculative decoding. The n-gram table and embeddings live in host memory, every hot expert stays on the cards, the experts a routing census never selects are parked in host memory behind a per-expert table in the MoE kernel, and the two Triton kernels the XPU port had replaced are restored to the model's own reference implementations. The record over the previous one is a removal: a per-phase MoE tile that had been adopted bundled with an unrelated change, and never isolated, cost 2.1% once it was measured alone. One line of a tuned configuration file, no source change, outputs bit-identical across three servers.

34.5 tok/s measured · 4× B70 · FP8 block-128 weights / BF16 KV