Candidate package · vLLM XPU · 4× Intel Arc Pro B70

Laguna S 2.1 INT4 with DFlash on four Intel Arc Pro B70 cards

Poolside's Laguna S 2.1 coding assistant on four Arc Pro B70 cards in 4-bit form, with a draft model speeding up generation. A replay of the lab's record run, quality-checked.

What we measured Lab-measured

125tok/s conventional decode median

Conventional 99-interval median across the sealed 13-prompt, one-start, cache-zero record gate. proof file

cache-zero = a fresh start, nothing pre-computed · 99-interval median = the middle rate across 99 measured stretches of a long answer

Model
Laguna S 2.1 INT4 target plus INT4 DFlash draft
Publisher
Poolside
Checkpoint
poolside/Laguna-S-2.1-INT4
Compression
INT4 / BF16 KV
Software
vLLM XPU
Cards
4× Intel Arc Pro B70 32 GiB
Model weight bytes
Operating systems
Linux
Delivery
native
Good for
general coding speculative decoding research
Published
2026-08-22
Clean-host replay
not yet

Still missing before this becomes an install guide

  • portable runtime rebuild
  • tested platform installation
  • record-specific model acquisition helper
  • non-originating-host replay
  • beginner recovery flow
  • decode, prefill, and TTFT context sweep

Measured performance profiles Not published

No qualified structured context or depth profile is published for this package. Diagnostic evidence may still be linked under “What to know” or in the full guide; nothing is estimated in its place — the clearly labeled projection block below is the current best guess.

Many people at once Not published

No qualified multi-user aggregate profile is published for this exact package. Diagnostic or unsupported boundaries may still appear under “What to know” or in the full guide. Nothing is interpolated or promoted from a different model, quantization, runtime, or card count; the projection below remains clearly labeled as projected.

How much faster could this get? Projected — not measured

The ML Bottleneck physics engine projects what stock software, a tuned run, and the physical ceiling look like for this exact model, compression, card count, and software. The grade is optimization headroom against the tuned-run target, not model quality.