Candidate package · llama.cpp SYCL · 1× Intel Arc Pro B70

LFM2.5 2.6B Q8_0 on one Intel Arc Pro B70

Liquid AI's compact LFM2.5 2.6B on one Arc Pro B70 in 8-bit form with stock, unpatched llama.cpp - the simplest single-command recipe on the site.

What we measured Lab-measured

132tok/s strict class-balanced decode

Median of two fresh-server class-balanced medians over the complete 12-prompt/six-class, 512-cap, cache-zero native HTTP suite; one B70, TP1, MTP0, reasoning off, F16 KV, 8K configured context. Both objective-canary batteries passed and complete token arrays matched 12/12 across servers. proof file

cache-zero = a fresh start, nothing pre-computed · MTP = multi-token prediction, a small draft the main model verifies

Model
LFM2.5 2.6B
Publisher
Liquid AI
Checkpoint
LiquidAI/LFM2.5-2.6B-GGUF
Compression
Q8_0
Software
llama.cpp SYCL
Cards
1× Intel Arc Pro B70 32 GiB
Model weight bytes
2.9 GB
Operating systems
Linux
Delivery
native
Good for
general chat starter
Published
2026-08-27
Clean-host replay
not yet

Still missing before this becomes an install guide

  • tested clean-host platform installation
  • output-qualified HTTP concurrency
  • beginner recovery flow

What to know

  • Emits untagged reasoning prose before final answers regardless of the reasoning flag; answers correct but verbose.
  • The quality gate establishes objective correctness and exact fresh-server repeatability for this identity; it is not a cross-model capability comparison.
  • Clean-host beginner flow not yet executed.

Measured performance profiles Lab-measured

Raw decode over existing context depth Lab-measured

13112712010898.189.92K4K8K16K24K32KExisting context depth before tg128 · y: tok/s
llama-bench raw engine rates (pp2048/tg128, fa on, 5 reps); see depth-sweep.svg and sweep JSON in the guide directory. The directly measured zero-depth point remains in the linked raw evidence; no missing depth is interpolated. evidence

Raw pp2048 over existing context depth Lab-measured

4,6994,5854,4003,7523,6982,8252K4K8K16K24K32KExisting context depth before pp2048 · y: tok/s
llama-bench raw engine rates (pp2048/tg128, fa on, 5 reps); see depth-sweep.svg and sweep JSON in the guide directory. The directly measured zero-depth point remains in the linked raw evidence; no missing depth is interpolated. evidence

Many people at once Not published

No qualified multi-user aggregate profile is published for this exact package. Diagnostic or unsupported boundaries may still appear under “What to know” or in the full guide. Nothing is interpolated or promoted from a different model, quantization, runtime, or card count; the projection below remains clearly labeled as projected.

How much faster could this get? Projected — not measured

The ML Bottleneck physics engine projects what stock software, a tuned run, and the physical ceiling look like for this exact model, compression, card count, and software. The grade is optimization headroom against the tuned-run target, not model quality.