LFM2.5 2.6B

Liquid AI's LFM2.5, a compact 2.6B model with a hybrid convolution-and-attention design. Very fast on one Arc Pro B70 and a good fit for lightweight assistants and edge-style workloads.

132.14tok/s
measured
≈ 99 words a secondlfm2.5-2.6b · Q8_0 · llama.cpp SYCL · 1× B70 · B70-verified
LFM2.5 2.6B · current measured weights
B70 fit
High
Quality evidence
Scoped deployment evidence
Measured results
2

Packets and recipes

The deployment variants of this family, at every maturity.

What has been classified

Dense scorecards summarize every declared combination; measured slices retain exact evidence links.

Q8_0 · target only

LFM2.5 2.6B Q8_0 · llama.cpp SYCL · f16 KV · MTP is not used. TP1 has serving and raw-depth evidence; TP2/TP4 have not been measured.

  • ✓ Measured1 card, no speculative decoding TP1·MTP0132.16 tok/sevidence
2 untested combinations
TP2·MTP0TP4·MTP0
Full matrix and exact selectors

Measured results

Every number links to its proof.

family data

Repeat runs

Q8_0 · TP1 · MTP0 · 8K configured · f16 KV · full 12-prompt/six-class cache-zero suite

fresh-server repeats · tensor parallel cards 1132.11–132.16tok/sDecodelab measurement; quality scope in evidence · evidence
fresh-server repeats
evidence 1

Speed as the conversation grows

Q8_0 · TP1 · f16 KV · raw llama-bench pp2048/tg128; raw engine rates are not serving medians

037.975.7113.6151.402K4K8K16K24K32Kexisting context tokens · tok/sQ8_0 · Decode (decode_tok_s) · existing context tokens=0.0 · value=135.198135 tok/sQ8_0 · Decode (decode_tok_s) · existing context tokens=2048.0 · value=131.259551 tok/sQ8_0 · Decode (decode_tok_s) · existing context tokens=4096.0 · value=127.211224 tok/sQ8_0 · Decode (decode_tok_s) · existing context tokens=8192.0 · value=120.195857 tok/sQ8_0 · Decode (decode_tok_s) · existing context tokens=16384.0 · value=107.980065 tok/sQ8_0 · Decode (decode_tok_s) · existing context tokens=24576.0 · value=98.138413 tok/sQ8_0 · Decode (decode_tok_s) · existing context tokens=32768.0 · value=89.93812 tok/s
Q8_0
Q8_0 0: 135.2 · 2K: 131.26 · 4K: 127.21 · 8K: 120.2 · 16K: 107.98 · 24K: 98.14 · 32K: 89.94
evidence 1
Fine print

LFM2ForCausalLM, 30 layers, 2048 hidden. one measured weight revision; quantization and runtime paths require independent gates. Measurements, artifact hashes, outputs, quality decisions, and speed stay pinned to their exact recorded identity.

Many people at once Not measured

Multi-user (aggregate) throughput has not been measured for this family. The lab’s one measured sweep so far is in the multi-user report; the projection block below includes a projected users curve where a projection exists.

How much faster could this get? Projected — not measured

The ML Bottleneck physics engine projects a tuned-run target and the physical ceiling for the headline setup (lfm2.5-2.6b · Q8_0 · llama.cpp SYCL · 1× B70). The grade is optimization headroom against the tuned-run target, not model quality.