Liquid AI's LFM2.5, a compact 2.6B model with a hybrid convolution-and-attention design. Very fast on one Arc Pro B70 and a good fit for lightweight assistants and edge-style workloads.
LFM2ForCausalLM, 30 layers, 2048 hidden. one measured weight revision; quantization and runtime paths require independent gates. Measurements, artifact hashes, outputs, quality decisions, and speed stay pinned to their exact recorded identity.
Many people at once Not measured
Multi-user (aggregate) throughput has not been measured for this family. The lab’s one measured sweep so far is in the multi-user report; the projection block below includes a projected users curve where a projection exists.
How much faster could this get? Projected — not measured
The ML Bottleneck physics engine projects a tuned-run target and the physical ceiling for the headline setup (lfm2.5-2.6b · Q8_0 · llama.cpp SYCL · 1× B70). The grade is optimization headroom against the tuned-run target, not model quality.