Learn · concept guides

The levers that set your tokens per second

Plain-language guides to the variables that actually move local-LLM speed and capacity. Each one explains what the knob is, shows what we measured when we turned it, tells you what to pick, and warns you where it bites. Numbers are labelled by where they came from.

How to read these

Every figure carries a source badge: Lab-measured we ran it on our own B70s; Community reported by a contributor, checked for consistency but not re-run here; Spec / vendor a published number we have not benchmarked. We never present a spec figure as a measured result.

Speed comes from a stack of choices — the model, how it's quantized, how the KV cache is stored, how long your context is, which decoding trick you use, which runtime, and which card. Pull the wrong lever and you can halve your rate without noticing. Start anywhere; they cross-link.

Pick your setup

Tune for speed

What to be aware of

These guides describe our lab's Intel Arc Pro B70 lane and the community reports around it. The directions (bigger KV precision costs more at long context, MTP2 is a common knee — but sweep it; some quant/runtime combinations regress, prefill tracks clock) generalize well; the exact numbers are specific to our tuned build and hardware. Re-measure on your own machine before treating any figure as a guarantee.