Candidate package · llama.cpp SYCL · 1× Intel Arc Pro B70

Ornith 1.5 35B-A3B Q4_K_M on one Intel Arc Pro B70

Ornith AI's 35B mixture-of-experts (about 3B active per word) on one Arc Pro B70 in 4-bit form - the fastest single-card result this lab has measured, using the lab's tuned kernel stack.

Public headline Pending

The 131.460231 tok/s two-server observation and matched patch A/B evidence remain valid scoped measurements, but the runtime produced 0/12 identical complete natural-response hashes across fresh stock servers. Keep the mechanism and context evidence; require a stable, registered cross-server oracle before assigning a strict package headline. Diagnostic measurements remain in the guide and evidence, but none is presented as the package headline.

Model
Ornith 1.5 35B-A3B
Publisher
Ornith AI
Checkpoint
ornith-ai/Ornith-1.5-35B-A3B-GGUF
Compression
Q4_K_M
Software
llama.cpp SYCL
Cards
1× Intel Arc Pro B70 32 GiB
Model weight bytes
21.7 GB
Operating systems
Linux
Delivery
native
Good for
general coding chat
Published
2026-08-22
Clean-host replay
not yet

Still missing before this becomes an install guide

  • tested clean-host platform installation
  • beginner recovery flow

What to know

  • MoE rate reflects ~3B active parameters; not comparable to dense 35B expectations.
  • Canary battery is objective self-consistency only; no cross-model oracle.
  • Stock two-card layer split measured 102.01/102.20 tok/s versus 104.84/104.81 on one card (-2.6%): use one card for single-stream decode.
  • Fresh stock servers matched 0/12 complete realistic-suite hashes across processes on this runtime. A realistic same-process repeat also produced four hashes across eight requests; use same-frozen-binary door-off/on exactness and activation counts for patch validation, while the short exact-answer 8x canary remains a narrower check.
  • The aggregate concurrency profile is a raw-engine continuous-batching measurement, not an HTTP serving users/sec result. The optional multi-row research patch improves the four-sequence aggregate rate by a confirmed 2.21% but is not enabled by the single-user package launch.
  • Stage the GGUF on local or sufficiently fast direct-attached storage; optimized first-token setup rereads roughly 20.0 GB of expert tensors, making a 100 Mb/s NFS mmap unsuitable for cold benchmarking.

Measured performance profiles Lab-measured

Current twelve-feature raw decode over existing context depth Lab-measured

14213713312711610799.602K4K8K16K24K32KExisting context depth before tg128 · y: tok/s
llama-bench raw engine rates using the exact twelve-feature source patch and accepted copy-offload setting, graph off, pp2048/tg128, flash attention on, F16 KV, and 5 repetitions at every displayed depth. No point is scaled and no missing depth is interpolated. evidence

Current twelve-feature raw pp2048 over existing context depth Lab-measured

1,4221,3441,3381,3011,2391,2131,11602K4K8K16K24K32KExisting context depth before pp2048 · y: tok/s
llama-bench raw engine rates using the exact twelve-feature source patch and accepted copy-offload setting, graph off, pp2048/tg128, flash attention on, F16 KV, and 5 repetitions at every displayed depth. No point is scaled and no missing depth is interpolated. evidence

Many people at once Lab-measured

Accepted-stack aggregate decode over concurrent sequences Lab-measured

98.010311914616321712481632Concurrent engine sequences · log2 spacing · y: tok/s
Direct llama-batched-bench raw-engine continuous batching with independent pp1024 prompts, interleaved tg256 decode, c65536, flash attention on, F16 KV, one B70, and the exact accepted twelve-feature stack. This excludes HTTP, JSON, queueing, and server-scheduler overhead. Per-user values are arithmetic aggregate/users. No point is scaled, interpolated, or extrapolated. evidence

Optimization grade Pending

No optimization grade is calculated because this package has no promoted measured headline. Diagnostic or scoped measurements never seed a headline projection.