Candidate package · llama.cpp SYCL · 1× Intel Arc Pro B70

Ornith 1.5 9B Q8_0 on one Intel Arc Pro B70

Ornith AI's dense 9B on one Arc Pro B70 in 8-bit form with stock llama.cpp. A small, official, beginner-friendly deployment.

Public headline Withheld

Strict headline withheld after output-gate failure. Two complete fresh-server, cache-zero, 12-prompt/six-class attempts measured 49.593582 and 49.515869 tok/s and passed both objective-canary batteries, but complete token arrays matched only 8/12 across servers. Diagnostic measurements remain in the guide and evidence, but none is presented as the package headline.

Model
Ornith 1.5 9B
Publisher
Ornith AI
Checkpoint
ornith-ai/Ornith-1.5-9B-GGUF
Compression
Q8_0
Software
llama.cpp SYCL
Cards
1× Intel Arc Pro B70 32 GiB
Model weight bytes
9.5 GB
Operating systems
Linux
Delivery
native
Good for
general chat starter
Published
2026-08-27
Clean-host replay
not yet

Still missing before this becomes an install guide

  • tested clean-host platform installation
  • additional operating points beyond the 8K standard
  • beginner recovery flow

What to know

  • The 2026-08-27 strict pair passed both objective-canary batteries but matched only 8/12 complete natural-response token arrays across fresh servers; no general single-user headline is published.
  • Clean-host beginner flow not yet executed.
  • The measured decode rate requires local or sufficiently fast direct-attached model storage; a 100 Mb/s NFS-backed mmap reduced the matched audit-host diagnostic from 50.149 to 25.642 tok/s.

Measured performance profiles Lab-measured

Raw decode over existing context depth Lab-measured

49.348.647.044.342.039.82K4K8K16K24K32KExisting context depth before tg128 · y: tok/s
llama-bench raw engine rates (pp2048/tg128, fa on, 5 reps). The directly measured zero-depth point remains in the linked raw evidence; no missing depth is interpolated. evidence

Raw pp2048 over existing context depth Lab-measured

1,6231,6011,5691,4891,4491,3142K4K8K16K24K32KExisting context depth before pp2048 · y: tok/s
llama-bench raw engine rates (pp2048/tg128, fa on, 5 reps). The directly measured zero-depth point remains in the linked raw evidence; no missing depth is interpolated. evidence

Many people at once Not published

No qualified multi-user aggregate profile is published for this exact package. Diagnostic or unsupported boundaries may still appear under “What to know” or in the full guide. Nothing is interpolated or promoted from a different model, quantization, runtime, or card count.

Optimization grade Pending

No optimization grade is calculated because this package has no promoted measured headline. Diagnostic or scoped measurements never seed a headline projection.