Ornith 1.5 · dense 9B + MoE 35B-A3B

Ornith AI's Ornith 1.5 line: a dense 9B and a 35B mixture-of-experts (about 3B active per word), both Qwen-derived. The 35B is the fastest single-card model this lab has measured; the 9B is a smaller stock deployment.

Ornith 1.5 · publisher release line containing architecture-distinct siblings
B70 fit
High
Quality evidence
Strong scoped deployment evidence
Interest
123.24K
Measured results
9

Packets and recipes

The deployment variants of this family, at every maturity.

What has been classified

Dense scorecards summarize every declared combination; measured slices retain exact evidence links.

serving profile × cards

Target-only · graph off · F16 KV · standard 8K serving profile. Rows keep stock and patched identities separate; a blank topology is never inferred from another row. Fixed: mtp=0 · graph=off · kv=f16 · configured_max_context_tokens=8192.

  • ✓ Measured1, 9B Q8_0 · stock 1·9B Q8_0 · stock49.59 tok/sguide · evidence
  • ✓ Measured1, 35B-A3B Q4_K_M · current optimized 1·35B-A3B Q4_K_M · current optimized132.47 tok/sguide · evidence
  • ✓ Measured1, 35B-A3B Q4_K_M · stock 1·35B-A3B Q4_K_M · stock104.84 tok/sevidence
  • ✓ Measured2, 35B-A3B Q4_K_M · stock 2·35B-A3B Q4_K_M · stock102.2 tok/sevidence

Codes in the rows: D = decode tok/s · P = prefill tok/s · T = ms to first token · AR = share of drafted tokens accepted · A = combined tok/s.

5 untested combinations
2·9B Q8_0 · stock4·9B Q8_0 · stock2·35B-A3B Q4_K_M · current optimized4·35B-A3B Q4_K_M · current optimized4·35B-A3B Q4_K_M · stock
Full matrix and exact selectors

9B Q8_0 context × cards

Stock 9B Q8_0 · target-only · graph off · F16 KV · raw pp2048/tg128 · five repetitions per marker. Fixed: revision=ornith-1.5-9b · variant=Q8_0 dense · runtime=llama.cpp SYCL 9fee29e · profile_id=ornith-9b-q8-stock-context · split_mode=single device · mtp=0 · graph=off · kv=f16.

  • ✓ Measured1, 0 1·0D50.29 · P3,184evidence
  • ✓ Measured1, 2048 1·2048D49.34 · P1,623evidence
  • ✓ Measured1, 4096 1·4096D48.55 · P1,601evidence
  • ✓ Measured1, 8192 1·8192D47.04 · P1,569evidence
  • ✓ Measured1, 16384 1·16384D44.35 · P1,489evidence
  • ✓ Measured1, 24576 1·24576D41.96 · P1,449evidence
  • ✓ Measured1, 32768 1·32768D39.84 · P1,314evidence

Codes in the rows: D = decode tok/s · P = prefill tok/s · T = ms to first token · AR = share of drafted tokens accepted · A = combined tok/s.

14 untested combinations
2·04·02·20484·20482·40964·40962·81924·81922·163844·163842·245764·245762·327684·32768
Full matrix and exact selectors

35B-A3B current context × cards

Current twelve-feature Q4_K_M overlay plus accepted copy-offload setting · target-only · graph off · F16 KV · raw pp2048/tg128 · five repetitions per marker. Fixed: revision=ornith-1.5-35b-a3b · variant=Q4_K_M MoE twelve-feature stack plus copy-offload setting · runtime=llama.cpp SYCL 9fee29e plus lab patch · profile_id=ornith-35b-q4-twelve-feature-context · split_mode=single device · mtp=0 · graph=off · kv=f16.

  • ✓ Measured1, 0 1·0D141.92 · P1,422evidence
  • ✓ Measured1, 2048 1·2048D136.85 · P1,344evidence
  • ✓ Measured1, 4096 1·4096D133.33 · P1,338evidence
  • ✓ Measured1, 8192 1·8192D126.83 · P1,301evidence
  • ✓ Measured1, 16384 1·16384D116.27 · P1,239evidence
  • ✓ Measured1, 24576 1·24576D106.97 · P1,213evidence
  • ✓ Measured1, 32768 1·32768D99.61 · P1,116evidence

Codes in the rows: D = decode tok/s · P = prefill tok/s · T = ms to first token · AR = share of drafted tokens accepted · A = combined tok/s.

14 untested combinations
2·04·02·20484·20482·40964·40962·81924·81922·163844·163842·245764·245762·327684·32768
Full matrix and exact selectors

35B concurrent sequences × cards

Accepted Ornith 35B Q4_K_M stack · one-card raw-engine continuous batching · independent 1,024-token prompts · 256 generated tokens per sequence · graph off · F16 KV. Fixed: revision=ornith-1.5-35b-a3b · variant=Q4_K_M MoE accepted twelve-feature stack · runtime=llama.cpp SYCL 9fee29e plus accepted lab patch · profile_id=ornith-35b-q4-accepted-stack-batched · measurement_class=raw-engine-continuous-batching · split_mode=single device · mtp=0 · graph=off · kv=f16 · configured_max_context_tokens=65536.

  • ✓ Measured1, c1 1·c1D98.02 · A98.02evidence
  • ✓ Measured1, c2 1·c2D51.34 · A102.67evidence
  • ✓ Measured1, c4 1·c4D29.72 · A118.88evidence
  • ✓ Measured1, c8 1·c8D18.29 · A146.28evidence
  • ✓ Measured1, c16 1·c16D10.16 · A162.5evidence
  • ✓ Measured1, c32 1·c32D6.77 · A216.51evidence

Codes in the rows: D = decode tok/s · P = prefill tok/s · T = ms to first token · AR = share of drafted tokens accepted · A = combined tok/s.

12 untested combinations
2·c14·c12·c24·c22·c44·c42·c84·c82·c164·c162·c324·c32
Full matrix and exact selectors

Measured results

Every number links to its proof.

family data

Speed as the conversation grows

Separate dense Q8_0 and MoE Q4_K_M architectures on one B70; shared axes do not imply transferable speed

039.779.5119.2158.902K4K8K16K24K32Kactive context tokens · tok/s9B dense · Q8_0 · Decode (decode_tok_s) · active context tokens=0.0 · value=50.291618 tok/s9B dense · Q8_0 · Decode (decode_tok_s) · active context tokens=2048.0 · value=49.338288 tok/s9B dense · Q8_0 · Decode (decode_tok_s) · active context tokens=4096.0 · value=48.554943 tok/s9B dense · Q8_0 · Decode (decode_tok_s) · active context tokens=8192.0 · value=47.044737 tok/s9B dense · Q8_0 · Decode (decode_tok_s) · active context tokens=16384.0 · value=44.349773 tok/s9B dense · Q8_0 · Decode (decode_tok_s) · active context tokens=24576.0 · value=41.957791 tok/s9B dense · Q8_0 · Decode (decode_tok_s) · active context tokens=32768.0 · value=39.83848 tok/s35B-A3B MoE · Q4_K_M · current 12-feature · Decode (decode_tok_s) · active context tokens=0.0 · value=141.917514 tok/s35B-A3B MoE · Q4_K_M · current 12-feature · Decode (decode_tok_s) · active context tokens=2048.0 · value=136.848608 tok/s35B-A3B MoE · Q4_K_M · current 12-feature · Decode (decode_tok_s) · active context tokens=4096.0 · value=133.329784 tok/s35B-A3B MoE · Q4_K_M · current 12-feature · Decode (decode_tok_s) · active context tokens=8192.0 · value=126.829116 tok/s35B-A3B MoE · Q4_K_M · current 12-feature · Decode (decode_tok_s) · active context tokens=16384.0 · value=116.267169 tok/s35B-A3B MoE · Q4_K_M · current 12-feature · Decode (decode_tok_s) · active context tokens=24576.0 · value=106.967195 tok/s35B-A3B MoE · Q4_K_M · current 12-feature · Decode (decode_tok_s) · active context tokens=32768.0 · value=99.614237 tok/s
9B dense · Q8_035B-A3B MoE · Q4_K_M · current 12-feature
9B dense · Q8_0 0: 50.29 · 2K: 49.34 · 4K: 48.55 · 8K: 47.04 · 16K: 44.35 · 24K: 41.96 · 32K: 39.84
35B-A3B MoE · Q4_K_M · current 12-feature 0: 141.92 · 2K: 136.85 · 4K: 133.33 · 8K: 126.83 · 16K: 116.27 · 24K: 106.97 · 32K: 99.61
evidence 1 · evidence 2

Tuned 35B speed as the conversation grows

Historical eleven-feature profile and current twelve-feature plus copy-offload profile; every marker is measured independently

039.779.5119.2158.902K4K8K16K24K32Kactive context tokens · tok/shistorical 11-feature · Decode (decode_tok_s) · active context tokens=0.0 · value=138.97752 tok/shistorical 11-feature · Decode (decode_tok_s) · active context tokens=2048.0 · value=133.765934 tok/shistorical 11-feature · Decode (decode_tok_s) · active context tokens=4096.0 · value=130.620489 tok/shistorical 11-feature · Decode (decode_tok_s) · active context tokens=8192.0 · value=124.209778 tok/shistorical 11-feature · Decode (decode_tok_s) · active context tokens=16384.0 · value=113.31086 tok/shistorical 11-feature · Decode (decode_tok_s) · active context tokens=24576.0 · value=104.315441 tok/shistorical 11-feature · Decode (decode_tok_s) · active context tokens=32768.0 · value=96.995532 tok/scurrent 12-feature + copy-offload · Decode (decode_tok_s) · active context tokens=0.0 · value=141.917514 tok/scurrent 12-feature + copy-offload · Decode (decode_tok_s) · active context tokens=2048.0 · value=136.848608 tok/scurrent 12-feature + copy-offload · Decode (decode_tok_s) · active context tokens=4096.0 · value=133.329784 tok/scurrent 12-feature + copy-offload · Decode (decode_tok_s) · active context tokens=8192.0 · value=126.829116 tok/scurrent 12-feature + copy-offload · Decode (decode_tok_s) · active context tokens=16384.0 · value=116.267169 tok/scurrent 12-feature + copy-offload · Decode (decode_tok_s) · active context tokens=24576.0 · value=106.967195 tok/scurrent 12-feature + copy-offload · Decode (decode_tok_s) · active context tokens=32768.0 · value=99.614237 tok/s
historical 11-featurecurrent 12-feature + copy-offload
historical 11-feature 0: 138.98 · 2K: 133.77 · 4K: 130.62 · 8K: 124.21 · 16K: 113.31 · 24K: 104.32 · 32K: 97
current 12-feature + copy-offload 0: 141.92 · 2K: 136.85 · 4K: 133.33 · 8K: 126.83 · 16K: 116.27 · 24K: 106.97 · 32K: 99.61
evidence 1 · evidence 2

35B: stock software by card count

Q4_K_M stock serving comparison; the optimized one-card packet is a separate patched identity

stock layer placement · B70 cards 1104.81–104.84tok/sDecodelab measurement; quality scope in evidence · evidence
stock layer placement · B70 cards 2102.01–102.2tok/sDecodelab measurement; quality scope in evidence · evidence
stock layer placement
evidence 1

35B: many users at once

Raw llama-batched-bench engine throughput; combined and per-sequence rates are both measured, and sequence count is not relabeled as HTTP users

060.6121.2181.9242.512481632concurrent engine sequences · tok/saccepted 12-feature stack · Combined (aggregate_tok_s) · concurrent engine sequences=1.0 · value=98.024651 tok/saccepted 12-feature stack · Combined (aggregate_tok_s) · concurrent engine sequences=2.0 · value=102.671738 tok/saccepted 12-feature stack · Combined (aggregate_tok_s) · concurrent engine sequences=4.0 · value=118.882942 tok/saccepted 12-feature stack · Combined (aggregate_tok_s) · concurrent engine sequences=8.0 · value=146.283051 tok/saccepted 12-feature stack · Combined (aggregate_tok_s) · concurrent engine sequences=16.0 · value=162.503937 tok/saccepted 12-feature stack · Combined (aggregate_tok_s) · concurrent engine sequences=32.0 · value=216.513077 tok/s
accepted 12-feature stack
accepted 12-feature stack 1: 98.02 · 2: 102.67 · 4: 118.88 · 8: 146.28 · 16: 162.5 · 32: 216.51
evidence 1
Fine print

Qwen3.5-derived sibling architectures. Publisher/version grouping only: the dense qwen35 9B and qwen35moe 35B-A3B variants are architecture-distinct, so no model-specific kernel, speed, output, or quality transfer is implied. Measurements, artifact hashes, outputs, quality decisions, and speed stay pinned to their exact recorded identity.

1 combination the lab has stopped pursuing: Both artifacts declare a 262,144-token native window, but these two packet profiles have measured active-context evidence only through 32,768 tokens. evidence Exact selectors are in the family data.

Many people at once Lab-measured

32 simultaneous users share 217 combined tok/s on 1 card (Q4_K_M MoE accepted twelve-feature stack); the full curve is under Measured results and in the multi-user report.

How much faster could this get? No projection

No projection is shown: the family has no curated headline measurement yet. The measured numbers above stand on their own.