Candidate package · llama.cpp SYCL · 1× Intel Arc Pro B70

Nemotron 3.5 Lightning 30B-A3B UD-Q4_K_M on one Intel Arc Pro B70

NVIDIA's Nemotron 3.5 Lightning, a 30B hybrid Mamba mixture-of-experts, on one Arc Pro B70 in 4-bit form with stock llama.cpp. Its speed barely drops as the conversation grows.

Public headline Pending

The 72.169452/72.035976 varied-suite observations and reasoning-off canary summary are preserved in the guide, but their raw operating-point and canary JSON files are not closed in this repository. Import and hash-bind those files, then replay the quality/determinism gate before assigning a strict headline. Diagnostic measurements remain in the guide and evidence, but none is presented as the package headline.

Model
Nemotron 3.5 Lightning 30B-A3B
Publisher
NVIDIA (GGUF by unsloth)
Checkpoint
unsloth/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF
Compression
UD-Q4_K_M
Software
llama.cpp SYCL
Cards
1× Intel Arc Pro B70 32 GiB
Model weight bytes
25.3 GB
Operating systems
Linux
Delivery
native
Good for
general long context chat
Published
2026-08-22
Clean-host replay
not yet

Still missing before this becomes an install guide

  • tested clean-host platform installation
  • context beyond 32K of the 1M native window

What to know

  • With reasoning ON, repeated identical requests are not hash-stable (thinking-channel sampling); deterministic with reasoning off - the recipe default.
  • MoE rate reflects ~3B active parameters.
  • 1M native context not exercised beyond 32K here.
  • Two-card layer split measured 69.45/69.49 tok/s vs 72.17/72.04 on one card (-3.7%): use one card for single-stream decode.

Measured performance profiles Lab-measured

Raw decode over existing context depth Lab-measured

72.471.870.768.666.664.62K4K8K16K24K32KExisting context depth before tg128 · y: tok/s
llama-bench raw engine rates (pp2048/tg128, fa on, 5 reps); only -11% decode from 0 to 32K depth. The directly measured zero-depth point remains in the linked raw evidence; no missing depth is interpolated. evidence

Raw pp2048 over existing context depth Lab-measured

1,1671,1531,1461,1061,1031,0192K4K8K16K24K32KExisting context depth before pp2048 · y: tok/s
llama-bench raw engine rates (pp2048/tg128, fa on, 5 reps); only -11% decode from 0 to 32K depth. The directly measured zero-depth point remains in the linked raw evidence; no missing depth is interpolated. evidence

Many people at once Not published

No qualified multi-user aggregate profile is published for this exact package. Diagnostic or unsupported boundaries may still appear under “What to know” or in the full guide. Nothing is interpolated or promoted from a different model, quantization, runtime, or card count.

Optimization grade Pending

No optimization grade is calculated because this package has no promoted measured headline. Diagnostic or scoped measurements never seed a headline projection.