Nemotron Cascade 2 · 30B-A3B

NVIDIA's Nemotron Cascade 2, a 30B hybrid attention/Mamba mixture-of-experts with about 3B active per word. Runs on one Arc Pro B70 in 4-bit form.

50.9tok/s
measured
≈ 38 words a secondnemotron-cascade-2-30b-a3b-931b595 · Q4_K_M · llama.cpp SYCL fdb1db877 · 1× B70 · Evidence D of A–D
Nemotron-Cascade-2 30B-A3B · measured GGUF weights
Read the lab reportAll packets and recipesNo step-by-step install guide is published for this model yet.
B70 fit
Runs on 1 card (measured)
Quality evidence
Rapid snapshot only
Measured results
1

Packets and recipes

The deployment variants of this family, at every maturity.

What has been classified

Dense scorecards summarize every declared combination; measured slices retain exact evidence links.

quant × TP

Exact measured revision and rapid runtime; TP2/4 are unmeasured gaps. Fixed: revision=nemotron-cascade-2-30b-a3b-931b595 · runtime=llama.cpp SYCL fdb1db877 · mtp=0.

2 untested combinations
TP2·Q4_K_MTP4·Q4_K_M
Full matrix and exact selectors

Measured results

Every number links to its proof.

family data

Measured speed (quick lab check)

Q4_K_M · one B70 · f16 KV · ctx=2048 configured; this one point is not a context curve

Q4_K_M · tensor parallel cards 150.9tok/sDecode449.16msTTFTstrict rapid snapshot · promoted standalone row · evidence
Q4_K_Mmissing
evidence 1
Fine print

Nemotron Cascade hybrid attention/Mamba MoE, 52 layers. one measured Cascade 2 GGUF revision. Measurements, artifact hashes, outputs, quality decisions, and speed stay pinned to their exact recorded identity.

Many people at once Not measured

Multi-user (aggregate) throughput has not been measured for this family. The lab’s one measured sweep so far is in the multi-user report; the projection block below includes a projected users curve where a projection exists.

How much faster could this get? No projection

No projection is shown: this model is not in the ML Bottleneck catalog yet. The measured numbers above stand on their own.