NVIDIA's Nemotron 3.5 Lightning, a 30B hybrid Mamba mixture-of-experts with about 3B active per word. A fast general assistant that runs on one Arc Pro B70 with stock software.
nemotron_h_moe, 53 layers, 2688 hidden. Only one model artifact and quantization are measured; no cross-quantization, cross-runtime, or later-weight transfer is claimed. Measurements, artifact hashes, outputs, quality decisions, and speed stay pinned to their exact recorded identity.
1 combination the lab has stopped pursuing: The artifact declares a 1,048,576-token native window, but this package has measured context-depth evidence only through 32,768 tokens. evidence Exact selectors are in the family data.
Many people at once Not measured
Multi-user (aggregate) throughput has not been measured for this family. The lab’s one measured sweep so far is in the multi-user report; the projection block below includes a projected users curve where a projection exists.
How much faster could this get? No projection
No projection is shown: the family has no curated headline measurement yet. The measured numbers above stand on their own.