Models·2 min read
By BitsMindsSource: NVIDIA Newsroom

NVIDIA Unveils Nemotron 3: Nano, Super, and Ultra Open Models for Agentic AI

NVIDIA debuted the Nemotron 3 family at GTC 2026, releasing Nano immediately and previewing Super (49B) and Ultra (253B) variants, along with three trillion tokens of pre-training data.

NVIDIA Unveils Nemotron 3: Nano, Super, and Ultra Open Models for Agentic AI
Share:

NVIDIA announced the Nemotron 3 family of open models at GTC 2026, positioning the lineup as the most efficient open models for building agentic AI applications. The family spans three tiers — Nano, Super, and Ultra — each tuned for different deployment contexts, from edge devices and multi-agent swarms to high-accuracy enterprise workloads running on multi-GPU clusters.

The Nemotron 3 Nano model launched immediately and delivers a headline benchmark: four times higher throughput than its Nemotron 2 predecessor, achieved through a hybrid mixture-of-experts architecture that dramatically reduces the compute required per token. For developers building multi-agent systems at scale, where thousands of inference calls may happen in parallel, this throughput improvement directly translates to lower latency and cost at runtime.

Nemotron 3 Super, at 49 billion parameters, and Nemotron 3 Ultra, at 253 billion parameters, are targeted for availability in the first half of 2026. Both are designed for applications requiring high reasoning accuracy — complex coding tasks, long-document analysis, and autonomous agent pipelines that need reliable, step-by-step problem decomposition. NVIDIA says the Super and Ultra models match or outperform much larger models from other providers on standard reasoning benchmarks.

In a notable move for the open-source community, NVIDIA is releasing not just the model weights but also the data used to train them: three trillion tokens of pre-training data and 18 million samples of post-training instruction data. This level of transparency is rare among frontier model releases and gives researchers and fine-tuners the ability to understand, audit, and extend the models in ways that are impossible with proprietary systems.

The release complements NVIDIA's broader push into physical AI and robotics, announced at the same GTC event alongside the GR00T N1.7 open model for humanoid robot training. Together, the announcements signal that NVIDIA sees open model releases as a strategic tool for expanding the AI software ecosystem around its hardware — a long-term bet that the more developers build with NVIDIA-origin models, the more they deploy on NVIDIA chips.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles

Many kinds of input, one Qwen context On a deep violet field, a filmstrip, an audio waveform, a landscape photograph and a text page curve toward a single luminous sphere bearing the official purple Qwen star. A label below reads “1M context”, representing text, images, audio and video sharing one million-token context window. Aa 1M context BITSMINDS.COM
Models

Qwen3.8-Omni-Flash Undercuts Gemini on Audio and Video

Hill Climb: GPT-6 Astra versus Claude Fable 5.1 A teal desert buggy climbs a sandstone ridge on the left. An orange jeep climbs a green hill on the right, under an arc of gold coins. The two landscapes meet at a diagonal divide. 7 GPT-6 ASTRA CLAUDE FABLE 5.1 VS HILL CLIMB BITSMINDS.COM
Models

GPT-6 Astra vs Claude Fable 5.1: Hill Climb

Gemini 3.8 Live: thinking while the conversation continues An editorial illustration in a dark blue and violet room. A carefully drawn studio microphone and a tilted smartphone flank a translucent speech bubble carrying the multicoloured Gemini star. A continuous luminous audio waveform travels between them. Above the bubble, a separate arc connects small search, reasoning and completion symbols, representing background work continuing during a spoken conversation. The phone screen and visual paths are conceptual, not a reproduction of Google's actual interface or internal reasoning. Original vector illustration for BitsMinds, gemini-3-8-live-extended-thinking-voice. 16 September 2026. GEMINI 3.8 LIVE EXTENDED THINKING Gemini 3.8 LIVE The conversation continues Thinking. Still talking. BITSMINDS.COM
Models

Gemini 3.8 Live Tops Voice AI and Undercuts GPT