Models·2 min read
By BitsMindsSource: NVIDIA Blog

NVIDIA Unveils Nemotron 3 Nano Omni: Open Multimodal Model with 9x Throughput

NVIDIA's new 30B-A3B mixture-of-experts model unifies vision, audio, and language into a single open system, delivering up to nine times the throughput of comparable open omni models for AI agents.

NVIDIA Unveils Nemotron 3 Nano Omni: Open Multimodal Model with 9x Throughput
Share:

NVIDIA on April 28, 2026 launched Nemotron 3 Nano Omni, a new open-weight multimodal model that fuses vision, speech, and language into a single system designed for autonomous AI agents. Built on a 30B-A3B hybrid mixture-of-experts architecture with a 256K-token context window, the model accepts text, images, audio, video, documents, charts, and graphical interfaces as input — a significant expansion over previous Nemotron releases, which were text-only across the Nano, Super, and Ultra tiers.

The headline figure is efficiency: NVIDIA says Nemotron 3 Nano Omni delivers up to nine times the throughput of comparable open omni models at the same level of interactivity, translating into lower inference costs and broader scalability. The model topped six leaderboards for document intelligence and combined audio-video understanding at launch, helped by new components including Conv3D and an Enhanced Visual System (EVS) that improve dense visual reasoning across long video clips and high-resolution screen captures.

Early adopters span both enterprise and AI-native companies. Aible, Applied Scientific Intelligence, Eka Care, Foxconn, H Company, Palantir, and Pyler are already deploying the model in production, while Dell Technologies, Docusign, Infosys, and Oracle are evaluating it. H Company CEO Gautier Cloix said in NVIDIA's announcement that "by building on Nemotron 3 Nano Omni, our agents can rapidly interpret full HD screen recordings — something that wasn't practical before," highlighting the appeal of unified perception for desktop-automation agents.

Nemotron 3 Nano Omni is available immediately through Hugging Face, OpenRouter, build.nvidia.com, and more than 25 partner platforms, with deployment options ranging from on-device and on-prem to public cloud. The release lands as the broader AI ecosystem pivots from chat-first models toward autonomous agents capable of seeing, hearing, and acting — a category in which efficient open multimodal foundations are quickly becoming strategic. By open-sourcing a model that runs roughly an order of magnitude faster than rivals, NVIDIA is moving to anchor the next wave of agent infrastructure around its own software stack.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles

Many kinds of input, one Qwen context On a deep violet field, a filmstrip, an audio waveform, a landscape photograph and a text page curve toward a single luminous sphere bearing the official purple Qwen star. A label below reads “1M context”, representing text, images, audio and video sharing one million-token context window. Aa 1M context BITSMINDS.COM
Models

Qwen3.8-Omni-Flash Undercuts Gemini on Audio and Video

Hill Climb: GPT-6 Astra versus Claude Fable 5.1 A teal desert buggy climbs a sandstone ridge on the left. An orange jeep climbs a green hill on the right, under an arc of gold coins. The two landscapes meet at a diagonal divide. 7 GPT-6 ASTRA CLAUDE FABLE 5.1 VS HILL CLIMB BITSMINDS.COM
Models

GPT-6 Astra vs Claude Fable 5.1: Hill Climb

Gemini 3.8 Live: thinking while the conversation continues An editorial illustration in a dark blue and violet room. A carefully drawn studio microphone and a tilted smartphone flank a translucent speech bubble carrying the multicoloured Gemini star. A continuous luminous audio waveform travels between them. Above the bubble, a separate arc connects small search, reasoning and completion symbols, representing background work continuing during a spoken conversation. The phone screen and visual paths are conceptual, not a reproduction of Google's actual interface or internal reasoning. Original vector illustration for BitsMinds, gemini-3-8-live-extended-thinking-voice. 16 September 2026. GEMINI 3.8 LIVE EXTENDED THINKING Gemini 3.8 LIVE The conversation continues Thinking. Still talking. BITSMINDS.COM
Models

Gemini 3.8 Live Tops Voice AI and Undercuts GPT