Skip to main content
Live
Main content

NVIDIA's Nemotron 3 Nano Omni claims 9x throughput edge over rival open multimodal models

The 30B-A3B mixture-of-experts model fuses vision, audio and text encoders, with a 256K context window and a 1920×1080 native input resolution.

Jaeden Schafer
Editor in Chief · · 5 min read
Nvidia logo

NVIDIA released Nemotron 3 Nano Omni on April 28, 2026, an open multimodal model that folds vision, audio and language into a single 30B-A3B hybrid mixture-of-experts architecture. The company says it delivers up to 9x higher throughput than other open omni models at the same interactivity, with a 256K context window and native input resolution of 1920×1080 pixels. NVIDIA also claims it tops 6 leaderboards spanning document intelligence and video and audio understanding.

The pitch is aimed squarely at agent builders. Most agentic stacks today chain separate models for vision, speech and language, paying a latency tax on every handoff and losing context across modalities. Nemotron 3 Nano Omni collapses that pipeline by housing vision and audio encoders inside the same MoE backbone, which NVIDIA frames as the perception sub-agent inside a larger system.

Nemotron 3 Nano Omni is available today on Hugging Face, OpenRouter and build.nvidia.com as an NVIDIA NIM microservice, plus 25+ partner platforms. The Nemotron 3 family — Nano, Super and Ultra — has crossed 50 million downloads in the past year, and the Omni release extends that footprint into multimodal territory.

Key facts

  • 01NVIDIA released Nemotron 3 Nano Omni on April 28, 2026 as a 30B-A3B hybrid mixture-of-experts model with a 256K context window.
  • 02NVIDIA claims up to 9x higher throughput than other open omni models at the same level of interactivity.
  • 03The model tops 6 leaderboards across document intelligence and video and audio understanding.
  • 04H Company's computer-use agent runs Nemotron 3 Nano Omni at a native 1920×1080 input resolution to read full HD screen recordings.
  • 05The Nemotron 3 family has logged 50 million downloads in the past year and ships across 25+ partner platforms.

NVIDIA is positioning the model as the perception layer that pairs with Nemotron 3 Super for high-frequency execution and Nemotron 3 Ultra for complex planning, or with proprietary cloud models from other vendors. The three target workloads are computer-use agents, document intelligence over charts and PDFs, and audio-video reasoning for customer service and monitoring.

Nemotron 3 Nano Omni tops six leaderboards for document intelligence and video and audio understanding, with up to 9x higher throughput than other open omni models at the same interactivity.
Jaeden Schafer

H Company is the first launch partner with a working demo. Its latest computer-use agent runs on Nemotron 3 Nano Omni at the model's native 1920×1080 resolution and, in preliminary evaluations on the OSWorld benchmark, posted what NVIDIA called a significant gain in navigating graphical interfaces.

"To build useful agents, you can't wait seconds for a model to interpret a screen," said Gautier Cloix, CEO of H Company. "By building on Nemotron 3 Nano Omni, our agents can rapidly interpret full HD screen recordings — something that wasn't practical before. This isn't just a speed boost: It's a fundamental shift in how our agents perceive and interact with digital environments in real time."

The early adopter list runs deeper than H Company. Aible, Applied Scientific Intelligence, Eka Care, Foxconn, Palantir and Pyler are already building on the model, while Dell Technologies, DocuSign, Infosys, K-Dense, Lila, Oracle and Zefr are evaluating it. The mix of enterprise software, healthcare, manufacturing and defense names suggests NVIDIA is courting buyers who want to keep weights in-house.

That open posture is the other piece of the strategy. Nemotron 3 Nano Omni ships with open weights, datasets and training recipes, and NVIDIA points customers to NeMo for fine-tuning. The model runs on local hardware like DGX Spark and DGX Station as well as the data center and cloud, which NVIDIA pitches as a fit for sovereignty and data-localization rules.

Related · from this week
Skild AI's S1 robot learns new factory tasks from a single video
Jaeden Schafer · 5 min read →

Author Kari Briski's blog post leans heavily on the 9x throughput claim, but the comparison set is unnamed open omni models at "similar interactivity," which is the kind of phrasing that gets sharper once independent evaluations land. The 6 leaderboard wins likewise need third-party replication, and OSWorld scores from H Company are still preliminary. Buyers weighing Nemotron against closed multimodal systems from OpenAI and Google will want benchmark numbers head-to-head, not just intra-open-source.

The release fits a pattern this week from NVIDIA, which also pushed an Omniverse manufacturing update touting 99% sim-to-real accuracy at ABB Robotics. Read together, the moves show NVIDIA pressing past chips and into the agent runtime — owning the perception model, the simulation stack and the deployment microservices that wrap them. If Nemotron 3 Nano Omni's efficiency claims hold up in the wild, it gives enterprises a credible reason to keep multimodal agent workloads on open weights rather than route every screen capture and call recording through a frontier API.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

Nvidia logo
Models

Skild AI's S1 robot learns new factory tasks from a single video

The startup hit a $100M revenue run rate 10 months after first deployment and is now installing Blackwell systems at Foxconn with Nvidia.

Jaeden Schafer5 min read
Nvidia logo
Models

Nvidia ships DLSS 5 on September 3rd, only for NBA 2K27 and RTX 50-series

The AI upscaler carries a 50–60% performance hit and needs 6x frame generation to hit 1080p on midrange cards.

Jaeden Schafer5 min read
Meta logo
Models

Meta opens Muse Spark 1.1 to developers via new Meta Model API

The upgraded model targets agentic coding and multimodal reasoning, arriving days after the controversial Muse Image launch.

Jaeden Schafer4 min read