NVIDIA released Nemotron 3 Nano Omni on April 28, 2026, an open multimodal model that folds vision, audio and language into a single 30B-A3B hybrid mixture-of-experts architecture. The company says it delivers up to 9x higher throughput than other open omni models at the same interactivity, with a 256K context window and native input resolution of 1920×1080 pixels. NVIDIA also claims it tops 6 leaderboards spanning document intelligence and video and audio understanding.
The pitch is aimed squarely at agent builders. Most agentic stacks today chain separate models for vision, speech and language, paying a latency tax on every handoff and losing context across modalities. Nemotron 3 Nano Omni collapses that pipeline by housing vision and audio encoders inside the same MoE backbone, which NVIDIA frames as the perception sub-agent inside a larger system.
Nemotron 3 Nano Omni is available today on Hugging Face, OpenRouter and build.nvidia.com as an NVIDIA NIM microservice, plus 25+ partner platforms. The Nemotron 3 family — Nano, Super and Ultra — has crossed 50 million downloads in the past year, and the Omni release extends that footprint into multimodal territory.
Key facts
- 01NVIDIA released Nemotron 3 Nano Omni on April 28, 2026 as a 30B-A3B hybrid mixture-of-experts model with a 256K context window.
- 02NVIDIA claims up to 9x higher throughput than other open omni models at the same level of interactivity.
- 03The model tops 6 leaderboards across document intelligence and video and audio understanding.
- 04H Company's computer-use agent runs Nemotron 3 Nano Omni at a native 1920×1080 input resolution to read full HD screen recordings.
- 05The Nemotron 3 family has logged 50 million downloads in the past year and ships across 25+ partner platforms.
NVIDIA is positioning the model as the perception layer that pairs with Nemotron 3 Super for high-frequency execution and Nemotron 3 Ultra for complex planning, or with proprietary cloud models from other vendors. The three target workloads are computer-use agents, document intelligence over charts and PDFs, and audio-video reasoning for customer service and monitoring.
“Nemotron 3 Nano Omni tops six leaderboards for document intelligence and video and audio understanding, with up to 9x higher throughput than other open omni models at the same interactivity.”— Jaeden Schafer
H Company is the first launch partner with a working demo. Its latest computer-use agent runs on Nemotron 3 Nano Omni at the model's native 1920×1080 resolution and, in preliminary evaluations on the OSWorld benchmark, posted what NVIDIA called a significant gain in navigating graphical interfaces.
"To build useful agents, you can't wait seconds for a model to interpret a screen," said Gautier Cloix, CEO of H Company. "By building on Nemotron 3 Nano Omni, our agents can rapidly interpret full HD screen recordings — something that wasn't practical before. This isn't just a speed boost: It's a fundamental shift in how our agents perceive and interact with digital environments in real time."
The early adopter list runs deeper than H Company. Aible, Applied Scientific Intelligence, Eka Care, Foxconn, Palantir and Pyler are already building on the model, while Dell Technologies, DocuSign, Infosys, K-Dense, Lila, Oracle and Zefr are evaluating it. The mix of enterprise software, healthcare, manufacturing and defense names suggests NVIDIA is courting buyers who want to keep weights in-house.
That open posture is the other piece of the strategy. Nemotron 3 Nano Omni ships with open weights, datasets and training recipes, and NVIDIA points customers to NeMo for fine-tuning. The model runs on local hardware like DGX Spark and DGX Station as well as the data center and cloud, which NVIDIA pitches as a fit for sovereignty and data-localization rules.
Author Kari Briski's blog post leans heavily on the 9x throughput claim, but the comparison set is unnamed open omni models at "similar interactivity," which is the kind of phrasing that gets sharper once independent evaluations land. The 6 leaderboard wins likewise need third-party replication, and OSWorld scores from H Company are still preliminary. Buyers weighing Nemotron against closed multimodal systems from OpenAI and Google will want benchmark numbers head-to-head, not just intra-open-source.
The release fits a pattern this week from NVIDIA, which also pushed an Omniverse manufacturing update touting 99% sim-to-real accuracy at ABB Robotics. Read together, the moves show NVIDIA pressing past chips and into the agent runtime — owning the perception model, the simulation stack and the deployment microservices that wrap them. If Nemotron 3 Nano Omni's efficiency claims hold up in the wild, it gives enterprises a credible reason to keep multimodal agent workloads on open weights rather than route every screen capture and call recording through a frontier API.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




