Skip to main content
Live
Main content

Nvidia ships physical AI agent skills at CVPR, anchored by 32B Alpamayo 2 Super

New agent skills built on Cosmos 3 automate scene reconstruction, simulation and policy training for AVs, robots and vision AI.

Jaeden Schafer
Editor in Chief · · 5 min read
Nvidia logo

Nvidia unveiled a suite of physical AI agent skills at CVPR 2026, headlined by Alpamayo 2 Super, an open 32-billion-parameter reasoning vision-language-action model aimed at level 4 autonomous driving. The skills sit on top of Nvidia Cosmos 3, the company's open frontier model for physical AI announced earlier in the week, and target the workflow stitching that has slowed robotics and AV research. Nvidia's Physical AI Dataset has now passed 15 million downloads on Hugging Face, a marker of how much of the field's training data already routes through the company's infrastructure.

The pitch is workflow automation, not just bigger models. Nvidia is packaging agent-callable tools for scene reconstruction, synthetic data generation, simulation, policy training and evaluation across autonomous vehicles, robotics and vision AI. AlpaGym, an open closed-loop reinforcement learning framework, scales across thousands of GPUs to connect policy rollouts with high-fidelity simulation. OmniDreams, an action-conditioned generative world model, renders camera frames in real time in response to policy actions.

Nvidia's framing, in a blog post by Pranjali Joshi, is that researchers spend more time gluing tools together than improving models.

The core challenge in physical AI research isn't simply developing stronger models. It's building a full workflow around them — reconstructing real-world scenes, generating edge-case scenarios, training policies, evaluating behavior and rapidly iterating.
Pranjali Joshi, Nvidia

Key facts

  • 01Nvidia released Alpamayo 2 Super, a 32-billion-parameter open reasoning vision-language-action model for level 4 autonomous driving.
  • 02The Nvidia Physical AI Dataset has crossed 15 million downloads on Hugging Face, with new GRAIL data adding roughly 50 hours of humanoid-object interaction.
  • 03AlpaGym, an open closed-loop reinforcement learning framework, scales policy rollouts across thousands of GPUs.
  • 04Nvidia is launching three CVPR benchmarks, including the 10th AI City Challenge and the new PAI-AV Reasoning Challenge.
  • 05CVPR 2026 runs June 3-7 in Denver, where Nvidia tech is referenced in the majority of accepted papers.

For autonomous vehicles, the problem Nvidia is targeting is the long tail of rare driving scenarios that can't be reliably collected on the road. Neural Reconstruction skills turn fleet data into editable 3D scenes for simulation, using Nvidia Omniverse NuRec, InstantNuRec, Harmonizer and the HiGS accelerated renderer. Alpamayo 2 Super then reasons, plans and acts across the full driving stack, and a new AlpaSim Closed-Loop End-to-End Driving Challenge benchmarks policies against real-world reconstructed scenarios.

On the vision AI side, new Nvidia Metropolis skills lean on Cosmos 3's mixture-of-transformers architecture to generate synthetic anomalies and augment data for inspection models. A Defect Image Generation skill creates examples of rare defects across surfaces by combining Nvidia Isaac Sim, Cosmos 3 and Nvidia OSMO for orchestration. For video reasoning, the Metropolis Blueprint for video search and summarization, paired with Nvidia TAO and Video Augmentation skills, automates the build-and-evaluate loop for agents that detect events and summarize activity.

Robotics gets the broadest skill coverage. Agents can author scenes, launch simulation sessions, capture data and validate environments inside Isaac Sim, while Isaac Lab skills handle reinforcement learning setup, training and evaluation. Specialized Isaac mobility skills cover scene search, USD conversion, residual reinforcement learning and policy evaluation. For healthcare robotics, Cosmos-H-Surgical-Simulator generates surgical robotics data by learning from real surgical footage rather than hand-engineered physics, aiming to narrow the sim-to-real gap.

Nvidia is also expanding the underlying data layer. New releases include GRAIL, with roughly 50 hours of humanoid-object interaction data, and six synthetic video datasets used to train Cosmos 3 spanning robotics, physics, digital humans, autonomous driving, warehouse safety and spatial reasoning. Nvidia Isaac GR00T X Embodiment Sim is among the most-downloaded robotics datasets on Hugging Face, a useful comparison against general-purpose LLM datasets that dominate the platform's top lists.

The CVPR push extends beyond product. Nvidia says its GPUs, open models and CUDA-accelerated libraries are referenced in the majority of accepted CVPR 2026 papers, with adoption at Carnegie Mellon University, Stanford University, UC Berkeley, Tsinghua University and Peking University. The conference runs June 3-7 in Denver. Nvidia is anchoring three open challenges there: the AI City Challenge in its tenth year, the new PAI-AV Reasoning Challenge testing chain-of-causation reasoning in driving models, and the AlpaSim closed-loop driving benchmark.

Related · from this week
Nvidia releases Cosmos 3, an open world model family for physical AI
Jaeden Schafer · 5 min read →

The open question is how much of this stack researchers will actually adopt end-to-end versus picking off individual pieces. Cosmos 3, Alpamayo 2 Super and the agent skills are open and live on GitHub, with Neural Reconstruction, Video Augmentation and Defect Image Generation available as preconfigured Physical AI Launchables on Nvidia Brev. But labs already running custom simulation pipelines won't migrate wholesale, and competing world-model efforts from Google DeepMind and academic groups will pull research attention in other directions. The reliance on Nvidia GPUs and CUDA is also a lock-in story, not a neutral one.

This is Nvidia building the same kind of moat in physical AI that CUDA built in deep learning a decade ago. The pattern is familiar: release the model, release the dataset, release the simulation framework, release the agent skills that connect them, and host the benchmark that measures progress on your terms. Every layer is open, but every layer assumes Nvidia silicon and Nvidia tooling underneath. For AV and robotics startups, the calculation is whether faster iteration on Nvidia's stack outweighs the cost of building portability for a future that may or may not arrive.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

Nvidia logo
Models

Nvidia releases Cosmos 3, an open world model family for physical AI

The three-tier release spans 4B to 64B parameters and tops five benchmarks for open-weights image, video, world and robot generation.

Jaeden Schafer5 min read
Nvidia logo
Models

Nvidia opens Alpamayo 2 Super for commercial robotaxi deployment

The 30B-parameter driving model tops LingoQA against Qwen, Gemini and GPT-4o, and ships under a permissive Linux Foundation license.

Jaeden Schafer5 min read
Nvidia logo
Models

Nvidia unveils Cosmos 3, a foundation model for physical AI reasoning

Nvidia Research positions Cosmos 3 as the bridge from scripted robotics demos to embodied autonomy in open-world environments.

Jaeden Schafer4 min read