Skip to main content
Live
Main content

Nvidia starts shipping Vera, its first CPU built for AI agents

Vera pairs with NVLink Fusion and a new NVHBM custom high-bandwidth memory tier aimed at trillion-parameter agent workloads.

Jaeden Schafer
Editor in Chief · · 4 min read
Nvidia logo

Nvidia has begun shipping Vera, its first CPU designed specifically for AI agents and trillion-parameter workloads, alongside an expansion of NVLink Fusion with a new custom high-bandwidth memory tier called NVHBM. The company framed the launch as a response to the shift from single-shot model inference to long-running, multi-step agent execution, where memory bandwidth and interconnect topology matter as much as raw compute. Vera is the CPU half of the Vera Rubin platform Nvidia has been previewing for over a year.

The pitch is architectural. As agent workloads move mainstream, the bottleneck stops being how many floating-point operations a GPU can grind through and starts being how quickly the system as a whole can feed those GPUs with context, tool outputs, and intermediate state. Nvidia's argument is that a CPU designed around that traffic pattern, tightly coupled to Rubin GPUs over NVLink Fusion and backed by NVHBM, closes the gap.

NVLink Fusion is the interconnect fabric Nvidia uses to stitch CPUs, GPUs, and memory into a single coherent domain. Adding NVHBM as a custom high-bandwidth memory tier inside that fabric gives the platform a place to park the long-lived state that agents accumulate — retrieved documents, tool call histories, planning traces — without paying the latency cost of round-tripping to standard system memory. For workloads that run for minutes rather than milliseconds, that latency delta compounds.

Key facts

  • 01Nvidia's Vera CPU, the company's first processor purpose-built for AI agents, is shipping now.
  • 02NVLink Fusion is expanding with NVHBM, a custom high-bandwidth memory tier designed for trillion-parameter workloads.
  • 03Vera is positioned as the CPU half of the Vera Rubin platform underpinning Anthropic's recently announced $45B compute deal with Nscale.

The commercial context is that Vera is not launching into a vacuum. Anthropic signed a $45B compute deal with Nscale earlier this quarter specifically for Nvidia Vera Rubin systems, one of the largest single infrastructure commitments in the AI market to date. Amazon separately tripled its Nvidia GPU order for AWS, adding roughly two million Blackwell and Rubin chips. Vera shipping now is what unlocks those orders from paper commitments into deployable capacity.

Nvidia positions the trillion-parameter workload as the target profile, which is a deliberate framing. Frontier models from OpenAI, Anthropic, and Google DeepMind have crossed or approached that scale, and the agent products those models power — Claude's computer use, ChatGPT's agent mode, Gemini's task-completion features — hold state across dozens of steps. Serving those products at scale is a different problem than serving a chatbot, and Nvidia is arguing Vera is the CPU designed for it.

The NVHBM naming signals customization. Standard HBM3E and the incoming HBM4 are commodity parts sourced from SK Hynix, Samsung, and Micron. A custom high-bandwidth memory variant tuned to Nvidia's fabric and Rubin's memory controller lets the company differentiate at a layer where competitors sourcing off-the-shelf HBM cannot easily match. It also deepens the lock-in for buyers who standardize on Vera Rubin — the memory tier is not a drop-in swap.

Vera also lands in a market where AMD's MI400 series and custom silicon from Google, Amazon, and Microsoft are all pushing into agent-scale inference. Nvidia's response has consistently been to sell the platform rather than the chip: CPU, GPU, interconnect, memory, and networking as a single procurement. Vera shipping completes that platform for the Rubin generation.

The financial stakes trace directly back to Nvidia's guidance. The company posted a $96.2B quarter last cycle and guided to $108B for the next, with data center revenue continuing to roughly double year over year. Vera Rubin systems are what those numbers are increasingly built on, replacing Hopper and eventually Blackwell as the workhorse for hyperscaler and neocloud buildouts.

Related · from this week
Nvidia says the harness, not the model, drives Claude Opus 5 to 100% on ARC-AGI-3
Jaeden Schafer · 5 min read →

Not everything is settled. Nvidia disclosed the shipping status and the NVHBM expansion but did not publish detailed specifications, pricing, or benchmark comparisons against its own Grace CPU or against AMD's Turin and Bergamo parts on agent-representative workloads. Independent evaluations will take months to surface, and early Vera Rubin deployments will be inside Nscale, AWS, and other hyperscaler facilities where third-party access is limited. Until customers publish real workload numbers, the trillion-parameter claim rests on Nvidia's own framing.

Vera shipping matters less as a single-product event and more as the piece that turns two years of Rubin-era announcements into deployable infrastructure. Every AI agent product now under development at frontier labs assumes a compute substrate that can hold state cheaply and route it fast; Vera is Nvidia's answer to that assumption, and the $45B and multi-million-GPU orders already signed against the platform mean the answer is going to be tested at production scale within the next few quarters. The competitive question for AMD and the in-house silicon efforts is no longer whether they can match Blackwell — it is whether they can match a platform where the CPU, the interconnect, and the memory tier were designed together.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

Nvidia logo
Models

Nvidia says the harness, not the model, drives Claude Opus 5 to 100% on ARC-AGI-3

New Nvidia research shows a custom harness with a supervisor agent lifted Claude Opus 5 from 30% to a perfect score on the interactive reasoning benchmark.

Jaeden Schafer5 min read
Nvidia logo
Models

Nvidia releases Nemotron 3.5 Lightning and NeMo Switchyard for agent workloads

The 30B mixture-of-experts model runs 4x faster on output, while the routing library cuts task cost to a third of Opus 4.8.

Jaeden Schafer5 min read
Nvidia logo
Business

Nvidia claims Vera CPU opens $200B agentic AI market, sells $20B this year

Jensen Huang positions the new CPU as purpose-built for agents, not traditional cloud cores.

Jaeden Schafer5 min read