Nvidia unveiled Vera, an 88-core Arm CPU built for agentic AI workloads, claiming a 50% instructions-per-cycle gain over its predecessor Grace and 1.8x the sustained per-core performance of x86 under load. Perplexity ran the chip through its production coding pipeline and completed a repository-clone-and-test workflow about 1.5x faster than x86, with concurrent sandbox startup up to 1.9x faster. The company said it is now looking to deploy Vera in its upcoming production system.
The design bet is that data center CPUs have spent a decade optimizing the wrong axis. Cloud economics pushed vendors toward higher core counts at lower cost per rentable core, and chiplet architectures cut die costs while starving individual cores of memory bandwidth. Nvidia argues that agent workloads, where each step in a reasoning loop depends on the previous result, punish that tradeoff. More cores do not shorten a sequential tool call, a code execution, or a data-processing pass.
Vera's custom Olympus core sits on a monolithic compute die paired with up to 1.2TB/s of LPDDR5X memory bandwidth at under 40 watts of memory power. Core-to-core bandwidth runs at 3.4TB/s, which Nvidia says is 3x any other data center CPU. The point of that fabric is to let all 88 cores hit full memory performance simultaneously without the contention penalty that shows up when workloads saturate a many-core chip.
“The world counts in seconds. Agents count in nanoseconds.”— Ian Buck, Nvidia VP of Hyperscale and HPC
Key facts
- 01Vera's Olympus core delivers 50% higher instructions per cycle than Nvidia Grace, with 88 cores per chip.
- 02The CPU pairs 1.2TB/s of LPDDR5X memory bandwidth at under 40W with 3.4TB/s of core-to-core bandwidth, 3x any other data center CPU.
- 03Perplexity ran a real coding workflow 1.5x faster than x86 on Vera and started concurrent sandboxes 1.9x faster.
- 04Partners measured 3x faster large-scale SQL analytics with Starburst and 6x lower streaming latency with Redpanda vs leading x86.
- 05Nvidia's next CPU, Rosa, will use the Rigel core built on Arm v9.2 with higher per-core performance than Olympus.
The framing matters for AI factory economics. GPUs are the expensive asset in a training or inference cluster, and any second a GPU spends idle waiting for a CPU-side tool call, data query, or verification pass is revenue foregone. Nvidia is positioning Vera as the way to keep GPU utilization high by shortening every non-GPU step in the agent loop. The compounding effect across millions of tool calls, sandbox spin-ups, and SQL queries is where the vendor is making its case.
Partner benchmarks focus on that data-adjacent work. Starburst measured 3x faster large-scale SQL analytics on Vera compared to leading x86 server CPUs. Redpanda reported up to 6x lower latency on real-time streaming. Both are workloads that sit inside agent execution paths, feeding retrieval, verification, and reinforcement-learning loops rather than model inference itself.
Vera is also the CPU inside the Vera Rubin GPU platform and the BlueField-4 STX storage processor, meaning Nvidia is trying to collapse an AI factory onto a single CPU architecture and toolchain. That vertical alignment is familiar Nvidia strategy — the same playbook that made CUDA sticky is now being extended down to the host CPU and out to storage. Customers running mixed workloads across training, inference, agents, and data pipelines get one target rather than several.
“AI factories need a CPU with max single-threaded performance to maximize AI factory revenue and agent performance.”— Ian Buck, Nvidia VP of Hyperscale and HPC
The roadmap does not stop at Vera. Nvidia said its next-generation Rosa CPU will use a new Rigel core built on Arm v9.2, promising higher per-core performance than Olympus in the same silicon footprint. Rosa gains include better instruction delivery, a larger L2 cache, and more efficient memory handling. Timing was not disclosed at GTC Berlin, which runs October 20–22.
The unresolved question is how much of the per-core advantage holds up outside Nvidia-supplied benchmarks. The 1.8x sustained per-core figure against x86 is measured under agentic workloads Nvidia selected, and the Perplexity, Starburst, and Redpanda numbers are partner-run. Intel and AMD have their own agent-oriented CPU roadmaps and will push back on comparisons they did not participate in. Independent testing on mixed enterprise workloads will decide whether Vera's design tradeoffs generalize or whether the wins are concentrated in the workloads Nvidia optimized for.
Vera is Nvidia's argument that the CPU is no longer a commodity host chip in an AI system. If agents multiply into the billions of concurrent loops that Nvidia is projecting, the CPU sitting next to each GPU becomes a direct lever on how much work an AI factory can bill for. Nvidia is now selling the GPU, the CPU, the DPU, the networking, and the software stack that ties them together — and Vera is the piece that makes the rest of that stack run harder. For AMD's Epyc and Intel's Xeon franchises, which have long treated hyperscale AI as GPU-adjacent revenue, that is the more consequential shift than any single benchmark number.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.



