Skip to main content
Live
Main content

NVIDIA Vera CPU delivers 1.5x performance edge over 128-core x86 in Phoronix tests

Custom Olympus cores and 1.2TB/s memory bandwidth push Vera past Intel and AMD in agentic AI workloads.

Jaeden Schafer
Editor in Chief · · 5 min read
Nvidia logo

NVIDIA Vera CPU delivered a 1.5x overall performance advantage over a latest-generation 128-core x86 processor in Phoronix benchmark testing published today, marking the strongest ARM-based challenge to Intel Xeon and AMD EPYC in the data center. The chip's 88 custom Olympus cores and 1.2TB/s memory bandwidth combine to outpace traditional x86 architectures on the sequential CPU tasks underpinning agentic AI — branch-heavy runtimes, sandboxed code execution, and large-scale orchestration.

Vera posted a 10% performance lead over the AMD EPYC 9575F running at 5.0 GHz, the fastest single-threaded x86 chip in Phoronix's test set. The gains showed up across workloads spanning code compilation, file compression, video transcoding, Python runtimes, Java execution, and database management — the same CPU-intensive tasks that AI agents and AI factories run daily.

NVIDIA custom-designed the Olympus cores for agentic workloads rather than licensing an off-the-shelf ARM core design. The cores are fully compatible with the Armv9.2 instruction set but feature wide pipelines, advanced branch prediction, and tight integration with NVIDIA's second-generation Scalable Coherency Fabric. That fabric keeps data moving across all 88 cores on a monolithic die, avoiding the latency penalties of chiplet-based x86 designs at this core count.

Key facts

  • 01NVIDIA Vera delivered 1.5x overall performance vs. a latest-generation 128-core x86 processor in Phoronix testing.
  • 02The CPU achieved 1.2TB/s of memory bandwidth — 2x traditional CPUs — using less than 30 watts of memory power vs. more than 100 watts for DDR5.
  • 03Vera sustained 90% of its peak memory bandwidth in STREAM TRIAD testing, the highest percentage Phoronix has measured.
  • 04Single-socket Vera compiled a Linux kernel in 20 seconds, 2x faster per core than a 128-core x86 chip.
  • 05The chip posted a 1.6x geometric mean performance increase over NVIDIA's prior-generation Grace CPU.

Vera's memory subsystem is the headline technical differentiator. The chip integrates a second-generation LPDDR5X controller delivering up to 1.2TB/s of bandwidth — 2x the peak of traditional DDR5-based CPUs — while consuming less than 30 watts for memory, compared to more than 100 watts on DDR5. That power efficiency matters in AI factories running hundreds or thousands of CPU sockets continuously.

Phoronix STREAM TRIAD testing showed Vera sustaining 90% of its rated peak bandwidth, the highest percentage of any CPU Phoronix has measured. Vera also delivered 4x the memory bandwidth per core compared to traditional x86 processors. Separate testing by Prime Intellect found that Vera maintained high bandwidth and low, consistent memory latency as workloads scaled — predictable performance under the parallel sandboxes and tool calls that agentic AI generates.

NVIDIA Vera with its LPDDR5X memory was showing its incredible advantage in memory performance over current Intel Xeon and AMD EPYC processors.
Michael Larabel, Phoronix founder

Compared to NVIDIA's prior-generation Grace CPU, Vera posted a 1.6x geometric mean performance increase. Grace launched in 2023 and was the company's first ARM-based data-center chip. The Vera-to-Grace generational leap exceeds the typical year-over-year gains seen in x86 processor families, where Intel and AMD routinely deliver 10–20% improvements per generation.

Practical developer benchmarks confirmed the Phoronix results. Single-socket Vera compiled a default Linux kernel in 20 seconds, the fastest result Phoronix has recorded in that test. On a per-core basis, Vera compiled the kernel 2x faster than a 128-core x86 processor. Compilation speed matters for AI development workflows where agents repeatedly build and test code.

NVIDIA announced ecosystem support for Vera at GTC earlier this year and has begun delivering chips to leading AI companies and cloud providers. The company expects partner availability in the second half of 2026, with dual- and single-socket systems available in both air-cooled and liquid-cooled configurations. The chip carries a 450-watt thermal design power rating.

On a [geometric] mean basis, the NVIDIA Vera delivered 10% better performance than the AMD EPYC 9575F 5.0 GHz high frequency processor.
Michael Larabel, Phoronix founder
Related · from this week
Nvidia ships DLSS 5 on September 3rd, only for NBA 2K27 and RTX 50-series
Jaeden Schafer · 5 min read →

The Vera results represent a structural shift in the data-center CPU market. ARM-based server chips have existed for over a decade — Amazon's Graviton, Ampere's Altra, and Marvell's ThunderX all targeted the market before Vera — but none achieved price-performance or absolute performance parity with x86 at scale. Vera is the first ARM chip to lead a major independent benchmark suite across a broad workload mix, not just on narrow power-efficiency or cost-per-core metrics.

The chip's real test will come under production agentic AI loads. Phoronix benchmarks are reproducible and well-understood, but they measure synthetic or general-purpose workloads. AI factories running tool-calling agents, sandboxed Python environments, and massively parallel inference orchestration will stress different parts of the architecture. Early customer deployments in the second half of 2026 will determine whether Vera's lead holds under those conditions.

Vera's pricing and total cost of ownership remain undisclosed. NVIDIA has not announced list prices, volume discounts, or partner pricing for systems integrating Vera. The economics matter as much as the performance — if Vera costs 50% more per socket than a comparable x86 chip, the 1.5x performance advantage compresses quickly once software licensing, power infrastructure, and operational overhead enter the calculation.

For NVIDIA, Vera is a strategic hedge and a margin play. The company dominates AI accelerators with its GPU lineup, but CPUs control the host server and orchestrate multi-GPU jobs. Owning the full stack — CPU, GPU, networking, and orchestration software — lets NVIDIA capture more of the AI factory buildout while reducing dependency on Intel and AMD as suppliers. If Vera hits its performance and cost targets at volume, it shifts the data-center CPU duopoly into a three-way race for the first time in two decades.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

Nvidia logo
Models

Nvidia ships DLSS 5 on September 3rd, only for NBA 2K27 and RTX 50-series

The AI upscaler carries a 50–60% performance hit and needs 6x frame generation to hit 1080p on midrange cards.

Jaeden Schafer5 min read
Nvidia logo
Models

Nvidia starts shipping Vera, its first CPU built for AI agents

Vera pairs with NVLink Fusion and a new NVHBM custom high-bandwidth memory tier aimed at trillion-parameter agent workloads.

Jaeden Schafer4 min read
Nvidia logo
Models

Nvidia Nemotron 3 Ultra hits closed-model parity at 10x lower cost on LangChain

LangChain tuned its Deep Agents harness for Nemotron 3 Ultra, matching top closed models on business tasks without retraining.

Jaeden Schafer5 min read