Nvidia's Vera Rubin NVL72 is in production, with racks running at CoreWeave, Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure and a first live benchmark showing 10x more tokens per second per megawatt than Grace Blackwell NVL72 on DeepSeek-R1. Nvidia disclosed the numbers Tuesday, alongside a claim of one-tenth the cost per million tokens versus GB200 NVL72. The platform is backed by 300 global partners and 350-plus factory sites across 30 countries, which Nvidia calls the largest rack-scale supply chain it has assembled.
The headline metric is tokens per megawatt, not raw FLOPs. For AI factories running into power constraints at the substation level, throughput per watt determines whether a build pencils out. CoreWeave's measured 10x gain over the prior Blackwell generation lands directly on that constraint, and Nvidia is positioning the platform's economics around it — a 10x throughput improvement paired with an order-of-magnitude drop in per-token cost, if the numbers hold across workloads beyond DeepSeek-R1.
Vera Rubin NVL72 is built from seven co-designed chips across five rack trays: the Vera Rubin GPU, Vera CPU, Groq 3 LPX, Spectrum-6 SPX and Vera BlueField-4 STX. At the center is the Vera CPU, whose custom Olympus core delivers what Nvidia claims is 2x single-threaded performance, 3x core-to-core bandwidth and 40% lower memory latency versus competing chiplet designs. That single-threaded focus is deliberate — agentic workloads, which Nvidia says can consume up to 15x more tokens than traditional AI applications, punish CPUs that stall on serial coordination.
Key facts
- 01CoreWeave's DeepSeek-R1 benchmark on Vera Rubin NVL72 delivered 10x more tokens per second per megawatt than Grace Blackwell NVL72.
- 02Vera Rubin NVL72 delivers one-tenth the cost per million tokens compared with GB200 NVL72, per Nvidia's figures.
- 03Production racks are live at CoreWeave, Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure, backed by 350-plus factory sites across 30 countries.
- 04Microsoft and Mistral signed a multibillion-dollar deal drawing on tens of thousands of Vera Rubin GPUs for European AI infrastructure.
- 05The rack's 260 TB/s all-to-all NVLink 6 fabric and 45C liquid-cooling inlet cut compute tray assembly to one minute and save millions of gallons of water per megawatt annually.
Networking is the other lever. The sixth-generation NVLink scale-up fabric delivers more than 2x throughput on complex workloads, 3x lower latency and 10x higher packet rates than off-the-shelf Ethernet, according to Nvidia. Each rack's 260 TB/s all-to-all NVLink 6 fabric lets 72 GPUs behave as a single accelerator — critical for mixture-of-experts models like DeepSeek-R1 that route every token across distributed expert sub-networks.
For scale-out between racks, Spectrum-X Ethernet pairs 102.4 Tb/s Spectrum-6 switch systems with 1.6T ConnectX-9 SuperNICs and claims 1.6x higher RDMA bandwidth than off-the-shelf Ethernet. CoreWeave is among the first deploying the SN6600-LD switch, hitting 1.64 Pb/s per rack with 100% more capacity than the prior generation of air-cooled switches. Nvidia Photonics with co-packaged optics — now in volume manufacturing — adds a claimed 5x lower power and 10x higher mean time between interruptions versus pluggable transceivers, with CoreWeave, Lambda and OCI as early adopters. This follows Nvidia's Spectrum-6 launch we covered earlier this month.
The physical design also shifted. Three generations into rack-scale co-design, the Vera Rubin NVL72 compute tray ships with no cables, fans or hoses, cutting assembly time from hours to one minute. A 45-degree Celsius liquid-cooling inlet temperature allows chiller-free dry-cooler operation, which Nvidia estimates saves millions of gallons of water per megawatt annually versus evaporative designs. That last figure matters for the siting fights that have started to slow US and European data center approvals.
In Europe, Microsoft and Mistral signed a new multibillion-dollar agreement built on Vera Rubin, with Mistral expanding its GPU capacity using tens of thousands of the new GPUs. Mistral Medium 3.5 and OCR 4 are now available in Microsoft Foundry, and Mistral models are integrated into Microsoft Copilot Studio, with Azure Local and Foundry Local extending the same stack into customer-controlled environments. The pitch to European governments and regulated industries is a sovereign-ready deployment path that doesn't force a tradeoff between open models, cloud economics and data control.
“The next era of research requires the next era of hardware.”— Lasse Espeholt, Cofounder of Ineffable Intelligence
Google Cloud's first A5X instance, powered by Vera Rubin NVL72, is running for London-based startup Ineffable Intelligence, which is building reinforcement-learning "superlearner" systems that train through continuous simulated interaction rather than static datasets. Those tightly coupled learning loops place unusually heavy demands on interconnect latency and memory bandwidth — the exact profile the NVL72 architecture is tuned for.
The caveats are worth naming. The 10x tokens-per-megawatt figure comes from a single benchmark on a single model, run by a partner with commercial interest in the result. DeepSeek-R1's MoE architecture stresses the all-to-all fabric that NVLink 6 was designed to accelerate, so gains on denser architectures may compress. Nvidia has not yet published third-party MLPerf results for Vera Rubin, and cost-per-token claims depend on utilization assumptions that hyperscalers rarely disclose. Buyers will want their own measurements before committing capital.
The commercial read is that Nvidia is defending its position not on peak performance but on total cost of intelligence per watt — the metric that actually gates AI factory economics as power becomes the binding constraint. If the 10x figure holds across production workloads, the platform resets the bar for what a competitive alternative from AMD, Google's TPUs or the wave of custom silicon has to deliver. And by opening NVLink Fusion to third-party XPUs, Nvidia is trying to make its scale-up fabric the default substrate even for chips it doesn't build — a subtler moat than raw silicon performance, and a harder one to displace.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




