
Nvidia ships Groq 3 LPX to accelerate Vera Rubin agentic inference
The new accelerator hits 3,400 output tokens per second on 100,000-token contexts, 4x the nearest platform, with Nebius as first cloud adopter.
Tag · 11 stories
Every story tagged Vera Rubin on AI Chat Daily.

The new accelerator hits 3,400 output tokens per second on 100,000-token contexts, 4x the nearest platform, with Nebius as first cloud adopter.

Vera Rubin and Grace Blackwell systems shipping early next year will cost hyperscalers more, with memory makers squeezing the world's most profitable chipmaker.

Ilya Sutskever's stealth lab gets Vera Rubin access and multi-billion Nvidia investment on top of a $32B valuation.

The 324,000-square-foot Texas facility already employs 500 workers and will scale to 1,000 by year-end, producing tens of thousands of boards a month.

CoreWeave's DeepSeek-R1 benchmark shows a 10x throughput-per-watt gain over Grace Blackwell, with racks now live at four major clouds.

The next-gen switch doubles capacity and lands first at CoreWeave, Microsoft, Nebius, SpaceXAI and Tesla to wire hundreds of thousands of GPUs.

GB300 NVL72 delivers up to 25x performance per watt over Hopper as power becomes the binding constraint on inference economics.

The joint platform adds Nvidia's first agent-tuned CPU, a new agent toolkit for Private Cloud AI, and confidential computing across the stack.

Nvidia's CEO arrives in South Korea after GTC Taipei with Grace Blackwell shipping and Vera Rubin in full production.

TSMC, Foxconn, Wistron, Pegatron and Inventec apply Nvidia software to the lines building 1 million MGX rack components.

Dell and NVIDIA unveiled Vera Rubin NVL72 systems with 10x lower cost-per-token for agentic AI, backed by 5,000 enterprises in production.
The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.
The briefing read inside teams at