Nvidia's DSX platform lifted token throughput 24% on a Lambda GPU cluster running inside a fixed power budget, the first production validation of software Nvidia is positioning as the answer to AI's grid problem. Lambda ran 19 HGX B200 nodes at 85% of full power and matched the electricity draw of 16 nodes at 100%, pushing throughput from roughly 4 million tokens per second to 5 million. Performance per watt improved 23%. The results, disclosed at the AI Infra Summit on Tuesday, put concrete numbers on a pitch Nvidia has been making since GTC Taipei in May: the way to build bigger AI factories is not to wait for new transmission lines, but to extract more useful work from every watt already flowing.
The demonstration is one piece of a broader stack Nvidia calls DSX — a suite spanning MaxLPS for dynamic power allocation, DSX Flex for grid-signal response, DSX OS for lifecycle management, DSX Sim for pre-deployment modeling, and reference designs for compute, networking, and cooling. MaxLPS monitors GPU and rack-level draw in real time and reallocates headroom across nodes based on whether they are training or serving inference, recovering capacity that static provisioning strands. Lambda, which serves more than 10,000 cloud customers, tested it on a five-rack, 19-node cluster.
Nvidia projects the same technique can unlock up to 40% more GPU capacity for its next-generation Vera Rubin NVL72 factories within an unchanged megawatt budget. That figure matters because power, not silicon, is now the binding constraint on hyperscaler expansion. Utility interconnection queues stretch years, and a fixed-cost AI factory that can do 40% more work is functionally a 40% cheaper one.
Key facts
- 01Lambda ran 19 HGX B200 nodes at 85% power inside the same budget as 16 nodes at full draw, hitting 5M tokens per second versus 4M — a 24% throughput gain.
- 02Performance per watt improved 23% on the same five-rack cluster, according to results Lambda released at the AI Infra Summit.
- 03Nvidia projects DSX MaxLPS can unlock up to 40% more GPU capacity for Vera Rubin NVL72 factories within the same megawatt budget.
- 04Silicon Valley Power has sent 200+ demand signals to Nvidia's Eos AI factory, cutting draw from 4MW to 3MW in under a minute each time without dropping jobs.
- 05Nvidia's first dedicated DSX Flex deployment will be a 96-megawatt Vera Rubin AI factory in Manassas, Virginia.
Lambda framed the result as a break from static power planning.
The other half of the pitch is grid participation. On a hot August evening in Silicon Valley, Silicon Valley Power sent a signal to Nvidia's Eos AI factory in Santa Clara asking it to shed load. Emerald AI's Conductor platform — running as part of SVP's Flexible Load Interconnect Program, the first commercial utility program that treats AI factories as dispatchable resources — dropped the factory's draw from 4 megawatts to 3 in under a minute. Lowest-priority jobs yielded; high-priority inference kept running. No operator touched anything.
Silicon Valley Power has since sent more than 200 such signals to the same facility. Every one worked. Varun Sivaram, whose team built Conductor, watched the first execution over Zoom with about forty engineers from Emerald AI, the data center, and the utility. "We were watching with bated breath," he said. "It was our first time deploying across thousands of NVIDIA GPUs." His head of product, Mansi Shah, compared it to a SpaceX launch.
Emerald AI's Conductor is set to integrate into DSX Flex as the platform matures. The first dedicated DSX Flex deployment will be a 96-megawatt Vera Rubin AI factory at Nvidia's AI Factory Research Center in Manassas, Virginia, building on five prior demonstrations across two continents. The commercial logic is straightforward: a facility that can throttle 40% of its draw on a utility signal gets to interconnect faster and at a bigger nameplate than one that cannot.
“A one-gigawatt factory will never become a two-gigawatt factory.”— Jensen Huang, NVIDIA founder and CEO
Underneath the software, Nvidia is also changing how power reaches the rack. The company is moving to an 800-volt DC power architecture, up from the 54V distribution common today, and folding it into DSX reference designs. The projected end-to-end efficiency gain is 3 to 5%, and the architecture ships with Vera Rubin NVL72 in 2027. The bigger point is that GB200 NVL72 racks with direct liquid cooling already carry roughly 120 kW of heat per rack — every conversion step and cooling overhead in the path from utility to GPU is now a first-order design problem.
The skeptics' case is that these are Nvidia's own numbers on Nvidia's own hardware, and that grid participation only works where utilities have programs like SVP's — which is nearly nowhere else in the United States. Emerald AI's Conductor has been proven at one facility so far. Whether utilities in Northern Virginia, Texas, or Arizona will build flexible-load tariffs quickly enough to matter is a policy question, not an engineering one, and the interconnection queue is not going to wait.
Still, the direction of travel is clear. Nvidia has been telling investors and customers for two years that compute-per-watt is the metric that will define the AI infrastructure buildout, and DSX is the first coherent product answer to it. Jensen Huang's line — that a one-gigawatt factory will never become a two-gigawatt factory — is a way of telling operators to stop waiting for grid upgrades and start extracting more tokens from the megawatts they already have permits for. If Lambda's 24% holds up across more deployments, the economics of every proposed AI factory get rewritten on the same slide.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.



