Nvidia is reframing the AI infrastructure conversation around a single metric: performance per watt. In a technical post published July 14, the company said its GB300 NVL72 rack-scale system delivers up to 25x the performance per watt of the prior Hopper generation on DeepSeek V4 Pro, 20x on GLM5.1, and 10x on Kimi K2.6, citing SemiAnalysis InferenceX benchmarks. The argument is blunt — in a power-constrained buildout, tokens per watt determine revenue per rack.
The pitch matters because power, not silicon supply, is now the binding constraint for AI factory operators. Nvidia estimates that in a typical AI factory, only about 60% of the electricity pulled from the grid converts into useful AI work, with the rest lost to cooling and rack-level inefficiencies. That gap is the target of a new software layer called Nvidia DSX MaxLPS, which shifts power between GPUs and racks in real time and, according to Nvidia, allows operators to run up to 40% more GPUs within the same power budget.
“Power is AI infrastructure's inescapable constraint. How many tokens an AI factory can generate within a fixed power budget determines its revenue and profitability.”— Shruti Koparkar, Nvidia
The performance gains lean heavily on codesign across the stack. Nvidia's inference software — Dynamo and TensorRT LLM alongside open-source stacks SGLang and vLLM — supports NVFP4 quantization, disaggregated serving, large-scale expert parallelism, KV-aware routing and KV cache offloading. Those optimizations compound over time: Nvidia says performance per watt on DeepSeek V4 improved by up to 5x in a single month on the same hardware.
Key facts
- 01Nvidia's GB300 NVL72 delivers up to 25x performance per watt over Hopper on DeepSeek V4 Pro, 20x on GLM5.1, and 10x on Kimi K2.6.
- 02Software optimizations improved performance per watt on DeepSeek V4 by up to 5x in a single month.
- 03Only about 60% of grid electricity converts to useful AI work in typical AI factories, per Nvidia.
- 04Nvidia DSX MaxLPS lets operators run up to 40% more GPUs within the same power budget by shifting power in real time.
- 05Anthropic and OpenAI use Nvidia [Blackwell](/openai) NVL72 systems to run inference, alongside CoreWeave, Perplexity, and Fireworks AI.
Rack-scale networking is doing much of the work. The Nvidia NVLink Switch, now in its sixth generation with the forthcoming Vera Rubin platform, is purpose-built for scale-up GPU domains rather than adapted from general-purpose networking. It supports SHARP, which performs in-network computing directly in the switch and offloads collective operations from the GPUs themselves — a design choice that gets more valuable as mixture-of-experts models push routing traffic higher.
Nearly every frontier model in production today uses a mixture-of-experts architecture, and serving those models at rack scale is a different engineering problem from single-node inference. Failure modes multiply, latency budgets tighten, and load-balancing across experts becomes a system-level concern rather than a model-level one.
“Serving MoE at rack scale demands codesign across every layer of the system and software stack, plus the operational depth earned from running these models under real production load.”— Shruti Koparkar, Nvidia
The customer roster Nvidia points to is the strongest part of the argument. Anthropic and OpenAI both use Blackwell NVL72 systems for inference, according to the post. CoreWeave has deployed Kimi K2.6 on GB300 NVL72, pairing NVFP4 quantization with EAGLE3 speculative decoding. Perplexity runs Qwen3 at 235 billion parameters and a post-trained Qwen3.5-397B-A17B on GB200 NVL72, serving millions of queries daily. Fireworks AI runs GLM 5.2 on Blackwell for customers including Cursor and Factory AI.
Those deployments matter because rack-scale reliability is not a spec-sheet claim. Systems at this density surface failure modes that single-node clusters never encounter, and the operational playbooks for handling them accumulate only through months of production traffic. Nvidia's implicit argument is that competitors selling raw FLOPS will spend years catching up on the software and operational layers that turn silicon into sustained tokens.
The metric shift also changes how buyers should think about competing accelerators. Peak FLOPS and memory bandwidth remain useful specs, but in a data center where the utility can only deliver a fixed megawatt allocation, the question is how many tokens leave the building per kilowatt-hour. That framing favors vertically integrated platforms — silicon, interconnect, cooling and orchestration software co-designed — over point solutions that win on a single benchmark.
The counterweight is that Nvidia is grading its own homework. The 25x, 20x and 10x figures come from SemiAnalysis InferenceX runs curated by Nvidia, on models Nvidia's software has been optimized against. Independent replication on production workloads at other operators will settle whether the gains hold outside the reference configurations, and rival accelerator vendors are pursuing their own performance-per-watt claims on the same models. The five-fold monthly improvement on DeepSeek V4 also cuts both ways — impressive velocity, but a reminder that today's numbers are a snapshot in a fast-moving software cycle.
For the AI market, the strategic read is that Nvidia is trying to move the goalposts before Vera Rubin ships. If buyers evaluate infrastructure on tokens per watt rather than dollars per GPU, the incumbent's codesign advantage compounds and the switching cost for hyperscalers considering custom silicon or AMD's MI400 line goes up. The company hosts GTC Berlin from October 20 to 22, where the Vera Rubin details will get their public airing — and where the next round of performance-per-watt claims will land.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




