Nvidia and Microsoft used the Microsoft Build stage on June 2, 2026 to roll out a single agentic AI stack spanning Windows laptops, Azure data centers, and on-prem hardware. The headline pieces: RTX Spark Windows PCs delivering 1 petaflop of AI performance with up to 128GB of unified memory, a DGX Station for Windows with 748GB of coherent memory and 20 petaflops of FP4, and GPU-accelerated Microsoft Fabric clocking SQL queries 6x faster than its CPU baseline. Jensen Huang joined Satya Nadella's keynote via livestream from Taipei to frame the expansion.
The pitch is that agentic AI needs the full stack — silicon, runtime, data layer, and models tuned for long-running reasoning — and that Microsoft and Nvidia intend to supply all of it from the same vendor pair. RTX Spark laptops and small desktops are positioned as the first Windows PCs purpose-built for personal agents, shipping this fall from Microsoft Surface, ASUS, Dell, HP, Lenovo, and MSI. The hardware carries Nvidia's accumulated CUDA, RTX, DLSS, and TensorRT software layers.
At the higher end, DGX Station for Windows is built on the Nvidia GB300 Grace Blackwell Ultra Desktop Superchip and targets frontier models up to 1 trillion parameters running locally for always-on enterprise agents. ASUS, Dell, GIGABYTE, HP, MSI, and Supermicro are expected to ship systems in Q4 2026. Both product lines run Nvidia OpenShell, a sandboxed runtime for autonomous agents.
“RTX Spark is a new beginning, powering the world's first Windows PCs purpose-built for personal agents, with 1 petaflop of AI performance, up to 128GB of unified memory, all-day battery life, and full AI and graphics performance unplugged.”— Jensen Huang, NVIDIA founder and CEO
Key facts
- 01RTX Spark Windows PCs deliver 1 petaflop of AI performance with up to 128GB unified memory, shipping this fall from Surface, ASUS, Dell, HP, Lenovo, and MSI.
- 02DGX Station for Windows packs 748GB of coherent memory and 20 petaflops of FP4, capable of running 1-trillion-parameter models, with systems in Q4 2026.
- 03Anthropic's Claude models will run natively on NVIDIA GB300 Blackwell Ultra systems on Azure in the weeks ahead.
- 04GPU-accelerated Microsoft Fabric Data Warehouse benchmarked 6x faster SQL than CPU baseline and 7x faster than three rival cloud warehouses.
- 05Microsoft's Fairwater Wisconsin AI factory went live early with hundreds of thousands of NVIDIA Grace Blackwell systems running as one cluster.
On the model layer, Microsoft Foundry is becoming a multi-vendor agent catalog. Anthropic's Claude models will run natively on Nvidia GB300 Blackwell Ultra systems on Azure within weeks, joining OpenAI and Nvidia's own Nemotron family in Foundry Agent Service. Nvidia Nemotron 3 Ultra, a new open reasoning model for long-running coding and research workflows, is landing on Foundry managed compute this month alongside Nemotron 3.5 ASR and Nemotron 3.5 Content Safety.
Nvidia is also pushing Cosmos 3 — its open omnimodel for physical AI covering vision reasoning, world simulation, and action generation — into Microsoft's Physical AI Toolchain. Earth-2 weather models will be available through Microsoft Planetary Computer Pro and Foundry for enterprise forecasting. The Nvidia Agent Toolkit and NemoClaw blueprints, which the company shipped to the edge earlier this month, give developers an open-source path to production agents on Foundry.
The data warehouse numbers are the sharpest competitive claim. Microsoft's internal benchmarks show GPU-accelerated Fabric Data Warehouse running SQL execution up to 6x faster than its CPU-powered baseline and up to 7x faster than three other leading cloud data warehouse providers under high-concurrency workloads. That matters because agents query data continuously, and warehouse latency has been a quiet ceiling on agentic throughput.
Security is the other half of the agent story. Nvidia OpenShell is being integrated into GitHub Copilot, isolating each agent in its own sandboxed container and evaluating every outbound call against policy before it touches files, networks, or credentials. Policies are written as code, versioned in the repo, and the runtime is open-source under Apache 2.0 across on-prem, hybrid, and cloud.
“As agents move from coding assistance to autonomous execution, they need real capability without real credentials.”— Jensen Huang, NVIDIA founder and CEO
For workloads that can't move to public cloud, Microsoft is bringing Foundry Local on Azure Local to the Nvidia RTX PRO 6000 Blackwell Server Edition platform, paired with Nemotron models. Foundry Local now supports multinode deployments and the vLLM runtime for manufacturing, energy, and sovereign data center scenarios. The same model family runs from a developer's RTX Spark laptop to a sovereign on-prem cluster.
Underneath all of it sits Microsoft's Fairwater Wisconsin AI factory, which went live ahead of schedule running hundreds of thousands of Nvidia Grace Blackwell systems as a single AI factory, linked to a sister facility in Georgia. Microsoft has also validated the Nvidia Vera Rubin platform for Azure deployment, with Nvidia citing up to 10x inference throughput per megawatt versus Blackwell and an order-of-magnitude reduction in cost per agentic token. Vera Rubin slots into existing Blackwell racks without retrofits.
The unanswered questions are mostly about choice. A unified stack from one vendor pair is fast to deploy but tightens the dependency: enterprises adopting Foundry, Fabric, OpenShell, and Nemotron together will find it hard to swap any one layer for AMD silicon, a different runtime, or a non-Microsoft data plane. Pricing for RTX Spark systems and DGX Station for Windows was not disclosed, and the 6x and 7x Fabric numbers come from Microsoft's own benchmarks rather than an independent test.
The strategic read is that Nvidia and Microsoft are racing to define what "agentic infrastructure" looks like before AWS, Google, and the open-source ecosystem coalesce around an alternative. By spanning a $1,000-class laptop to a Vera Rubin data center on one toolchain, they're betting developers will pick the path of least friction — and that the lock-in pays off across every deployment tier. For enterprises already standardized on Azure and Windows, the calculus just got simpler; for everyone else, the pressure to match this stack just got harder.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.



