Nvidia's Nemotron 3 Ultra matched the highest-scoring closed models on LangChain's Deep Agents benchmark at 10x lower inference cost per run, after LangChain tuned its Deep Agents harness for the open model. The result, disclosed by Nvidia on July 8, 2026, was achieved without retraining Nemotron 3 Ultra — every gain came from engineering the environment around the model. It's the sharpest public data point yet for the argument that open models plus a well-tuned agent harness can compete head-on with closed frontier systems on enterprise workloads.
LangChain, whose agent engineering platform draws more than 200 million monthly downloads, ran Nemotron 3 Ultra against its public Deep Agents benchmark suite, analyzed the execution traces to identify where the agent lost points, and then adjusted system prompts, tool descriptions and middleware. The tuned harness delivered the highest accuracy among open models, completed more tasks at higher throughput than the closed field, and hit business-task parity with the top closed scorers. At a tenth of the running cost, teams can now run evaluations continuously and iterate faster on specialized agents.
The reference implementation shipped alongside the benchmark result. Nvidia NemoClaw for LangChain Deep Agents packages the tuned Deep Agents Code with the Nvidia OpenShell secure runtime for executing agent actions, and is available now. Developers can pull the tuned profile directly from LangChain or start from the NemoClaw blueprint. Hosted access to Nemotron 3 Ultra with the tuned harness is live on Baseten, Crusoe Cloud, DeepInfra, Fireworks, Nebius and Together AI.
Key facts
- 01Nemotron 3 Ultra matched top closed models on LangChain's Deep Agents benchmark at 10x lower inference cost per run.
- 02No model retraining was required — every gain came from tuning the harness around the model.
- 03LangChain's agent engineering platform has more than 200 million monthly downloads.
- 04Abridge, Amdocs, Box and EY are already building on the NemoClaw for LangChain Deep Agents blueprint.
- 05The tuned Deep Agents profile is available now via Baseten, Crusoe Cloud, DeepInfra, Fireworks, Nebius and Together AI.
Harrison Chase, LangChain's cofounder and chief executive, framed the work as validation of a harness-first approach to agent quality rather than another round of model scaling.
Enterprise adopters are already in flight. Abridge, Amdocs and Box are embedding specialized agents built on the stack into their platforms, and global systems integrator EY is expanding its Nvidia implementation practice around NemoClaw blueprints to help clients customize, evaluate and govern agents in high-value workflows. Nvidia founder and chief executive Jensen Huang recently sat down with Chase to discuss why the last six months have produced a step-change in useful enterprise AI — a claim the benchmark numbers now underwrite.
The pitch to buyers is ownership. An open model, an open harness and an open secure runtime means the enterprise controls the full stack: they can run it on their own infrastructure, in their own cloud, under their own governance, and keep tuning it around workflows that differentiate the business. That's a materially different posture from renting agent capability from a closed API vendor, particularly as agents move from answering questions to taking action inside core systems.
The cost gap matters most for the workloads that were previously uneconomic. A 10x reduction in inference cost per run turns continuous evaluation, exhaustive test-time search and always-on background agents from budget-line items into operational defaults. It also changes the calculus for smaller teams that couldn't afford to run a closed frontier model in a loop against every ticket, every contract, every log file.
There are caveats. LangChain's Deep Agents benchmark is one test suite among several, and parity with closed models on business tasks doesn't mean parity on every category — coding, multimodal reasoning and long-horizon planning have separate leaderboards where the closed labs still lead. The harness advantage also compounds only for teams that actually invest in tool descriptions, memory design and middleware; buyers looking for a single-API drop-in will find the open-stack story more work, not less. And Nvidia's own hardware position means the company's incentives around promoting open-model economics are not disinterested.
For the AI agent market, the Nemotron-plus-LangChain result narrows a question that has hung over the space for a year: can an open stack ship agents that enterprises will actually deploy against closed alternatives from OpenAI, Anthropic and Google? The answer, at least on LangChain's benchmark and at a tenth of the cost, is yes — and the tuned profile is a download away rather than a fine-tuning project. That shifts leverage toward customers who want to own their agent stack end to end, and it raises the pressure on closed-model providers to justify their premium on something other than raw benchmark scores.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




