NVIDIA and Ineffable Intelligence announced an engineering-level partnership on May 13, 2026 to codesign reinforcement learning infrastructure, putting Jensen Huang's chip roadmap directly behind David Silver's bet that the next jump in AI capability comes from systems that learn by trial and error rather than from scraped human text. Ineffable, the London-based lab Silver founded, came out of stealth one week earlier. The work begins on NVIDIA Grace Blackwell and will be among the first deployments on the upcoming Vera Rubin platform.
Silver is one of the architects of AlphaGo, the system that beat the world's top Go players using reinforcement learning. His new lab is structured around a single thesis: that pretraining on human-generated data is approaching its ceiling, and the harder problem — systems that discover new knowledge on their own — requires a fundamentally different compute pipeline.
"The next frontier of AI is superlearners — systems that learn continuously from experience," Huang said in NVIDIA's announcement. "We are thrilled to partner with Ineffable Intelligence to codesign the infrastructure for large-scale reinforcement learning as they push the frontier of AI and pioneer a new generation of intelligent systems."
Key facts
- 01NVIDIA announced a codesign partnership with Ineffable Intelligence on May 13, 2026, one week after the London lab exited stealth.
- 02Ineffable Intelligence was founded by David Silver, the AlphaGo architect and one of the pioneers of reinforcement learning.
- 03The collaboration starts on NVIDIA Grace Blackwell and will be among the first workloads to run on the upcoming Vera Rubin platform.
- 04Jensen Huang framed the partnership around 'superlearners' — systems that learn continuously from experience rather than fixed datasets.
- 05NVIDIA's next platform event runs June 1-2 in Taipei.
Silver framed the technical problem in the same announcement. "Researchers have largely solved the easier problem of AI: how to build systems that know all the things humans already know," he said. "But now we need to solve the harder problem of AI: how to build systems that discover new knowledge for themselves. That requires a very different approach — systems that learn from experience."
“Researchers have largely solved the easier problem of AI: how to build systems that know all the things humans already know. But now we need to solve the harder problem.”— Jaeden Schafer
The infrastructure demands diverge sharply from pretraining. In a pretraining run, a fixed dataset of human-generated text and images flows through the system in predictable batches. Reinforcement learning workloads generate their training data on the fly: the agent acts in an environment, observes the result, scores it, and updates — over and over, in tight loops. That pattern stresses interconnect bandwidth, memory subsystems, and serving infrastructure in ways pretraining does not.
The data itself is also different. Rich forms of experience drawn from simulation and interaction look nothing like tokenized human language, which is why NVIDIA and Ineffable expect the work to surface new model architectures and new training algorithms alongside the hardware codesign. Engineers from both sides are now working on the pipeline together.
Starting on Grace Blackwell gives the joint team current-generation silicon to validate workloads, while Vera Rubin gives them a runway to influence the next platform before it ships at scale. For NVIDIA, having a frontier reinforcement learning customer stress-testing Rubin early is how the company has historically locked in the next cycle of hyperscaler demand.
For Ineffable, the value of the partnership is access. Reinforcement learning at frontier scale is not something a stealth-stage lab can build alone; the tight feedback loops between actor, environment, and learner demand interconnect and memory engineering at the rack and cluster level. Codesigning with NVIDIA is the fastest path to that.
The unknown is whether the underlying bet holds. The view that pretraining is hitting diminishing returns is not universal — OpenAI, Anthropic, and Google continue to scale pretraining-heavy systems alongside their own reinforcement learning work, and none has conceded that the human-data paradigm is exhausted. Ineffable has not published a model, a benchmark, or a research paper yet. The lab's case rests on Silver's track record and a directional thesis.
NVIDIA's next platform showcase is June 1-2 in Taipei, where more detail on Vera Rubin and the surrounding software stack is expected. The same announcement window referenced AI Hermes, a separate effort on self-improving agents running on NVIDIA RTX PCs and DGX Spark, suggesting reinforcement learning infrastructure is becoming a deliberate product line rather than a one-off research collaboration.
The strategic read for the AI market is that NVIDIA is hedging its dependence on the pretraining cycle. Frontier labs spending tens of billions on pretraining runs is a great business, but it is one business. By embedding with Silver early, NVIDIA positions Rubin and its successors as the default substrate for whichever paradigm wins next — and ensures that if reinforcement learning at scale is the answer, the silicon was designed for it from the start.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




