Skip to main content
Live
Main content

DiG-bench and Faraday mark AI's slow climb toward recursive self-improvement

New benchmarks show frontier models cracking discovery games and replicating research papers, with Opus 5 and a 27B scientist model leading the way.

Jaeden Schafer
Editor in Chief · · 5 min read
DiG-bench and Faraday mark AI's slow climb toward recursive self-improvement

Two new benchmarks published this month sharpen the picture of how close AI models are to autonomous scientific work. DiG-bench, a suite of 70 hidden-rule discovery games, and Replica, a 310-task dataset of research replications drawn from 100 papers, both target the same underlying capability: figuring out what to do when nobody tells you the rules. On DiG-bench's hardest tier, only Claude Opus 5 and Fable 5 scored above zero, at 0.2, while individual humans hit 100%.

DiG-bench comes out of a collaboration among Thinking About Thinking, the University of Oxford, Princeton University, King Abdullah University of Science and Technology, the Swiss AI Lab, Inria, and MIT, with Juergen Schmidhuber among the authors. The games are text-based, handcrafted, and most are kept private so models cannot train on them. Each game presents between 2 and 34 possible actions per step, with rules and objectives hidden until the player deduces them through interaction.

each game is a self-contained miniature world with its own laws, but both the rules and the objective are hidden from the player and must be uncovered through interaction
Juergen Schmidhuber, co-author, DiG-bench

Only 21 of the 70 games have been released publicly. Across the seven-tier structure, Opus 5 and Fable 5 running with Claude Code led the leaderboard, followed by GPT-5.5. Opus 5, GPT-5.5, and Kimi K3 cleared some Tier 6 tasks when equipped with a harness. GLM-5.2 and Gemini 3.1 Pro managed levels in Tier 4. No frontier model came close to human performance on the hardest tier, where the 20% success rate looks especially thin against the 100% humans posted.

Key facts

  • 01DiG-bench spans 70 hidden-rule games across seven tiers; only Opus 5 and Fable 5 scored above zero on Tier 7, at 0.2.
  • 02Faraday, a 27B model post-trained on Qwen-3.6-27B, beat Opus 4.8 and GPT-5.5 on 73% of in-distribution ML replication tasks.
  • 03Faraday scored 60% on held-out AI-for-science tasks drawn from a 310-task dataset built from 100 papers published between 1990 and 2026.
  • 04Individual humans hit 100% on DiG-bench tests where frontier models managed just 20% on the hardest tier.
  • 05Import AI author Jack Clark predicts DiG-bench human parity by mid-2027, which he ties to serious recursive self-improvement.

The point is to isolate discovery as a discrete skill, distinct from pattern-matching or memorized knowledge. If a model can figure out an unfamiliar system by poking at it, that generalizes to research, coding under ambiguous specs, and any environment where the operator does not hand over documentation. Import AI author Jack Clark guesses DiG-bench sees human parity by the middle of 2027, and ties that milestone directly to when recursive self-improvement becomes plausible.

The Faraday project from AI startup Inherent attacks the same question from the research-replication angle. Faraday is a 27B model post-trained on top of Qwen-3.6-27B that acts as a supervisory harness over a larger frontier model, in this case OpenAI Codex. The team assembled a dataset called Replica consisting of 100 ML and AI-for-science papers published between 1990 and 2026, then knocked out individual results to create 310 replication tasks.

For grading, Inherent used Claude Opus 4.7 prompted with a meta-rubric to generate task-specific grading criteria, then a Codex-based judge model to provide reward and per-turn credit assignment for training via a modified GRPO. Faraday beat standard Opus 4.8 and GPT-5.5 on 73% of in-distribution ML tasks and 60% of held-out AI-for-science tasks. The pattern is a small, cheap outer agent steering a big, expensive frontier model to better outcomes than the frontier model alone.

The skills that allow Faraday to fill in vaguely-specified details may be the very same skills that would allow it to advance the state of the art by designing its own experiment
Inherent, AI startup building Faraday

Inherent frames the implication directly. The company argues that scoping experiments to a budget, deciding what to investigate, and judging replications are the same skills required to design original experiments — and that a post-trained outer agent should be able to track the frontier as underlying coding models improve. That is a specific, testable claim about how quickly AI research assistance becomes AI research.

Against this technical backdrop, Meta CEO Mark Zuckerberg published an essay titled "The Future is for Everyone" laying out Meta's AI philosophy. The core proposition is proliferation: distribute AI capabilities widely to prevent concentration of power. Zuckerberg's framing puts individual empowerment, invention, and balance of power at the center, with an exceptionally capable personal agent for every user as the deliverable.

The defining questions of our age are who will have access to superintelligence and what will we direct it towards
Mark Zuckerberg, Meta CEO
Related · from this week
Inherent's Faraday agent beats Claude and GPT-5.5 at replicating research on a 27B model
Jaeden Schafer · 5 min read →

What the essay does not address is what superintelligent systems capable of independent invention might themselves want to do — a gap Clark flags directly. If DiG-bench and Faraday are read together, the trajectory they sketch is one where models increasingly set their own research agendas, and the question of alignment stops being abstract.

The counterweight is that current numbers still show real gaps. A 20% Tier 7 score against 100% human performance is a wide margin, and Faraday's uplift, while real, comes from a narrow rubric-based judge on a specific dataset. Held-out generalization at 60% is progress, not parity. And every claim about recursive self-improvement remains a projection, not a demonstrated capability.

What the two benchmarks change is the resolution of the debate. Instead of arguing over whether AI can "do science," the conversation now runs on specific tier scores, specific task-completion rates, and specific replication rubrics. That is how a field converges on a milestone — by naming it, measuring it, and watching the number climb. On current slopes, the interesting inflection point is closer than the labs building toward it have publicly acknowledged, and the companies that ship the first genuinely autonomous research agent will reshape the economics of every downstream science-heavy business.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

Inherent's Faraday agent beats Claude and GPT-5.5 at replicating research on a 27B model
Models

Inherent's Faraday agent beats Claude and GPT-5.5 at replicating research on a 27B model

The London lab, fresh off a $50M seed, says its DeepMind-alumni-built agent matches frontier systems using a fraction of the parameters.

Jaeden Schafer5 min read
Nvidia logo
Models

Nvidia says the harness, not the model, drives Claude Opus 5 to 100% on ARC-AGI-3

New Nvidia research shows a custom harness with a supervisor agent lifted Claude Opus 5 from 30% to a perfect score on the interactive reasoning benchmark.

Jaeden Schafer5 min read
Meta logo
Models

Meta launches Muse Spark 1.1 to challenge Claude and GPT-5.6 on coding

The agentic coding model prices at $1.25 per million input tokens, slightly above Claude Haiku 4.5 and GPT-5.6 Luna.

Jaeden Schafer4 min read