Skip to main content
Live
Main content

Jack Clark: 60% chance AI builds its own successor by end of 2028

Anthropic co-founder lays out the benchmark trail — SWE-Bench from 2% to 93.9%, METR horizons from 30 seconds to 12 hours — and calls the takeoff.

Jaeden Schafer
Editor in Chief · · 5 min read
Anthropic logo

Jack Clark, co-founder of Anthropic and author of the Import AI newsletter, now puts the probability of "no-human-involved AI R&D" by the end of 2028 at 60% or higher. In Import AI 455, published May 4, Clark argues that the engineering pieces required for an AI system to autonomously build its successor are already in place, and that creative research taste is the only remaining gap. He expects a non-frontier proof-of-concept — a model end-to-end training its successor — within a year or two.

The case rests on a stack of benchmark trajectories. SWE-Bench, the GitHub-issue coding test, has moved from a roughly 2% score for Claude 2 in late 2023 to 93.9% for Claude Mythos Preview, a number Clark says effectively saturates the benchmark. For comparison, about 6% of ImageNet validation labels are wrong or ambiguous, suggesting Mythos is now bumping against the noise floor of the test itself.

METR's time-horizon plot tells the second half of the story. The length of task an AI can complete with 50% reliability has gone from ~30 seconds for GPT 3.5 in 2022, to 4 minutes for GPT-4 in 2023, 40 minutes for o1 in 2024, ~6 hours for GPT 5.2 (High) in 2025, and ~12 hours for Opus 4.6 in 2026. Ajeya Cotra of METR thinks ~100-hour tasks are reachable by the end of 2026.

Key facts

  • 01Jack Clark puts the probability of fully autonomous AI R&D by end of 2028 at 60%+ in Import AI 455.
  • 02Claude Mythos Preview scores 93.9% on SWE-Bench, up from ~2% for Claude 2 in late 2023.
  • 03METR's 50%-reliable task horizon went from ~30 seconds (GPT 3.5, 2022) to ~12 hours (Opus 4.6, 2026).
  • 04CORE-Bench moved from ~21.5% (GPT-4o, Sept 2024) to 95.5% (Opus 4.5, Dec 2025) — declared solved.
  • 05MLE-Bench, covering 75 Kaggle competitions, jumped from 16.9% (o1, Oct 2024) to 64.4% (Gemini3, Feb 2026).

Clark's argument is that this combination — high coding accuracy plus long autonomous horizons — is exactly the profile needed to delegate AI research itself. Cleaning data, launching experiments, reading results, refining kernels: most of the day-to-day labor of an AI researcher fits inside a 12-hour task envelope. "AI systems have gotten good enough to automate a major component of AI R&D, speeding up all the humans that work on it," he writes, adding that most engineers he meets in Silicon Valley now code entirely through AI systems.

SWE-Bench has gone from a 2% score for Claude 2 in late 2023 to 93.9% for Claude Mythos Preview, effectively saturating the benchmark Jack Clark calls a proxy for coding competency.
Jaeden Schafer

Two scientific benchmarks back the claim. CORE-Bench, which scores agents on reproducing research-paper results from a code repository, launched in September 2024 with a GPT-4o-based CORE-Agent scoring ~21.5% on the hardest tasks. By December 2025, an Opus 4.5 model hit 95.5%, and one of the benchmark's authors declared it solved.

OpenAI's MLE-Bench, which pits agents against 75 Kaggle competitions across NLP, computer vision and signal processing, tells the same story on a longer arc. The top score at launch in October 2024 was 16.9%, from an o1-based agent. By February 2026, Google's Gemini3 inside a search-equipped harness reached 64.4%.

Kernel design — the unglamorous work of mapping operations like matrix multiplication onto specific hardware — sits closer to the heart of AI R&D and is starting to fall to automation as well. Clark cites work using DeepSeek models to generate better GPU kernels and ongoing efforts to automatically convert PyTorch modules. There's no widely tracked benchmark yet, which is itself a tell: the field is moving faster than its measuring sticks.

Clark is careful about the caveats. Each benchmark has idiosyncratic flaws, and frontier-model self-improvement is materially harder than the non-frontier case because frontier training runs are expensive and reflect thousands of human decisions. Saturating SWE-Bench is not the same as inventing a new architecture, and a 12-hour horizon does not mean a model can sustain a six-month research agenda. "If that happens, we will cross a Rubicon into a nearly-impossible-to-forecast future," Clark writes — a hedge as much as a warning.

Related · from this week
Anthropic's Amodei lays out a three-part plan to slow AI progress
Jaeden Schafer · 5 min read →

The skeptic's read is that benchmark saturation is a story about benchmarks, not about science. Reproducing a paper from its own repo is a tractable engineering task; coming up with the paper is not. Whether models can substitute for human researchers on the creative end — picking which problems matter, designing experiments that haven't been run — is the open question Clark himself flags, and nothing in the public benchmark set settles it.

Still, the trajectory matters because of who is drawing it. Clark sits inside Anthropic and reads the same arXiv, bioRxiv and NBER feeds the labs do, and his 60% number is a public commitment from someone with line of sight into the next two model generations. If he is even directionally right, the most consequential AI release of the next 24 months will not be a chatbot or an agent product — it will be the first model that meaningfully closes the loop on training the next one, and every capex, hiring and policy decision being made today is being made on the wrong side of that line.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Analysis

Anthropic logo
Analysis

Anthropic's Amodei lays out a three-part plan to slow AI progress

Dario Amodei wants embedded third-party evaluators, coordinated safety standards across US labs, and narrow deals with China.

Jaeden Schafer5 min read
AIUC raises $40M Series A to certify enterprise AI agents against rogue behavior
Security

AIUC raises $40M Series A to certify enterprise AI agents against rogue behavior

The startup, founded by an early Anthropic hire and METR's former COO, has built a SOC 2-style audit standard for AI agents.

Jaeden Schafer5 min read
Anthropic logo
Models

Anthropic logs 8x code merge jump as researchers benchmark AI gaming society's rules

Anthropic sees early signs of recursive self-improvement, a new benchmark tests AI loophole-hunting, and RL drones beat a human champion at 22 m/s.

Jaeden Schafer5 min read