Jack Clark, co-founder of Anthropic and author of the Import AI newsletter, now puts the probability of "no-human-involved AI R&D" by the end of 2028 at 60% or higher. In Import AI 455, published May 4, Clark argues that the engineering pieces required for an AI system to autonomously build its successor are already in place, and that creative research taste is the only remaining gap. He expects a non-frontier proof-of-concept — a model end-to-end training its successor — within a year or two.
The case rests on a stack of benchmark trajectories. SWE-Bench, the GitHub-issue coding test, has moved from a roughly 2% score for Claude 2 in late 2023 to 93.9% for Claude Mythos Preview, a number Clark says effectively saturates the benchmark. For comparison, about 6% of ImageNet validation labels are wrong or ambiguous, suggesting Mythos is now bumping against the noise floor of the test itself.
METR's time-horizon plot tells the second half of the story. The length of task an AI can complete with 50% reliability has gone from ~30 seconds for GPT 3.5 in 2022, to 4 minutes for GPT-4 in 2023, 40 minutes for o1 in 2024, ~6 hours for GPT 5.2 (High) in 2025, and ~12 hours for Opus 4.6 in 2026. Ajeya Cotra of METR thinks ~100-hour tasks are reachable by the end of 2026.
Key facts
- 01Jack Clark puts the probability of fully autonomous AI R&D by end of 2028 at 60%+ in Import AI 455.
- 02Claude Mythos Preview scores 93.9% on SWE-Bench, up from ~2% for Claude 2 in late 2023.
- 03METR's 50%-reliable task horizon went from ~30 seconds (GPT 3.5, 2022) to ~12 hours (Opus 4.6, 2026).
- 04CORE-Bench moved from ~21.5% (GPT-4o, Sept 2024) to 95.5% (Opus 4.5, Dec 2025) — declared solved.
- 05MLE-Bench, covering 75 Kaggle competitions, jumped from 16.9% (o1, Oct 2024) to 64.4% (Gemini3, Feb 2026).
Clark's argument is that this combination — high coding accuracy plus long autonomous horizons — is exactly the profile needed to delegate AI research itself. Cleaning data, launching experiments, reading results, refining kernels: most of the day-to-day labor of an AI researcher fits inside a 12-hour task envelope. "AI systems have gotten good enough to automate a major component of AI R&D, speeding up all the humans that work on it," he writes, adding that most engineers he meets in Silicon Valley now code entirely through AI systems.
“SWE-Bench has gone from a 2% score for Claude 2 in late 2023 to 93.9% for Claude Mythos Preview, effectively saturating the benchmark Jack Clark calls a proxy for coding competency.”— Jaeden Schafer
Two scientific benchmarks back the claim. CORE-Bench, which scores agents on reproducing research-paper results from a code repository, launched in September 2024 with a GPT-4o-based CORE-Agent scoring ~21.5% on the hardest tasks. By December 2025, an Opus 4.5 model hit 95.5%, and one of the benchmark's authors declared it solved.
OpenAI's MLE-Bench, which pits agents against 75 Kaggle competitions across NLP, computer vision and signal processing, tells the same story on a longer arc. The top score at launch in October 2024 was 16.9%, from an o1-based agent. By February 2026, Google's Gemini3 inside a search-equipped harness reached 64.4%.
Kernel design — the unglamorous work of mapping operations like matrix multiplication onto specific hardware — sits closer to the heart of AI R&D and is starting to fall to automation as well. Clark cites work using DeepSeek models to generate better GPU kernels and ongoing efforts to automatically convert PyTorch modules. There's no widely tracked benchmark yet, which is itself a tell: the field is moving faster than its measuring sticks.
Clark is careful about the caveats. Each benchmark has idiosyncratic flaws, and frontier-model self-improvement is materially harder than the non-frontier case because frontier training runs are expensive and reflect thousands of human decisions. Saturating SWE-Bench is not the same as inventing a new architecture, and a 12-hour horizon does not mean a model can sustain a six-month research agenda. "If that happens, we will cross a Rubicon into a nearly-impossible-to-forecast future," Clark writes — a hedge as much as a warning.
The skeptic's read is that benchmark saturation is a story about benchmarks, not about science. Reproducing a paper from its own repo is a tractable engineering task; coming up with the paper is not. Whether models can substitute for human researchers on the creative end — picking which problems matter, designing experiments that haven't been run — is the open question Clark himself flags, and nothing in the public benchmark set settles it.
Still, the trajectory matters because of who is drawing it. Clark sits inside Anthropic and reads the same arXiv, bioRxiv and NBER feeds the labs do, and his 60% number is a public commitment from someone with line of sight into the next two model generations. If he is even directionally right, the most consequential AI release of the next 24 months will not be a chatbot or an agent product — it will be the first model that meaningfully closes the loop on training the next one, and every capex, hiring and policy decision being made today is being made on the wrong side of that line.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




