Skip to main content
Live
Main content

Intology's Locus beats human baseline on PostTrainBench, signaling AI R&D automation

Locus hits 51.6% on PostTrainBench+ with 4,000+ H100 hours, surpassing the 51.1% human baseline and every frontier-agent competitor.

Jaeden Schafer
Editor in Chief · · 5 min read
Anthropic logo

Intology's Locus system scored 51.6% on PostTrainBench+, edging past the 51.1% human baseline and beating Anthropic's Opus 4.8 at 44.3% and GLM 5.2 at 42.7%. The result required over 4,000 hours of H100 GPU time and marks the first time an automated agent has exceeded the human bar on the benchmark, which measures how well AI systems can take an open-weight model and improve its performance. Intology, whose stated goal is to automate R&D, released the numbers this week alongside externally verified results from the PostTrainBench authors.

On the standard PostTrainBench, Locus paired with Claude Opus 5 scored 44.7%, versus 34.1% for Opus 5 running without the harness and 41.8% for Fable 5. The 10-point absolute jump from adding a research harness is the more revealing number: it suggests frontier models are being significantly under-elicited when run bare. PostTrainBench+ removes the 10-hour wall-clock limit on a single GPU that constrains the base benchmark, allowing systems to spend far more compute per problem.

The trajectory on PostTrainBench itself is steep. When the benchmark launched in March 2026, the top score was Opus 4.6 at 23.2%, up from Claude Sonnet 4.5 at 9.9% in September 2025. Locus at 44.7% roughly doubles the March state of the art in five months. Intology says the results underwent stringent contamination and cheating checks.

outperforms every frontier-agent baseline on PostTrainBench, and given greater compute, post-trains models that collectively surpass both the baselines and the official human instruction-tuned Qwen3-1.7B release across the benchmark suite
Intology, AI research startup

Key facts

  • 01Intology's Locus scored 44.7% on PostTrainBench using Opus 5, up from 34.1% for Opus 5 without the harness and 41.8% for Fable 5.
  • 02On PostTrainBench+, Locus hit 51.6% with 4,000+ hours of H100 time, exceeding the 51.1% human baseline and beating Opus 4.8 at 44.3%.
  • 03PostTrainBench scores have jumped from 9.9% for Claude Sonnet 4.5 in September 2025 to 23.2% for Opus 4.6 in March 2026 to 44.7% now.
  • 04For no-code startup Bubble, Locus trained a production language model with 2.8× lower error, 5.4× lower latency, and 105× lower cost.
  • 05Think tank IFP published 23 policy recommendations across 7 categories aimed at helping governments respond to automated AI R&D.

Locus is not confined to benchmark work. Intology says the system discovered and trained a language model end-to-end for no-code app builder Bubble that now runs in production at 2.8× lower error, 5.4× lower latency, and 105× lower cost than the prior model. That is the kind of number that turns automated R&D from a research curiosity into a procurement decision.

The broader claim from Jack Clark's Import AI newsletter is that AI systems are capable of substantially more AI research work than most operators assume, and that the human baseline on PostTrainBench v1.1 will fall before the end of 2026. Given that Locus already crossed it on the extended-compute variant, that prediction now looks conservative for the base benchmark too.

The policy side of the same conversation is moving in parallel. Think tank IFP published 23 specific recommendations across 7 categories aimed at helping policymakers, especially in the United States, respond to accelerating AI R&D automation. The categories include transparency into automated AI R&D, state capacity to understand it, verification technology, AI resilience, and international coordination options.

it's as if the world is driving AI development in a car that only has an accelerator pedal and no brake pedal, let alone any kind of sophisticated telemetry for knowing things ranging from the speed of the car to the properties of the engine to the wear on the tires
Jack Clark, Import AI author and Anthropic co-founder

Separately, researchers at MIT and Columbia published Racing to Ruin, a game-theoretic analysis of whether frontier firms can coordinate a slowdown. Their model treats scaling as an activity that raises the hazard of an event that permanently drives all firms' payoffs to zero, and finds that stable coordination equilibria require both transparency about technology development and trust that rivals are rational actors. With low trust, every equilibrium races to ruin. With high trust, the probability of two rational firms racing forever vanishes quadratically.

The paper also flags that transparency alone is double-edged. Faster detection makes it cheaper for a firm to wait for confirmation that a rival has stopped before stopping itself, which can destroy early-stopping equilibria at intermediate trust levels before restoring them once detection becomes fast enough to make stopping self-enforcing. The parallels to nuclear arms verification regimes are explicit.

Related · from this week
Anthropic ships Fable and Mythos 5.1 with cheaper tokens and looser guardrails
Jaeden Schafer · 5 min read →

The caveat on the Locus numbers is compute cost. Beating the human baseline required more than 4,000 H100-hours on a single benchmark run, which is not a workflow any lab will deploy casually. The PostTrainBench+ variant was designed specifically to test what happens when the wall-clock ceiling is removed, so the human-beating score is a demonstration of what compute buys, not a claim about efficient automated research. Fable was not tested on the plus variant, leaving one competitive gap unmeasured.

The direction of travel matters more than any single score. A benchmark that ran at 9.9% eleven months ago now has an automated agent above the human bar, a production deployment showing a 105× cost reduction, and a policy community publishing 23-item playbooks for how governments should respond. Automated AI R&D is no longer a thought experiment about 2028; it is a line item on Intology's blog with externally verified numbers attached, and the labs still holding the frontier models will have to decide whether to build their own harnesses or watch startups elicit more from their models than they do.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

Anthropic logo
Models

Anthropic ships Fable and Mythos 5.1 with cheaper tokens and looser guardrails

The twinned 5.1 release cuts token costs, reduces false-positive refusals, and finally brings Zero Data Retention to Fable.

Jaeden Schafer5 min read
Anthropic logo
Models

Anthropic merges Claude and Claude Cowork memory into one system

Claude will now carry context across chat and Cowork, and users can read, edit, or delete stored memories on any topic.

Jaeden Schafer4 min read
Anthropic logo
Models

Anthropic ships Claude Sonnet 5 at $2 per million input tokens

Sonnet 5 hits 63.2% on agentic coding, close to Opus 4.8's 69.2%, at a fraction of the price.

Jaeden Schafer5 min read