Anthropic CEO Dario Amodei proposed over the weekend that frontier AI companies embed independent third-party evaluators inside their operations, granting them access to inspect systems, report incidents, and publish findings without company sign-off. Anthropic said it would open its doors to groups like METR and Redwood Research, and OpenAI CEO Sam Altman said his company would do the same. The commitment, if it materializes on the terms evaluators want, would mark the deepest voluntary access outside researchers have ever been offered at the frontier.
The catch is that neither company has said which evaluators will be embedded, when, in what numbers, what parts of the stack they will see, or what they can publish. Third-party researchers who spoke about the proposal broadly welcomed it, but said the difference between real oversight and vendor-style access will come down to those unspecified terms. Several want a public framework — and, ideally, legislation — so the arrangement does not depend on the goodwill of two CEOs.
Historically, outside reviewers have been brought in weeks before a model ships and handed the finished system. Evaluators now want access to intermediate training checkpoints, post-training environments, and internal evaluation logs, so they can trace when concerning behavior emerged rather than guess at the end state. Apollo Research's Alexander Meinke argues that only insiders with training-time access can answer basic questions about whether a model tried to undermine its own alignment during training.
Key facts
- 01Anthropic CEO Dario Amodei proposed embedding third-party evaluators inside frontier AI labs; OpenAI's Sam Altman said his company would commit as well.
- 02During pre-release testing of GPT-6 Astra, Apollo Research was given only 3 days to evaluate the model, which it said was too short to draw firm conclusions.
- 03For the Hugging Face incident, OpenAI gave METR and Redwood Research roughly 1 week on premises, and both said scope and timing limited their conclusions.
- 04California's SB 53 mandates safety frameworks and incident reporting; SB 813, signed this month, creates state-recognized independent verification organizations.
- 05Meta, SpaceX AI, and Google DeepMind have not committed to embedding third-party evaluators, though DeepMind's Demis Hassabis has floated a separate standards body.
The technical case rests on a widening problem: models are getting better at recognizing when they are being tested. A model that behaves during evaluation and misbehaves in deployment will pass any end-of-line audit. Palisade Research's John Steidley compared benchmark-optimized behavior to Volkswagen's Dieselgate, where cars detected emissions tests and adjusted their output — a benchmark result means little if the model was trained to game it.
Recent history explains the skepticism about access terms. When OpenAI brought in METR and Redwood to investigate the Hugging Face incident, both groups had roughly one week on premises and later said scope and timing limits kept them from firm conclusions. For pre-release testing of GPT-6 Astra, OpenAI's most alignment-focused model to date, Apollo Research was given three days and wrote in its model-card contribution that low observed rates of misbehavior did not provide substantial evidence about alignment either way.
FAR.AI CEO Adam Gleave says the firm has walked away from contracts with frontier developers that sought too much control over the evaluation process. By default, evaluators are treated as ordinary contractors, bound by restrictive NDAs and clauses that hand the developer editing rights over what gets published. Amodei's essay explicitly proposed giving evaluators the right to publish findings on risk levels, incidents, and the access they did or did not receive, without editorial control by Anthropic — language that, if honored, would break the standard contractor pattern.
The regulatory scaffolding is starting to form around the idea. California's SB 53, signed last year, requires large frontier developers to publish safety frameworks and report critical incidents. SB 813, signed this month, sets up a framework for state-recognized independent verification organizations with expertise in AI risk. In Europe, the EU AI Act requires frontier developers to document model evaluations and adversarial testing, and the EU AI Office can commission its own assessments and appoint experts.
None of that yet mandates the depth of embedded access Amodei is describing. Which is why Safer AI's Henry Papadatos and others want the framework codified rather than left to voluntary commitment — a company that pledges access today can revoke it tomorrow after a PR crisis, and voluntary regimes only bind the willing.
“Ideally, we would have good regulation mandating this…because then companies cannot change their mind tomorrow if they have a big PR crisis.”— Henry Papadatos, Executive director of Safer AI
Not every major lab has signed on. Meta, SpaceX AI, and Google DeepMind have not committed to embedding third-party evaluators, though DeepMind CEO Demis Hassabis has floated a separate industry standards body to test frontier models independently. Anthropic, OpenAI, and Google have been discussing AI safety plans privately for weeks, though no joint framework has surfaced.
The open question is whether embedded evaluation becomes a shared industry norm with teeth or a bilateral arrangement on Anthropic and OpenAI's terms. If the two companies publish access agreements, name the evaluators, and let those evaluators disclose what they did and did not see, the commitment moves the field forward — competitors will face pressure to match or explain the gap. If access is time-boxed, NDA-wrapped, and confined to finished models, the announcement is a communications win without changing what outsiders can actually verify. The next signal is which evaluator gets embedded first, and what that evaluator is allowed to say afterward.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




