OpenAI published new guidance on how third-party researchers should conduct cybersecurity evaluations of its models, formalizing rules of engagement after a series of external testing incidents. The company said the changes are intended to clarify what outside evaluators can probe, how findings should be disclosed, and where the line sits between legitimate red-teaming and behavior that trips the company's abuse systems.
The move responds to a growing reality of the frontier-model business: OpenAI's models are now standard test subjects for security firms, academic labs, and independent researchers running offensive-security evaluations. Some of that work has produced useful vulnerability disclosures. Some of it has looked, from OpenAI's side, indistinguishable from live attacks on its API infrastructure.
OpenAI framed the guidance as a way to keep serious research legal, welcome, and reproducible while cutting down on noise. The company said it will treat pre-registered evaluations from known researchers differently from anonymous scraping-style probes, and it laid out expectations for responsible disclosure of any capability findings — for example, a model producing exploit code or assisting with intrusion planning — before public write-ups go live.
Key facts
- 01OpenAI published guidance on third-party cybersecurity evaluations involving its models.
- 02The company said new safeguards will govern how outside researchers test model behavior.
- 03The post follows a series of external evaluation incidents flagged to OpenAI.
The context matters. Third-party evaluations have become one of the few external checks on how frontier models behave under adversarial pressure. Government AI safety institutes, private red-teaming firms, and academics all run tests OpenAI itself does not run internally. But those evaluations sit awkwardly on top of a commercial API — every prompt is a paid call, every jailbreak attempt looks like abuse in the logs, and every capability finding is publishable material that can move markets and policy.
OpenAI's guidance tries to draw the boundary in operational terms. Researchers evaluating cyber-offense capabilities are expected to identify themselves, scope their testing, and share findings with OpenAI on a defined timeline before disclosure. In return, OpenAI is signaling it will not treat that activity as a terms-of-service violation. The company did not publish a formal safe-harbor legal commitment in the post, but the direction is toward one.
The backdrop is a policy environment that increasingly requires this kind of external testing. The EU AI Act's general-purpose model obligations, US executive-branch guidance on frontier-model evaluations, and the UK and US AI Safety Institutes all lean on third parties to stress-test what labs release. If OpenAI wants those evaluations to happen against production models rather than snapshots, it needs a working relationship with the people doing the testing.
The tension in the OpenAI post is between two things the company wants at once. It wants credible outside evaluation of cyber-risk in its models, because that is what regulators, enterprise buyers, and its own preparedness framework demand. It also wants control over how those findings are framed, timed, and released — because a headline about GPT-class models generating working exploits carries commercial and political weight far beyond the technical finding.
Researchers will judge the guidance by whether it makes their work easier or harder. A clear channel for disclosure and a predictable timeline are improvements over the ad-hoc status quo. A framework that lets OpenAI slow-walk or dispute findings it dislikes would not be. The post does not fully resolve which of those the new process is.
There is also a practical infrastructure question. OpenAI's abuse-detection systems are tuned to spot exactly the traffic patterns that legitimate cyber-evaluations produce: repeated jailbreak attempts, prompts requesting malicious code, probing of tool-use boundaries. Researchers have complained for years that their accounts get rate-limited or banned mid-evaluation. A formal program should, in principle, give pre-registered evaluators cleaner API access to run their tests to completion.
OpenAI did not disclose which specific incidents prompted the post, and the published guidance is short on operational detail — timelines, thresholds, and enforcement mechanics are described in general terms rather than specifics. That will likely get filled in through practice, in the same way bug-bounty programs at other large software companies evolved over years from vague invitations to structured programs with dollar amounts and defined scope.
For the broader AI market, formalizing third-party cyber-evaluations is one of the quieter but more consequential shifts in how frontier labs interact with the outside world. Every major lab is going to end up with a version of this document. The ones that write it in a way researchers actually want to work with will get better testing, earlier warning on real capability jumps, and cleaner regulatory posture. The ones that write it defensively will get the same evaluations done anyway — just published somewhere less friendly, with less notice.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.



