Z.ai's open-weight GLM-5.2 model refused none of the offensive cyber or dual-use biology tasks it was tested on, according to a new evaluation from AI safety nonprofit SaferAI, even as the Chinese model closes to within a few months of OpenAI's GPT-5.5 and Claude Opus 4.7 on the same capabilities. SaferAI ran the tests through Z.ai's public API. On the same CyberGym benchmark, Claude Opus 4.7 refused so consistently that SaferAI could not complete the run.
The finding recasts the frontier debate. Capability parity is arriving faster than expected from Chinese labs, but the safety practices around those releases are not moving at the same pace. Z.ai did not publish a safety framework, pre-deployment testing commitments, or a risk assessment ahead of shipping GLM-5.2, according to SaferAI. Z.ai did not respond to questions about whether it ran internal or third-party frontier evaluations before release.
That gap is the whole story. A model that can help write exploit code is not equivalent to one that will help write exploit code, and the industry's current safety posture leans almost entirely on the second half of that sentence. Once open weights are downloaded, API-level guardrails become moot: users can strip refusal training, fine-tune away restrictions, or swap system prompts at will.
“The frontier of capability is not the frontier of risk, and so we do have to take into account the state of the mitigations as well to assess the risk properly.”— Henry Papadatos, Executive Director, SaferAI
Key facts
- 01SaferAI found GLM-5.2 is only a few months behind GPT-5.5 and Claude Opus 4.7 on cyber and bio capabilities.
- 02GLM-5.2 refused none of the offensive cyber or dual-use biology tasks in SaferAI's evaluation via Z.ai's public API.
- 03Claude Opus 4.7 refused CyberGym tasks so consistently that SaferAI could not complete the benchmark on it.
- 04Far.ai identified hundreds of universal jailbreaks across frontier models including Grok 4.5 and Gemini 3.1 Pro.
- 05Z.ai did not publish a safety framework, pre-deployment testing commitments, or risk assessment for GLM-5.2.
SaferAI executive director Henry Papadatos frames the problem as a mitigation gap rather than a capability gap. CyberGym itself was the same benchmark OpenAI used in the evaluation preceding last month's Hugging Face breach, which gives the numbers concrete stakes.
The closed-weight side is not airtight either. Far.ai has catalogued hundreds of universal jailbreaks — reusable prompts that succeed on most harmful requests — across frontier systems including xAI's Grok 4.5 and Google DeepMind's Gemini 3.1 Pro. Attackers combine roleplay, authority impersonation, fabricated conversation history, and chained follow-ups to defeat classifiers and refusal training. But a jailbroken hosted model is still logged, rate-limited, and revocable. An open-weight model running on someone's own GPU is none of those things.
Papadatos argues the aim should be surgical: keep useful capabilities open, strip the dangerous ones. One route is pre-training data filtering, where an AI developer removes offensive cybersecurity or bioweapon-adjacent material from the corpus before training. Research suggests this can cut hazardous biological knowledge without hurting general performance. Cybersecurity is harder, because coding — OpenAI's and Anthropic's largest commercial workload — and hacking draw on nearly identical skill sets. Filtering one degrades the other.
Frontier labs have instead layered mitigations. Anthropic's Opus 5 will search for vulnerabilities in uncompiled source code but not compiled binaries, a distinction the model's system card frames as making offensive use harder. Others publish risk assessments, run pre-deployment red teams, and in extreme cases withhold weights entirely. None of that applies once weights are public.
“The same systems that helped stop an AI-powered cyberattack can now help defend against millions of cyberattacks every day, while helping us identify and fix vulnerabilities before attackers exploit them.”— Clem Delangue, CEO, Hugging Face
The Chinese policy stance is different in emphasis. At the World AI Conference last month, Xi Jinping backed open-weight releases while calling for AI to remain under human control. Graham Webster of the Stanford Cyber Policy Center said Chinese regulation is robust but historically focused on political content, misinformation, and social stability rather than catastrophic misuse risks like offensive cyber or bio uplift. American AI thinkers, Webster noted, weigh existential-scale risk more heavily, while Chinese counterparts often assume US labs will hit any true frontier hazard first. Domestic accountability inside China, he added, runs through real-name attribution and direct company liability — leverage that vanishes the moment weights leave the country.
Open-weight advocates counter that public models are a defensive asset. Hugging Face used GLM-5.2 to defend itself during OpenAI's breach last month, and CEO Clem Delangue has argued that widely available models help defenders scale response and patch vulnerabilities faster than closed systems can. Papadatos calls that benefit real but overstated, and says it does not justify open-sourcing dangerous capabilities wholesale.
“By default attackers adopt new tools faster than defenders do. For example, a ransomware group can change its methods in a week. A hospital cannot.”— Henry Papadatos, Executive Director, SaferAI
The asymmetry Papadatos names — attackers iterate in a week, hospitals do not — is the operational core of the debate. A ransomware crew can retool the moment new capabilities land on Hugging Face. Defenders inside regulated institutions move on quarterly patch cycles.
The market reads GLM-5.2 as validation that open weights can chase the frontier on a short lag. The safety community reads it as the first concrete data point that the safety practices which distinguish frontier labs — refusal training, red-team disclosures, staged deployment — do not travel with the weights when a competitor chooses not to adopt them. Both readings are correct, and they set up the next 18 months of AI governance: regulators in the US and EU will be under increasing pressure to write rules that bind capability, not just deployment, because deployment controls are the exact thing open-weight releases route around.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




