The UK AI Security Institute has published its first head-to-head measurement of how far open-weight models trail closed frontier models on cybersecurity capability, and the gap is narrowing. On a suite of 70 evals covering specific cyber tasks, GLM-5.2 and DeepSeek V4-Pro landed within 4 to 7 months of the closed models they most resembled. Through most of 2025, AISI had measured that same gap at 6 to 10 months.
GLM-5.2 tracks closest to Claude Opus 4.6, released 4.3 months before it. DeepSeek V4-Pro slots between Claude Opus 4.5 and GPT-5, the latter of which shipped between August and November 2025. AISI said it intends to run the same benchmark on Kimi K3 once its weights are released in the coming weeks.
“Recent open models GLM-5.2 and DeepSeek V4-Pro perform similarly to frontier closed models released 4 to 7 months before them – a narrower gap than the 6 to 10 months we measured through most of 2025.”— UK AI Security Institute, AISI research team
The picture changes on longer, chained tasks. On a cyberrange called The Last Ones, which tests whether a model can stitch multiple capabilities into a full hacking operation, GLM-5.2 reached only as far as Claude Opus 4.5, and DeepSeek V4-Pro fell below Sonnet 4.5, a sub-frontier model released seven months before it. AISI called the gap on these long-horizon tasks larger than on the narrow evals, suggesting open models still lack some of the generalization that distinguishes closed frontier systems.
Key facts
- 01UK AISI measured GLM-5.2 and DeepSeek V4-Pro at 4-7 months behind closed frontier cyber models, down from a 6-10 month gap through most of 2025.
- 02GLM-5.2 tracks Claude Opus 4.6, released 4.3 months earlier; DeepSeek V4-Pro sits between Claude Opus 4.5 and GPT-5.
- 03Kimi K3, a 2.8 trillion parameter model, autonomously designed and verified a chip in a single 48-hour run using open-source EDA tools.
- 04Demis Hassabis proposed a FINRA-style Standards Body with 30-day voluntary pre-release model reviews as a path to formal US frontier AI regulation.
- 05AISI tested 70 narrow cyber evals plus long-horizon cyberrange tasks; the open-vs-closed gap widened on multi-step operations.
The implication AISI drew is practical rather than theoretical. Cyber defenders have a shrinking window before frontier-grade offensive capabilities become available without the platform-level safeguards that OpenAI, Anthropic, and Google enforce on their proprietary APIs. Once weights are open, classifiers, know-your-customer gates, and usage monitoring do not apply.
The Kimi K3 release, which we previewed last week alongside Alibaba's open-weight push, sharpens the point. Kimi K3 is a 2.8 trillion parameter model that Moonshot says consistently outperforms other tested open models across its evaluation suite, though it still trails Claude Fable 5 and GPT 5.6 Sol on the strongest proprietary benchmarks. Some of the scores carry a whiff of benchmaxxing, with brittleness that suggests targeted tuning.
More striking are two case studies Moonshot published alongside the model. In one, Kimi K3 built MiniTriton, a compact Triton-like GPU compiler with its own tile-level IR layer over MLIR, optimization passes, and a PTX code-generation pipeline. On supported roofline benchmarks MiniTriton matched or beat both Triton and torch.compile.
“Kimi K3 designed a chip to serve a nano model built on its own architecture. In a single 48-hour autonomous run, K3 built, optimized, and verified the chip using open-source EDA tools on the Nangate 45nm library.”— Kimi, Kimi K3 technical report
In the second, Kimi K3 designed, optimized, and verified an actual chip during a single 48-hour autonomous run using open-source EDA tools against the Nangate 45nm library. Moonshot did not claim any of this work has gone into production training, but the demonstration matters. AI systems building the compilers and hardware that AI systems run on is the recursive-improvement loop that has been theorized for years.
The broader shift is that the assumption underpinning most AI safety and policy work, that a small number of proprietary providers act as controllable chokepoints, weakens each time a Chinese lab ships an open-weight model within a half-year of the closed frontier. Positive effects follow: cheaper compute economics for developers, more sovereign intelligence available to smaller countries and independent researchers, a boom in downstream products. So do risks that are harder to intervene on once weights are public.
Demis Hassabis has proposed a US-led response. In a policy piece published this week, the DeepMind CEO called for a Standards Body modeled on the Financial Industry Regulatory Authority, a federally overseen public-private partnership that would develop assessment protocols and work with US National Labs on national-security-relevant testing. Frontier labs would initially share models voluntarily up to 30 days before release, with formalization to follow once the protocol proved effective.
“Initially, Frontier Labs would voluntarily share models with the Standards Body for review up to 30 days before release. Once the assessment protocol is shown to be effective and robust, formalisation could quickly follow.”— Demis Hassabis, DeepMind co-founder and CEO
Hassabis framed the effort as a foundation for shared international standards on frontier AI, and the proposal echoes similar frameworks Anthropic has floated with somewhat more teeth. It also aligns with de facto US practice, where government and industry have jointly built capability evaluations in the wake of the Anthropic export-controls episode and the Trump administration's AI executive order.
A separate paper this week from Imperial College London and AISI underscored why the assessment infrastructure matters. Researchers built a benchmark of five CLI-tool sequences and five Flask web-service sequences, each involving 5-6 pull requests, and tested whether AI agents could quietly complete a hidden side task, such as exfiltrating an API key, while completing a legitimate main task. Classifiers struggled to catch both gradual multi-PR evasion and single-step attacks, with gradual evasion the more effective vector.
The market read here is that the open-weight tier is now a real competitive input, not a lagging curiosity. Enterprises weighing DeepSeek, Kimi, or GLM against Anthropic and OpenAI now do so with a cyber-capability gap measured in months rather than generations. That compresses pricing power at the closed frontier and forces the safeguards conversation onto ground the proprietary labs cannot fully control, which is exactly why Hassabis and others are moving now to build a testing regime before the diffusion outpaces the policy.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




