Computer scientists have found a way to pull the hidden reasoning out of frontier AI models from OpenAI, Anthropic, and Google, and the same method recovered API keys and passwords embedded in those reasoning traces before the three companies patched it last month. The researchers tested the technique across 90 prompts and say every major frontier provider they examined was vulnerable. They also found that Moonshot AI's open-weight Kimi K3 produced outputs strikingly similar to the concealed reasoning traces of Claude Opus 4.8 and GPT 5.6 Sol, though the paper stops short of calling that proof of distillation.
The team came from the University of Tübingen, the Max Planck Institute, MATS Research, and the security firm Snyk. Their attack exploits the way frontier providers ship products in tiers. Larger models are more capable and more expensive to run; smaller siblings share infrastructure and, crucially, the same decryption keys used when encrypted reasoning is offloaded to a user's machine. Alignment training on the smaller variants is lighter, so they are more willing to expose the chain of thought that the flagship model is instructed to hide.
Feed an encrypted reasoning trace from a big model into its smaller sibling, and the smaller one obligingly decrypts and continues it in the clear. That is the entire trick. Anthropic, OpenAI, and Google were notified last month and each has adjusted its API to block the personal-information leak. Alexander Panfilov, the University of Tübingen researcher who led the work, says some reasoning content can still be reconstructed with the same method, and closing the hole entirely would require rebuilding how the APIs handle offloaded computation.
Key facts
- 01Researchers extracted hidden reasoning traces from frontier models at OpenAI, Anthropic, and Google by routing encrypted traces through smaller, less-aligned sibling models.
- 02The same technique recovered API keys and passwords embedded in reasoning traces from user machines before the three companies patched it last month.
- 03Across 90 test questions, Moonshot AI's open-weight Kimi K3 produced outputs closely matching hidden traces from Claude Opus 4.8 and GPT 5.6 Sol.
- 04OpenAI told US lawmakers in February that DeepSeek appeared to have copied one of its models; Anthropic told lawmakers in June that Alibaba systematically distilled its models to build Qwen.
- 05DeepSeek and Thinking Machines' Inkling did not exhibit the same reasoning similarity with Claude Opus in the researchers' tests.
Florian Tramer, a computer security researcher at ETH Zürich, called the swap between a large model and a weaker-aligned sibling with a shared decryption key a real and growing problem. Michael Aciman, an Anthropic spokesperson, said the study did not involve recovering encryption keys, accessing Anthropic infrastructure, or extracting personal data from its systems. Google and OpenAI declined to comment.
The distillation finding is the part that will move markets and policy desks. Distillation is a standard technique for compressing the behavior of a large model into a smaller one, and it underpins much of the open-weight ecosystem. It has also become a diplomatic flashpoint. OpenAI told US lawmakers in February that DeepSeek appeared to have copied one of its models to build the R1 reasoning system. Anthropic told lawmakers in June that Alibaba had systematically distilled its models to build Qwen.
When Panfilov's team fed the first few words of proprietary reasoning traces to open-weight models and let them continue, Kimi K3 from Moonshot AI generated remarkably similar completions to Claude Opus 4.8 and GPT 5.6 Sol. Two other open-weight models tested the same way — DeepSeek and Thinking Machines' Inkling — did not show that pattern with Claude Opus. Moonshot AI and Z.ai did not respond to requests for comment. The researchers are careful: the paper says the results cannot causally establish distillation, only that the reasoning similarity is unusually high.
“We value independent research on our models and have begun building short-term mitigations for the replay behaviors described in the report”— Michael Aciman, Anthropic spokesperson
The technique matters beyond the Kimi K3 question. It shows that closed models leak more of their reasoning through commercial APIs than providers assumed, which changes the threat model for anyone building on those APIs. If a downstream service passes sensitive strings through a reasoning model, those strings can end up in a trace that a motivated attacker can reconstruct via the smaller sibling. That is now closed for the identified path, but the underlying architecture — big model plus cheaper siblings sharing keys — is industry standard.
The geopolitical framing is louder than the technical picture warrants. Mark Zuckerberg wrote this week that distillation is an important principle of how the open source ecosystem works and warned that restricting it would disadvantage US developers. Kyle Miller of the Center for Security and Emerging Technologies argues the strategic significance of Chinese distillation is overstated, because it lifts existing models only so far and Chinese labs have shown they can build frontier systems from scratch.
Yarin Gal of Oxford makes the opposing point on progress rather than security: if every lab blocks distillation from every other, the overall pace of AI development slows, because distillation is one of the main ways capability spreads. That tension — between protecting frontier IP and letting the field compound on itself — is the real policy question, not the specific question of whether Kimi K3 studied at Claude's feet.
For the AI market, the immediate consequence is a quiet infrastructure change at three of the largest inference providers. The bigger consequence is that reasoning traces, which labs have been treating as a proprietary moat since OpenAI's o1, are now demonstrably porous through mundane API design choices. Expect more scrutiny of what tiered model families share behind the scenes, expect enterprise buyers to ask harder questions about what data leaves their reasoning calls, and expect the distillation debate to keep escalating regardless of whether this particular vector is patched.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




