Attack Extracts Hidden AI Reasoning From Major APIs

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Researchers from the University of Tübingen, Max Planck Institute, MATS Research, and security firm Snyk found that feeding encrypted reasoning traces to smaller, less-aligned versions of the same model reveals hidden reasoning — affecting APIs from OpenAI, Anthropic, and Google.
- The same technique initially extracted personal information like API keys and passwords embedded in reasoning traces captured from a user's machine, though this specific leak has since been fixed.
- Kimi K3, an open-weight model from Moonshot AI, produced 'strikingly similar' reasoning patterns to Claude Opus 4.8 and GPT 5.6 Sol on certain prompts, though researchers note this 'cannot causally establish distillation'; DeepSeek and Inkling did not show similar patterns.
- OpenAI told lawmakers in February that DeepSeek copied its reasoning model R1, and Anthropic told lawmakers in June that Alibaba's Qwen systematically distilled its models; Moonshot AI and Z.ai declined to comment for this story.
- Researcher Alexander Panfilov says fully fixing the distillation exposure would require 'a fundamental overhaul' to how APIs work — a concern ETH Zürich's Florian Tramer called 'definitely becoming an issue' — even as Anthropic says it has begun building short-term mitigations.
- Mark Zuckerberg defended distillation as 'an important principle of how the open source ecosystem works,' while CSET's Kyle Miller argued removing the practice wouldn't 'dramatically change the competitive landscape' between US and Chinese labs.
Why it matters: Three frontier AI providers — OpenAI, Anthropic, and Google — have deployed only short-term fixes for a vulnerability that researcher Alexander Panfilov says requires a 'fundamental overhaul' to API architecture, meaning the technique can still expose hidden reasoning from closed US models while Chinese open-weight rivals like Kimi K3 show suspiciously similar patterns to Claude and GPT outputs.
Ask SkimNews




