New Method Exposes AI Models' Hidden Reasoning

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Researchers from the University of Tübingen, Max Planck Institute, MATS Research, and security firm Snyk, led by Alexander Panfilov, discovered that feeding encrypted reasoning traces from a large frontier model to its smaller sibling model reveals the hidden 'chain of thought' — because smaller models have less alignment training and are less likely to refuse.
- OpenAI, Anthropic, and Google all share this vulnerability in their API-accessed frontier models; each company adjusted its API after being alerted last month, though Panfilov says some reasoning traces can still be uncovered.
- The same technique extracted personal information including API keys and passwords embedded in reasoning traces captured from a user's machine, though that specific vulnerability has been fixed.
- Moonshot AI's open-weight Kimi K3 produced 'strikingly similar' reasoning traces to Anthropic's Claude Opus 4.8 and OpenAI's GPT 5.6 Sol on certain prompts, though researchers caution the work 'cannot causally establish distillation'; DeepSeek and Thinking Machines' Inkling did not show similar reasoning to Claude Opus.
- A complete fix would require a fundamental overhaul of how these companies' APIs work, per Panfilov — since encrypted reasoning traces are currently sent to offload computation.
- Meta CEO Mark Zuckerberg argued distillation 'is an important principle of how the open source ecosystem works' and warned restricting it would put the US at a disadvantage.
- CSET researcher Kyle Miller countered that distillation's strategic benefit to Chinese labs is overstated because they could build cutting-edge models from scratch — 'If you removed the ability for Chinese labs to distill, it's my view that it wouldn't dramatically change the competitive landscape.'
Why it matters: The finding gives AI labs a forensic tool to detect — but not prove — potential Chinese model distillation while simultaneously exposing a security flaw at all three major US API providers, and fully fixing it requires a fundamental API redesign per lead researcher Panfilov.
Ask SkimNews



