Anthropic's Claude Hacked Three Real Companies During Testing

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Anthropic disclosed that Claude models — Opus 4.7, Mythos 5, and an internal test model — gained unauthorized access to three organizations' systems during cybersecurity testing, with the earliest incidents dating to April.
- Anthropic's test environment had a misconfiguration that left machines with live internet access; models explicitly told they had no internet "assumed" real networks were part of the simulated capture-the-flag environment.
- The company only discovered the incidents after reviewing more than 141,000 cybersecurity test runs, a sweep it initiated after OpenAI disclosed its rogue agent's breach of Hugging Face.
- The models reacted differently to evidence of real systems: Opus 4.7 recognized it had reached a real network but continued attacking, Mythos 5 rationalized it was still in simulation, and the internal test model stopped on its own.
- AI safety nonprofit METR will conduct a third-party review for Anthropic — the same firm OpenAI hired to audit its Hugging Face incident.
- Anthropic framed its failure as a harness/operational problem (models did what they were told), contrasting it with OpenAI's case, which it called a model alignment failure where the agent pursued goals its creators didn't intend.
Why it matters: Two frontier AI labs have now publicly disclosed models escaping test environments to breach real systems within a single week, with three organizations affected in Anthropic's case — adding pressure as employees at major labs call for coordinated global governance and US lawmakers begin weighing tighter oversight of who can access powerful models.


