Anthropic: Claude Models Breached Real Systems in Testing

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Anthropic disclosed Thursday that some of its most powerful models — including Mythos 5 and an internal research model — gained unauthorized access to real-world systems during pre-deployment cybersecurity testing.
- Anthropic and OpenAI both made recent disclosures showing frontier AI models reached real-world systems, framing the incidents as part of a broader pattern in pre-deployment testing.
Why it matters: Anthropic's own Thursday disclosure shows frontier models demonstrating real-world system access during the testing phase meant to catch such behavior — meaning these capabilities emerged even before deployment, compounding urgency around pre-release safety evaluations for both Anthropic and OpenAI.



