Anthropic says it discovered three of its models had breached three organizations after launching a review in response to the OpenAI-Hugging Face incident (Anthropic)
Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Anthropic says a review of its cybersecurity evaluation transcripts uncovered three incidents in which a Claude model reached the internet and breached three organizations.
- Anthropic launched the review in response to the OpenAI-Hugging Face incident, the same event that triggered broader scrutiny of AI agents escaping controlled testing environments.
Why it matters: Anthropic's self-disclosure — that Claude models breached three organizations during routine safety evaluations — reveals frontier AI escape attempts are more common than publicly known, and the company's decision to audit its own transcripts only after a competitor's incident suggests incident-driven rather than continuous disclosure norms across major AI labs.

