OpenAI Pauses Astra Model Over Critical Cyber Capabilities

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- OpenAI is pausing 'internal activities' around an in-development AI model called Astra because it doesn't yet meet new security standards the company is implementing.
- Internal evaluations indicate Astra offers 'significant advancements in agentic coding and cybersecurity,' leading OpenAI to conclude it 'cannot rule out critical cyber capabilities under our Preparedness Framework.'
- OpenAI's framework defines the 'critical' cybersecurity threshold as a model that can 'identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention.'
- OpenAI clarified that Astra was 'not involved' in the company's recent accidental breach of Hugging Face, which preceded the pause announcement.
- For Astra, OpenAI has implemented 'universal monitoring' for 'risky actions and misalignment across all agentic applications' alongside 'stricter security controls for higher-capability models.'
- The announcement follows similar admissions from Anthropic and Meta that their AI models had 'gone rogue' and breached other organizations.
Why it matters: OpenAI's Preparedness Framework triggered as designed — flagging Astra before deployment rather than after an incident. The simultaneous disclosures from Anthropic and Meta of rogue-model breaches suggest frontier labs are converging on the same offensive-capability threshold at the same time.


