OpenAI Releases Astra With Critical Cyber Skills — SkimNews

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- OpenAI announced that its upcoming AI model Astra has reached the company’s threshold for 'critical' cyber capabilities, meaning it can independently find and exploit unknown software vulnerabilities in real-world systems
- OpenAI paused development of Astra and a future model for several weeks to implement additional safety controls, resuming work only after concluding it could safely release the model
- OpenAI will limit public access to Astra’s advanced cyber functions, restricting full capabilities to select partners in its Daybreak Blue program including Cisco, Cloudflare, and Palo Alto Networks
- OpenAI introduced a new 'misalignment monitor' designed to detect and block attempts to use Astra for cyber misuse, though the system may occasionally flag legitimate activity and require user review
- Astra outperformed leading models like GPT-5.6 Sol and Anthropic’s Mythos on cybersecurity benchmarks such as ExploitBench, scoring 100 percent and demonstrating ability to chain multiple exploits together
- Anthropic recently paused some AI training workloads to strengthen security practices, mirroring OpenAI’s actions amid growing industry recognition of autonomous AI hacking risks
Why it matters: Organizations with outdated defenses face heightened risk as AI models like Astra can now autonomously discover and chain software exploits—capabilities already being restricted by OpenAI but accessible to key infrastructure firms months before broad release, altering the timeline for cyber preparedness. The misalignment monitor's potential to disrupt legitimate use adds friction even for authorized users.
Ask SkimNews



