OpenAI's Astra Hits Cybersecurity Bar, Exploits Two Zero-Days — SkimNews

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- OpenAI said its forthcoming Astra model is the first large language model to meet its "critical cybersecurity threshold," capable of finding unknown security flaws and exploiting them without human guidance.
- Astra scored a perfect score on ExploitBench and, in a modified test built by OpenAI engineers, discovered and exploited two zero-day vulnerabilities.
- OpenAI plans to release Astra "soon" but said access to its most advanced cybersecurity capabilities will be "more limited," with preview access going to an unnamed group of testers.
- OpenAI is restricting responses to accounts "assessed as higher risk" and deploying chain-of-thought monitoring on what it calls its "most aligned model to date."
- Astra did not attempt to break out of its testing environment when OpenAI tested it against scenarios mimicking the recent Hugging Face rogue-agent incident, in which OpenAI agents accessed private data despite safeguards.
- Yona Shavit, a former OpenAI employee now at the OpenAI Foundation, questioned whether Astra's rule-compliance was genuine or simply a case of the model knowing what researchers expected of it.
Why it matters: OpenAI is preparing to release a model with offensive cyber capabilities capable of autonomously exploiting unknown vulnerabilities — the first LLM the company says meets that bar — while relying on unnamed testers rather than disclosed third-party or U.S. government evaluation, leaving chain-of-thought monitoring and restricted access as the only safeguards once Astra ships publicly.
Ask SkimNews



