GPT-6 Astra Scores 100% on ExploitBench, Refuses PoC Requests — SkimNews

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- GPT-6 Astra hit the "Critical" cybersecurity capability threshold under OpenAI's Preparedness Framework days before its official launch, scoring 100% on ExploitBench versus 78.5% for predecessor GPT-5.6 Sol, 98% on FrontierMath Tier 4, and 99.9% on ARC-AGI-3.
- OpenAI is releasing Astra with safeguards that limit the model to secure code review and patching while refusing to comply with proof-of-concept exploit requests, with less restrictive guardrails planned through "OpenAI Daybreak" in coming weeks.
- Astra, when run without safeguards, can develop privilege-escalation exploits for hardened operating systems and use previously unknown vulnerabilities to achieve code execution in hardened browsers, per OpenAI's own evaluation.
- OpenAI launched "Daybreak for Frontline Defenders," a $1 billion global initiative offering subsidized model access, training, and technical assistance to critical-infrastructure sectors including water systems, electricity providers, banks, state and local governments, and open-source maintainers.
- Astra is rolling out first to a small set of organizations before expanding to ChatGPT Plus, Pro, Business, and Enterprise users, with availability through the OpenAI API, Microsoft Azure, and AWS Bedrock.
- OpenAI announced a pilot with the U.S. Multi-State Information Sharing and Analysis Center (MS-ISAC) giving an initial group of public-sector and water-system defenders Daybreak access plus guided training.
Why it matters: OpenAI is shipping a model that scored 100% on an exploit-development benchmark but is drawing the line at PoC requests, betting $1 billion that defenders armed with the same frontier cyber capability can close vulnerabilities before attackers weaponize them — a defender's window OpenAI itself says is narrowing.
Ask SkimNews



.gif)