OpenAI Launches GPT-5.6-Cyber with Reduced Safeguards for Exploit Development

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- OpenAI launched GPT-5.6-Cyber, built on GPT-5.6 Sol and trained to find zero-day vulnerabilities and develop exploit chains with reduced refusals for dual-use cyber tasks, building on its GPT-5.5-Cyber model from June 2026.
- GPT-5.6-Cyber completes 95.0% of advanced cyber requests on OpenAI's internal "Advanced Cybersecurity Completion Rate" evaluation, versus 1.5% for GPT-5.6 Sol, 2.0% with Daybreak Blue access, and 57.3% for GPT-5.5-Cyber.
- The model discovered CVE-2026-15903 (CVSS 8.8), an out-of-bounds read/write in Google's V8 JavaScript engine patched in mid-July 2026, and also flagged 400+ privilege-escalation flaws in a popular OS kernel plus five mobile OS and three critical database vulnerabilities.
- Daybreak Red, part of OpenAI's Daybreak initiative introduced in May 2026, is the new tier distributing the model to authorized firms; launch partners include Accenture, Akamai, Cisco, Cloudflare, CrowdStrike, Fortinet, IBM, Palo Alto Networks, PwC, and Sophos.
- GPT-5.6-Cyber underperforms GPT-5.6 Sol on open-ended tasks like developing working proofs-of-concept and writing vulnerability reports, with OpenAI noting "the model sometimes producing shorter, less detailed vulnerability reports."
- OpenAI acknowledged the trade-off directly: "Models running with reduced safeguards carry risks beyond standard model usage, whether from misuse or misalignment."
- A 1Password study cited in the article found LLM-generated patches fully resolve vulnerabilities only 26.0% of the time without altering application behavior, with 53.9% failing to resolve or introducing new flaws.
Why it matters: For partners Accenture, Cisco, CrowdStrike, IBM, Palo Alto Networks, and six others, GPT-5.6-Cyber's 95% exploit-request completion rate is a material capability gain over GPT-5.6 Sol's 1.5%. OpenAI itself flags misuse risk; the cited 1Password study shows LLM patches fully resolve vulnerabilities only 26.0% of the time, with 53.9% failing or introducing new flaws — meaning the same model that finds bugs at scale also fails to patch them reliably.
Ask SkimNews




