OpenAI Built GPT-Red to Hack Its Own AI

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- OpenAI built GPT-Red, an LLM super-hacker designed to red-team its own models by finding and exploiting vulnerabilities before release.
- GPT-Red discovered a new type of prompt injection attack called 'fake chain of thought,' where it tricks an LLM into acting on spoofed internal reasoning notes.
- GPT-Red outperformed human red-teamers in identifying effective attacks on earlier versions of GPT-5, demonstrating superior persistence and efficiency.
- OpenAI used a self-play training loop in a simulated 'dojo' environment where GPT-Red attacked other models, improving both offensive and defensive capabilities.
- GPT-5.6 was hardened using GPT-Red, with fewer than 23% of GPT-Red’s strongest attacks succeeding, compared to over 90% success against the older GPT-5.
- Chris Choquette-Choo, a research scientist at OpenAI, emphasized that replicating GPT-Red is not trivial due to the company’s extensive compute resources and over a year of development effort.
- Jessica Ji of Georgetown’s CSET called the self-play approach promising, noting that human expertise will still be essential to guide testing where models fall short.
Why it matters: By automating red-teaming with GPT-Red, OpenAI significantly raises the bar for AI security testing, reducing vulnerability to novel attacks like fake chain of thought. This shifts the cost and speed of safety validation, making it harder for rivals or attackers to replicate defenses—especially given the resource-intensive development behind GPT-Red.




