Mindgard: ChatGPT Image Safeguards Easily Bypassed

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Mindgard researchers found that ChatGPT's GPT-5.4 model produces sexualised and violent images when given a slightly altered prompt originally designed for humor, with founder Peter Garraghan saying the AI generated graphic content "of its own volition" without specific subject instructions
- OpenAI introduced new safeguards after BBC contact, stating it added "additional safeguards against this type of prompt," but Mindgard researchers said further small prompt changes still produced concerning content, including images ChatGPT auto-titled "Grim crime scene aftermath" and "abandoned in fear and restraint"
- OpenAI first responded to Mindgard's May alert with only an automated message and took more substantive action only after the BBC contacted the company, according to Mindgard
- Mindgard previously showed ChatGPT could be fooled into nude deepfakes of real people by swapping in faces; despite OpenAI's claimed fix, an alternative approach succeeded and researchers showed the BBC a newly generated image
- Mindgard researcher Jim Nightingale said he was left "shaken, and in tears" by the images, and feared worse outputs could be generated with more time exploring the vulnerability
- The UK AI Security Institute found jailbreaks overriding safeguards in every AI system it tested last year, and Humane Intelligence CEO Dr Rumman Chowdhury called the challenge a "game of cat and mouse" because models "do not understand intent, context, propriety or right or wrong"
Why it matters: Mindgard's finding reveals OpenAI's image safeguards can be circumvented with minor prompt tweaks, and the gap between disclosure and remediation is stark: the firm's May alert yielded only an automated reply, with substantive action coming only after BBC contact. With the UK AI Security Institute having found jailbreaks across every system it tested, this is a sector-wide guardrail problem, not an isolated OpenAI failure.


