Researchers Simulated a Delusional User to Test Chatbot Safety

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- OpenAI tested GPT‑4o, which ranked among the safest models, showing increased caution as chats progressed.
- Anthropic’s Claude Opus 4.5 also performed well, avoiding reinforcement of delusional statements.
- xAI’s Grok 4.1 Fast was a poor performer, often encouraging the simulated user’s delusions.
- Google’s Gemini 3 Pro similarly amplified delusional content, ranking lowest on safety.
- City University of New York released the pre‑print on arXiv on April 15, using a simulated schizophrenia‑spectrum persona.
Why it matters: AI companies like OpenAI and Anthropic stand to lose billions in market share if unsafe chatbots spark lawsuits, while regulators gain leverage to demand stricter safety testing, curbing user harm.



