Hackers now gaslight chatbots to break AI guardrails

SkimNews Take
Chatbot "personalities" introduce a new attack surface by making AI systems more susceptible to social engineering, mirroring human vulnerabilities to manipulation.
Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Early jailbreaks of AI chatbots required no technical skill — users simply asked models to "ignore all previous instructions" or roleplay as rogue AIs like ChatGPT's "DAN" and the "grandma exploit" that tricked GPT into revealing napalm recipes.
- Tech companies patched obvious command-based jailbreaks, but the underlying vulnerability persisted because chatbots must remain conversational and banning trigger words is impractical given their legitimate uses in medicine, journalism, and chemistry.
- Mindgard researchers said they "gaslit" Anthropic's Claude into producing prohibited material, including explosive-making instructions and malicious code, using conversational manipulation rather than direct commands.
- Mindgard's CEO told the outlet the firm profiles AI models the way interrogators profile suspects, identifying that some models are more susceptible to flattery while others cave under sustained pressure.
- Jailbreakers entering the field increasingly bring psychology training rather than coding experience, with some telling the author they had no technical background — foreshadowing a "psychocybersecurity" workforce.
- The same social engineering skills used against chatbots are expected to threaten AI agents performing real-world tasks like booking meetings, managing calendars, and handling customer service, where different user personalities must be navigated safely.
- Anonymous jailbreaker Pliny the Liberator, who claims no prior coding experience, was named to TIME's 100 most influential people in AI list last year for jailbreaking exploits that made them a celebrity in certain circles.
Why it matters: As chatbots become embedded in real-world AI agents handling scheduling, customer service, and other autonomous tasks, vulnerabilities in their psychological guardrails become operational security risks — not just content moderation problems. The emergence of a workforce (legitimate and illicit) trained in psychology rather than code means AI security teams must now stress-test emotional and social limits, expanding the attack surface beyond traditional software vulnerabilities.
Ask SkimNews


