AI Guardrails Push U.S. Cyber Researchers to Chinese Models

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Anthropic's Mythos and Fable models faced U.S. export controls in June after a report claimed their guardrails could be bypassed for cyberattacks; Fable 5 returned to general access July 1, while Mythos 5 was reintroduced only to vetted U.S. organizations.
- OpenAI and Anthropic both run vetted researcher programs — Trusted Access for Cyber and the Cyber Verification Program — but researchers say the guardrails still remain overly strict even inside them.
- Mark Dowd, a veteran zero-day researcher, said it's "not really comfortable" that large AI companies make "arbitrary decisions about what is safe in security and what's not."
- Chris Anley, chief scientist at NCC Group, argued that prompts like "fix this code" are "irreducibly" both offensive and defensive tools, and guardrails can't cleanly separate the two.
- Paolo Stagno of CrowdFense said AI companies treat customers "like children who need babysitting" and only uses frontier models for reverse engineering, turning to local open-source models for sensitive vulnerability work due to data-leak concerns.
- Chris Thompson of RemoteThreat said guardrails are inconsistent even inside vetted programs, pushing responsible researchers toward Chinese open-source models like GLM that run locally with no restrictions.
Why it matters: Per the researchers cited, U.S. export controls on Anthropic's Mythos and Fable, combined with strict cyber guardrails at OpenAI and Anthropic, are pushing legitimate defenders toward unsupervised open-source alternatives — including Chinese models like GLM. Cybersecurity researchers are warning of an incoming wave of AI-powered attacks while the tools to prepare for them stay restricted.


