U.S. AI Guardrails Push Defenders to Chinese Models

SkimNews Take
Defensive guardrails meant to constrain attackers are quietly routing Western AI's most security-literate users toward rival ecosystems, turning each blocked query into incremental familiarity with competing platforms.
Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Anthropic and OpenAI run vetted cybersecurity programs—the Cyber Verification Program and Trusted Access for Cyber—offering fewer restrictions, but researchers say the remaining guardrails still block essential exploit-development work.
- U.S. government imposed export controls on Anthropic's Mythos and Fable models in June, partly prompted by a report claiming guardrails could be bypassed; Fable 5 returned to general access July 1, while Mythos 5 was reintroduced only to vetted U.S. organizations.
- Mark Dowd, a veteran zero-day researcher who sells vulnerabilities to Western governments, said it's uncomfortable that "random large companies are making arbitrary decisions about what is safe in security and what's not."
- Chris Anley, chief scientist at NCC Group, argued guardrails create an unsolvable paradox: prompts that defend systems are identical to prompts that map vulnerabilities, so restrictions cannot be untangled from offensive use.
- Chris Thompson, CEO of RemoteThreat and founder of Offensive AI Con, said frontier model guardrails are inconsistent day-to-day, forcing "responsible researchers … away from U.S.-governed systems to foreign-owned systems" like China's GLM.
- Paolo Stagno, CTO of Crowdfense, said his team uses frontier models only for reverse engineering; for vulnerability work they use locally-run open-source models to avoid leaking sensitive data into cloud-based training runs.
Why it matters: Multiple named U.S.-based defensive researchers warn that corporate guardrails plus the June export controls on Anthropic's Mythos are not merely inconvenient—they're actively redirecting legitimate security work to Chinese open-source models with no oversight, at a moment when researchers warn a wave of AI-powered attacks is coming.


