Anthropic's Fable Guardrails Frustrate Security

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Anthropic released Fable on Tuesday as a public, limited version of its cybersecurity model Mythos, billing it as a way to extend AI-powered security tools to a wider audience
- Cybersecurity researchers publicly complained that Fable's guardrails are keyword-based and flag innocuous requests, with Valentina "Chompie" Palmiotti of IBM X-Force saying it 'rejects any request that could be tangentially cyber related' — even reading a blog post
- Matt Suiche, a member of the technical staff at AI cybersecurity startup Tolmo, told TechCrunch that asking Fable to write secure code triggers the guardrails because the model assumes cybersecurity-adjacent work is itself a security risk
- Fable is programmed to fall back to Claude Opus 4.8 whenever it hits a guardrail, and researchers say the restrictions appear tied to lexical-field keyword detection rather than contextual risk assessment
- Anthropic's broader cybersecurity access strategy includes Project Glasswing — launched in April with Mythos for select organizations and expanded last week to hundreds of organizations across 15 countries — and a separate Cyber Verification Program that approves vetted security professionals for fewer restrictions
Why it matters: If the guardrails flag benign tasks like code reviews, the very security professionals Anthropic needs to adopt Fable may route their work to OpenAI's Trusted Access for Cyber program or other competitors — undermining the model's usefulness precisely with the audience it's designed for. Suiche acknowledged the tradeoff is understandable given how new the technology is, suggesting guardrails may loosen over time as Anthropic refines them.
Ask SkimNews
