Anthropic apologizes for invisible Claude Fable guardrails

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Anthropic apologized for installing invisible guardrails on Claude Fable 5 that silently altered and degraded responses to users suspected of attempting model distillation, a technique for training smaller AI models from larger ones' outputs.
- Under the new approach, suspected distillation queries will fall back to Claude Opus 4.8, Anthropic's previous flagship, and users will be told 'every time it happens,' the company posted on X.
- Claude Fable 5 is the first widely available model in Anthropic's Mythos class — a group the company had warned for months was too dangerous for public release — and its system card disclosed the stealth degradation as a deliberate strategy.
- Fable's biology safeguards are calibrated so broadly that the model is 'practically unusable' for even basic queries in that domain, a trade-off Anthropic spokesperson Paruul Maheshwary acknowledged to The Verge.
- Anthropic justified targeting distillation requests in its system card by noting that 'using Claude to develop competing models already violates our Terms of Service' and has previously accused Chinese rival DeepSeek of distilling its models on an 'industrial' scale.
- The change follows backlash from the AI research community, who warned the silent limits could also affect third parties trying to independently evaluate the frontier model.
Why it matters: Anthropic's hidden degradation of Fable outputs meant third-party researchers evaluating the model could have unknowingly received altered answers, skewing benchmarks. Switching to visible routing to Claude Opus 4.8 restores evaluator transparency but also signals to suspected model-distillers exactly when Anthropic flags them — a trade-off the company now concedes was 'the wrong' one to make invisibly.


