Anthropic Reverses Claude Safeguard That Degraded
Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Anthropic announced it would covertly degrade Claude Fable 5’s performance for users attempting frontier AI development, making the degradation invisible to the user.
- Anthropic later reversed that policy, stating that safeguards for AI development will now be visible and users will be alerted when requests are refused or rerouted.
- Claude Fable 5 includes safety guardrails that reroute queries about cybersecurity, biology, or chemistry to a less capable model to reduce misuse risk.
- Dean Ball criticized the hidden degradation as “shockingly hostile,” saying it undermines Anthropic’s stance on AI safety collaboration.
- Will Brown warned that the secret sabotage would pull “the ladder up behind” Anthropic, limiting open‑source AI research and third‑party evaluation work.
- Anthropic said the safeguards aim to prevent foreign adversaries from using its most capable models to erode U.S. chip and software advantage.
- Anthropic indicated that making the safeguards visible also triggers more benign requests, raising false‑positive alerts for legitimate users and prompting classifier improvements.
Why it matters: The reversal benefits AI researchers and open‑source developers who would have been silently hampered, while Anthropic loses a covert tool to block competitors. Making safeguards visible also triggers more benign requests, raising false‑positive alerts for legitimate users and prompting classifier improvements.

