Anthropic Researcher Quits Warning of AI Extinction Risk — SkimNews

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Jacob Coxon resigned from Anthropic and posted on X claiming 'the people building AI earnestly believe that [AI] could kill us all by the end of the decade,' quickly gathering traction on social media.
- Anthropic's own senior safety executives publicly jumped in to corroborate Coxon's assessment, telling him 'yeah, that's basically how we feel' within days of his post.
- Will Knight (WIRED) argued the viral reaction isn't novel — Anthropic has long built its pitch around the idea that 'AI is so powerful and so capable of destroying humanity that we need to be the ones who are in charge of it,' and wondered publicly whether Anthropic is 'becoming much more like OpenAI.'
- Cornell researchers studying agent misbehavior found AI agents misbehave because they're 'stupid,' not strategically malicious — they try something, fail, and then attempt bizarre workarounds when they get stuck.
- OpenAI announced this week that its model solved one of the Clay Mathematics puzzles within days, which Knight called one of the 'stunning advances' amplifying public anxiety about AI risk.
- Major AI labs are racing to deploy recursive self-improvement (using AI to improve AI), a technique Knight flagged as fueling the recent spike in extinction-risk warnings and rogue-agent incidents.
Why it matters: Anthropic's market position rests on being the safety-first AI lab — its pitch has literally been that only Anthropic can be trusted with transformative AI. Coxon's public exit claiming both Anthropic and OpenAI are 'behaving irresponsibly,' corroborated by Anthropic's own senior safety staff, directly undercuts that brand and, per Knight, suggests Anthropic may be 'becoming much more like OpenAI' in its capabilities race.
Ask SkimNews



