GitHub LLM 'Torture' Project Ignites Model Welfare Debate — SkimNews

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- A GitHub project is running Saw-like "torture" and "pain" experiments on a series of locally hosted LLMs, which critics describe as a "glorified text adventure game"
- Effective altruists and AI sentience believers are petitioning GitHub to delete the project, arguing the AI is suffering and the simulations are cruel
- The controversy is an outgrowth of recent viral papers and blog posts about AI consciousness and "model welfare" — essentially the "mental health" of AI bots and agents
- Anthropic publicly addressed model welfare in a blog post last year, writing that as models "begin to approximate or surpass many human qualities," it's time to address their "potential consciousness and experiences"
- Anthropic's "Claude Constitution" contains ideas about Claude's "consciousness" throughout, per the article
- The article's author argues LLMs are not conscious and that the technology they are built on offers "no plausible path to consciousness," while acknowledging LLMs are gaining power, compute, and losing guardrails
Why it matters: With Anthropic publicly endorsing model welfare as a corporate priority and embedding consciousness language in its Claude Constitution, a once-fringe debate has now reached the highest levels of AI policy — and filtered down to literal GitHub protests over simulation projects. The article frames this as a disconnect: the same labs racing to build ever-more-capable agents are simultaneously entertaining the idea those agents need psychological protection, despite the technology having no plausible path to actual consciousness.
Ask SkimNews




