Princeton Study: AI Agents Lack Creativity for Self-Improvement — SkimNews

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Princeton researchers led by Sayash Kapoor and Peter Kirgis tested AI agents on two unpublished NeurIPS 2026 questions using a new "shadow evaluation" method, and the original authors rejected both agent-produced papers as nowhere close to conference quality.
- Anthropic's Claude Opus 4.8, running on open-source software OpenClaw, was given six days, $3,000 in API credits, a GPU budget, and web access — it handled literature reviews and ran hundreds of experiments but committed to unpromising approaches, couldn't backtrack, and made no novel contribution.
- Anthropic cofounder Jack Clark wrote in his Import AI newsletter that AI systems' lack of creativity is a "bearish signal on short recursive self-improvement timelines," echoing what the company found internally when trying to automate AI safety research.
- Kapoor attributes the gap to training methods: reinforcement learning works well on tasks with checkable answers but is hard to apply to open-ended research requiring judgment and taste.
- The team is now repeating the experiment with Mythos, Anthropic's most advanced model launched in April, which the Trump administration subsequently restricted to approved organizations under new safety requirements.
- OpenAI has made building an automated AI researcher an explicit goal, while Anthropic identifies self-improving AI as the industry's next milestone — though the study suggests current systems are nowhere near delivering it.
Why it matters: Recursive self-improvement is the premise underwriting massive AI infrastructure spending and the race between OpenAI, Anthropic, and others. If current models can't do the open-ended research that produced transformative advances like the transformer, the timeline for self-accelerating AI stretches out — and the competitive moat may belong to whichever lab solves the creativity gap first.
Ask SkimNews



