AI Agents Failed Most Real-World Workplace Tasks

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Anthropic's study predicted managers, architects, and media workers will be most affected by LLMs while groundskeepers, construction workers, and hospitality staff will be least affected—but the article notes these are guesses based on which tasks LLMs seem suited to, not how they perform in actual workplaces.
- Mercor's February study tested AI agents powered by top-tier models from OpenAI, Anthropic, and Google DeepMind on 480 workplace tasks performed by bankers, consultants, and lawyers; every agent failed to complete most of its duties.
- OpenAI chief scientist Jakub Pachocki called AI an "economically transformative technology," but the article observes boosters are hazy on the actual path from current capabilities to that promised future.
- The hype-reality gap extends well beyond coding: most predictions rest on rapid coding-tool improvement, yet LLMs are bad at strategic judgment calls and can't be dropped into messy, human-laden workflows without sometimes making things worse.
- Pause AI, an international activist group, co-organized a February anti-AI march in London and distributed flyers calling for a pause until the path from AI technology to transformation—what they call "Step 2"—is understood.
- The information vacuum around AI's actual workplace impact is so wide that a single social media post can shake markets, the article argues, because no one yet has evidence-based answers to how the technology will be deployed.
Why it matters: The Mercor study tested 480 real workplace tasks across banking, consulting, and law using leading models from OpenAI, Anthropic, and Google DeepMind—and every agent failed most duties. That gap between hype and demonstrated capability means companies and workers are making bets on AI transformation based largely on guesses about what tasks LLMs might handle, leaving the global economy's AI promises resting on unverified claims rather than reproducible results.



