AI benchmarks fail in real-world teams, researcher argues

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Angela Aristidou, a professor at University College London and Stanford faculty fellow, argues that current AI benchmarks test models in isolation and miss how they actually perform within human teams and organizational workflows.
- Aristidou proposes 'HAIC benchmarks' (Human–AI, Context-Specific Evaluation), which shift the unit of analysis from individual task performance to team and workflow performance, expand the time horizon beyond one-off tests, and measure organizational outcomes and error detectability.
- In her research, FDA-approved radiology AI models with high benchmark scores—including 98% accuracy on technical tests—introduced delays in hospital radiology units because staff had to reconcile outputs with hospital-specific reporting standards and multidisciplinary team decision-making.
- A UK hospital system from 2021–2024 shifted its evaluation of a medical AI application from whether it improved diagnostic accuracy to how it affected coordination, deliberation, and collective reasoning within multidisciplinary teams.
- Over 18 months, a humanitarian-sector organization tracked an AI system's 'record of error detectability'—how easily human teams could identify and correct mistakes—to design context-specific guardrails and maintain trust.
- When highly scored AI fails to translate into real-world performance, it lands in what Aristidou calls the 'AI graveyard,' wasting time, money, and effort while eroding organizational confidence in AI and, in health settings, public trust in the technology.
Why it matters: If benchmarks continue to measure AI in sanitized isolation, organizations will keep deploying models that pass tests but fail in practice—wasting procurement budgets and, in healthcare, eroding public trust. Aristidou's case studies from hospitals and humanitarian groups show that evaluating AI within real workflows and over months, not minutes, surfaces coordination breakdowns and downstream inefficiencies that current scores never capture.




