Harvard AI Outperforms Doctors in ER Triage Study

Get the Health newsletter
Daily health & science — research, biotech, public health, the studies worth knowing. Free.
- Harvard researchers published a Science study in which OpenAI's o1 reasoning model diagnosed 67% of 76 Boston ER cases correctly, compared with 50-55% accuracy for paired human doctors reading the same electronic records.
- OpenAI's o1 model rose to 82% accuracy when given more clinical detail, versus 70-79% for expert humans — though the study noted this wider-information gap was not statistically significant.
- The AI scored 89% on long-term treatment plans across five case studies, far outpacing the 34% achieved by 46 doctors working with conventional resources like search engines.
- Lead authors Arjun Manrai of Harvard Medical School and Dr. Adam Rodman of Beth Israel Deaconess framed the findings as pointing to a 'triadic care model' of doctor, patient, and AI — not physician replacement.
- In one lupus case the AI caught what humans missed: a patient with a blood clot whose lung inflammation was driven by autoimmune history rather than failing anti-coagulants.
- The study tested only text-based patient data, not visual or behavioral cues like distress levels, meaning the AI functioned as a paperwork-based second opinion rather than a bedside substitute.
- Dr. Wei Xing of the University of Sheffield flagged a concern absent from the headline: doctors may unconsciously defer to AI answers rather than think independently, a risk likely to grow as clinical AI use expands.
Why it matters: With one in five US physicians and 16% of UK doctors already using AI clinically, hospitals need accountability rules fast — Dr. Rodman admitted 'there is not a formal framework right now' for AI error, and the 89%-to-34% treatment-plan gap suggests current clinical resources are leaving significant quality on the table.




