Clinical AI Chatbots Underperformed General Models in Study

Get the Health newsletter
Daily health & science — research, biotech, public health, the studies worth knowing. Free.
- OpenEvidence, Doximity, and UpToDate market their clinical large language models to hundreds of thousands of U.S. doctors as safer alternatives to hallucination-prone Big Tech generalist models.
- NYU Langone Health researchers tested general and clinical models — including OpenEvidence and UpToDate Expert AI — on three sets of clinical questions, with results published in Nature Medicine in June.
- The study found that the clinical AI models performed worse than the general models, contradicting the central marketing pitch of the clinical-AI category.
- Kaiser Permanente's vice president of AI and emerging technologies wrote on LinkedIn: "I've never seen a single paper trigger the kind of reactions this one has in the health AI community."
- The paper's findings, like all science, are subject to interpretation and debate — but many online reactions treated the results more like a clear victory for general frontier models.
- Few studies had previously pitted clinical chatbots against each other, making this summer's head-to-head a rare empirical test in a fast-growing market.
Why it matters: Hundreds of thousands of U.S. doctors rely on clinical AI tools whose core selling point — safety versus hallucination-prone Big Tech models — was directly challenged by a peer-reviewed study, and the online reaction treated the result as decisive before scientific debate could play out.



