Clinical AI Chatbots Lost to General AI in NYU Study

Get the Health newsletter
Daily health & science — research, biotech, public health, the studies worth knowing. Free.
- NYU Langone Health researchers tested general and clinical AI models — including OpenEvidence and UpToDate Expert AI — on three sets of clinical questions, publishing in Nature Medicine in June that clinical models performed worse than general ones.
- OpenEvidence, Doximity, and UpToDate pitch their clinical LLMs as antidotes to hallucination-prone generalist models from Big Tech, and hundreds of thousands of U.S. doctors already use them.
- Kaiser Permanente's vice president of AI and emerging technologies wrote on LinkedIn he'd never seen a single paper trigger the kind of reactions this one did in the health AI community.
- The source notes the paper's findings are "subject to interpretation and debate," yet many online discussions treated the results as a clear victory for general frontier models.
- Few studies had previously pitted clinical LLMs against each other before this summer's high-profile head-to-head.
Why it matters: Hundreds of thousands of U.S. doctors rely on clinical LLMs sold as safer alternatives to general AI — OpenEvidence, UpToDate Expert AI, and Doximity market themselves specifically as antidotes to hallucination. A Nature Medicine paper from NYU Langone finding those specialized tools scored worse than general-purpose models on three sets of clinical questions undermines the core sales pitch and reframes a market that has so far operated on limited head-to-head evidence.



