Clinical AI Lost to General Models in NYU Head-to-Head

Get the Health newsletter
Daily health & science — research, biotech, public health, the studies worth knowing. Free.
- NYU Langone Health researchers tested general and clinical AI models—including OpenEvidence and UpToDate Expert AI—on three sets of clinical questions, finding clinical AI performed worse than general models, according to results published in Nature Medicine in June.
- OpenEvidence, Doximity, and UpToDate pitch their clinical large language models to hundreds of thousands of U.S. doctors as an antidote to hallucination-prone generalist models from Big Tech, even though few studies had previously pitted the tools against each other.
- Kaiser Permanente's vice president of AI and emerging technologies wrote on LinkedIn that he had "never seen a single paper trigger the kind of reactions this one has in the health AI community," underscoring the outsized response.
- Many online reactions treated the findings as a clear victory for general frontier models—though the article notes the paper's results, like all science, remain subject to interpretation and debate.
- The study marks one of the first head-to-head comparisons of clinical AI systems, arriving just as adoption of these specialized tools accelerates across U.S. medical practice.
Why it matters: Clinical AI products are sold to hundreds of thousands of doctors on the promise that specialization reduces dangerous hallucinations compared to general-purpose models. If general models actually outperform clinical ones on clinical questions, the value proposition and safety argument underpinning the entire specialized medical AI market is called into question.



