AI Health Chatbots Launch Without Independent Testing

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Microsoft launched Copilot Health earlier this month, with VP Dominic King saying the company fields 50 million health questions daily on Copilot and that health is the platform's most popular mobile topic.
- Amazon opened Health AI to the general public this month, expanding a tool previously limited to One Medical members, joining OpenAI's ChatGPT Health (released in January) and Anthropic's Claude, which can access health records with permission.
- Mount Sinai researchers published a study finding ChatGPT Health sometimes recommends excessive care for mild conditions and fails to identify emergencies — a finding OpenAI's Karan Singhal disputes but that all six academic experts interviewed said reflects a broader lack of independent pre-release testing.
- OpenAI's HealthBench benchmark uses LLM-generated conversations, and a separate Oxford study led by doctoral candidate Andrew Bean found that non-expert users working with an LLM identified medical conditions from a scenario only about a third of the time.
- OpenAI's GPT-5.4 performs worse at soliciting contextual information from users than the earlier GPT-5.2, according to the company's own reporting, a weakness Bean flagged as especially risky for patients without medical expertise.
- Google's AMIE medical chatbot matched physicians' diagnostic accuracy in a controlled study with real patients, but the company said it has no plans to release the tool publicly pending further equity, fairness, and safety research.
Why it matters: Microsoft alone says its Copilot app fields 50 million health questions a day, so the gap between company self-evaluation and independent testing now reaches a massive patient base. The Mount Sinai and Oxford studies both show these tools can over-recommend care, miss emergencies, and fail non-expert users in ways that compound when patients lack medical literacy.


