Health

Can you trust an AI chatbot's medical advice? What new benchmarks show

STAT News2 h ago
A doctor looking at a computer screen in a clinic office
A doctor looking at a computer screen in a clinic officePhoto: Tima Miroshnichenko / Pexels

A few years ago, AI chatbots were a curiosity in medicine. Today doctors consult them to draft notes, patients describe their symptoms to them, and hospital systems are deploying them to speed up triage decisions. Behind that rapid adoption sits an uncomfortable question: how do we actually know how reliable these systems are?

In theory, the answer should be benchmarks — standardized tests measuring how often an AI model gets medical questions right. But some developers in the industry argue the benchmarks themselves are seriously flawed. A model can score well on a multiple-choice medical exam and still fall short when faced with a real patient's messy, contradictory account of their symptoms.

Part of the problem comes from comparing general-purpose large language models against systems purpose-built for medicine using the same yardstick. A general chatbot, trained on a vast pool of text, can produce fluent, confident-sounding answers — and that fluency can be mistaken for medical accuracy. Clinically focused systems, trained on narrower datasets, may hedge more — which can make them look "less helpful" on the same tests.

That distinction carries real consequences outside the lab. A chatbot that helps a doctor broaden a differential diagnosis and an app that talks directly to a patient assessing their own symptoms carry very different risk profiles. An error in the first case gets filtered through an experienced clinician; an error in the second can reach the patient directly.

There's a growing consensus among developers that current benchmarking systems fail to capture the messiness of real clinical settings. Standard tests tend to feature clean, single-correct-answer questions — but actual medical practice is full of incomplete information, conflicting symptoms and time pressure. A model acing exam questions doesn't reveal how it handles uncertainty.

Some researchers are proposing safety-focused tests instead: measuring when a model knows to say "I don't know" or to refer a patient to a human doctor, rather than just its raw accuracy rate. That's a harder capability to measure, but a far more clinically important one.

On the patient side, the stakes are more concrete still. Someone describing symptoms to a chatbot can't always tell whether the response they're getting is expert judgment or a statistical guess. That ambiguity becomes especially dangerous with rare conditions or situations requiring urgent care.

Health systems adopting these tools are trying to strike a careful balance: capturing efficiency gains without compromising patient safety. Some hospitals use AI strictly as a "second opinion" layer, leaving the final call to a clinician in every case.

Experts broadly agree the industry needs shared, independently audited benchmarking standards going forward. Without them, every company designs its own test and declares its own success — making it nearly impossible for patients and doctors to know how much to trust any given system.

For now, the most practical advice for patients is this: treat medical information from a chatbot as a starting point, not the final word. The technology may be advancing quickly, but the tools built to certify it as safe haven't yet caught up.

This article is an AI-curated summary based on STAT News. The illustration is a stock photo by Tima Miroshnichenko from Pexels.

Read next