Can you trust an AI chatbot's medical advice? What new benchmarks show

A few years ago, AI chatbots were a curiosity in medicine. Today doctors consult them to draft notes, patients describe their symptoms to them, and hospital systems are deploying them to speed up triage decisions. Behind that rapid adoption sits an uncomfortable question: how do we actually know how reliable these systems are?
In theory, the answer should be benchmarks — standardized tests measuring how often an AI model gets medical questions right. But some developers in the industry argue the benchmarks themselves are seriously flawed. A model can score well on a multiple-choice medical exam and still fall short when faced with a real patient's messy, contradictory account of their symptoms.
Part of the problem comes from comparing general-purpose large language models against systems purpose-built for medicine using the same yardstick. A general chatbot, trained on a vast pool of text, can produce fluent, confident-sounding answers — and that fluency can be mistaken for medical accuracy. Clinically focused systems, trained on narrower datasets, may hedge more — which can make them look "less helpful" on the same tests.
That distinction carries real consequences outside the lab. A chatbot that helps a doctor broaden a differential diagnosis and an app that talks directly to a patient assessing their own symptoms carry very different risk profiles. An error in the first case gets filtered through an experienced clinician; an error in the second can reach the patient directly.
There's a growing consensus among developers that current benchmarking systems fail to capture the messiness of real clinical settings. Standard tests tend to feature clean, single-correct-answer questions — but actual medical practice is full of incomplete information, conflicting symptoms and time pressure. A model acing exam questions doesn't reveal how it handles uncertainty.
Some researchers are proposing safety-focused tests instead: measuring when a model knows to say "I don't know" or to refer a patient to a human doctor, rather than just its raw accuracy rate. That's a harder capability to measure, but a far more clinically important one.
On the patient side, the stakes are more concrete still. Someone describing symptoms to a chatbot can't always tell whether the response they're getting is expert judgment or a statistical guess. That ambiguity becomes especially dangerous with rare conditions or situations requiring urgent care.
Health systems adopting these tools are trying to strike a careful balance: capturing efficiency gains without compromising patient safety. Some hospitals use AI strictly as a "second opinion" layer, leaving the final call to a clinician in every case.
Experts broadly agree the industry needs shared, independently audited benchmarking standards going forward. Without them, every company designs its own test and declares its own success — making it nearly impossible for patients and doctors to know how much to trust any given system.
For now, the most practical advice for patients is this: treat medical information from a chatbot as a starting point, not the final word. The technology may be advancing quickly, but the tools built to certify it as safe haven't yet caught up.
Read next

Could a daily multivitamin help older adults stay independent longer?
In a three-year trial, older adults taking a daily multivitamin better preserved the strength and stamina needed for everyday activities compared with those on a placebo. The findings haven't been peer-reviewed yet, but they add a new dimension to the long-running debate over supplements.

Does intermittent fasting protect the aging brain? What a new trial found
In a small six-month trial, older women who confined their eating to an 8-to-9-hour daily window performed better on planning and problem-solving tests than those who ate over 12 hours — even though both groups lost similar amounts of weight. The findings suggest timing may matter as much as quantity.

Traffic noise and Parkinson's disease: what a 3-million-person study found
A pioneering study that tracked more than three million people in Denmark over 18 years has found a consistent link between traffic noise and Parkinson's disease risk. Here's what the research found, and why it should change how we think about noise pollution.

What is the UK's long-running social care funding crisis, and why hasn't it been fixed?
The Conservative Party leader has written to the prime minister offering to work together on reforming adult social care, while urging him to rule out tax rises to fund it. Here's the story behind a funding crisis that has gone unresolved in England for three decades.

What is the new hybrid PhD model connecting academia and industry?
The US National Science Foundation is piloting a new PhD model that gives doctoral trainees both academic and industry research experience. Students in fields eligible for NSF funding — including chemistry, engineering and basic life science research not directly tied to disease treatment — will be able to take part.