We have spent years asking whether artificial intelligence can perform like a doctor. The more important question now is whether doctors and AI, working together, can deliver better care.
For much of the past decade, medical AI has been judged like a student sitting an examination.
Can it identify an abnormal scan?
Can it predict deterioration?
Can it recognize a disease as accurately as a clinician?
Can it outperform one?
These were necessary questions. Before introducing any technology into healthcare, we should know whether it can perform the task it claims to perform.
But medicine is not an examination hall.
A patient does not benefit because an algorithm achieved an impressive score on a benchmark. A patient benefits if something meaningful happens because of it: a cancer is detected earlier, deterioration is recognized sooner, unnecessary testing is avoided, treatment becomes safer, or a clinician gets enough time back to actually speak to the person sitting in front of them.
That distinction is becoming increasingly important.
A new Nature Medicine Comment argues that the next generation of medical AI should be judged not simply by whether an algorithm can match clinicians, but by whether carefully designed human–AI systems improve patient outcomes.
I think that is exactly where the conversation needs to go.
Accuracy is the entrance exam, not the final result
Imagine an AI system that detects a particular condition with extraordinary accuracy.
That sounds impressive.
Now imagine it produces so many alerts that clinicians begin ignoring them.
Or it works beautifully in the population on which it was trained but less reliably in another hospital.
Or it identifies something correctly, but there is no pathway for the healthcare system to act on that information quickly.
Or clinicians become overly dependent on its recommendation and stop challenging it when the context does not fit.
The algorithm may still be “accurate.”
The healthcare system may still not be better.
That is why evaluating medical AI cannot stop at sensitivity, specificity or an area-under-the-curve score.
Once an algorithm enters clinical practice, the intervention is no longer simply:
AI
It becomes:
AI + clinician + workflow + patient + healthcare system.
And every part of that equation can influence the eventual outcome.
We are beginning to see what better evaluation looks like
Breast-cancer screening offers an interesting example.
The large randomized MASAI trial in Sweden compared AI-supported mammography screening with standard double reading by radiologists.
Its latest results, published in The Lancet in 2026, found an interval-cancer rate of 1.55 per 1,000 participants in the AI-supported group versus 1.76 per 1,000 with standard screening, meeting the study’s criterion for non-inferiority.
Sensitivity was higher with AI-supported screening—80.5% versus 73.8%—while specificity was essentially the same at 98.5% in both groups. The AI-supported approach also reduced screen-reading workload.
That does not mean we can suddenly conclude that AI has solved breast-cancer screening.
Nor does it tell us everything about longer-term outcomes.
What makes the study important is something broader: AI was evaluated prospectively, inside an actual screening system, against clinically meaningful endpoints, rather than only being tested retrospectively on a curated dataset.
That is a much higher bar.
And healthcare AI needs more of it.
The unit we should be studying is not the machine. It is the partnership.
There is a recurring temptation in technology conversations to frame progress as a competition:
AI versus doctor.
I have never found that particularly useful.
The better question is:
What can the combination do that neither can do as well alone?
AI may be extraordinarily good at scanning enormous amounts of information, identifying patterns, prioritizing cases or performing repetitive tasks consistently.
Clinicians bring something different: context, judgment, communication, responsibility and the ability to understand when a patient’s situation does not fit neatly into the data.
These strengths are complementary.
But there is another side to that partnership.
Humans can over-trust automation.
We can also under-trust it.
We can misunderstand why a recommendation was produced.
We can become fatigued by alerts.
An algorithm can technically perform beautifully while being incorporated into clinical practice badly.
That means human factors are part of AI safety.
If we want to understand whether medical AI works, we need to study the behaviour of the entire human–technology system.
So what should the next generation of medical-AI studies measure?
I would like to see us become much more demanding.
Instead of asking only:
“How accurate is the model?”
we should increasingly ask:
Did it help detect disease earlier?
Did fewer diagnoses get missed?
Did patients receive treatment sooner?
Did it reduce unnecessary investigations?
Did it improve clinician workload without compromising safety?
Did it improve access to expertise?
Did the benefit persist outside the institution where the algorithm was developed?
Did it work equally well across age groups, sexes, ethnicities and different levels of disease prevalence?
Did clinicians know when to disagree with it?
And ultimately:
Did patients do better?
Those questions are much harder than running an algorithm against a benchmark.
They are also the questions medicine actually cares about.
The efficiency question matters too
Not every valuable medical technology will extend life directly.
Sometimes the gain is time.
If AI safely reduces hours spent on repetitive screening, documentation or administrative work, that can matter enormously.
But even here, we need to measure the downstream effect.
Did the saved time translate into shorter waiting lists?
More patients being seen?
More thoughtful consultations?
Less clinician burnout?
Or did the healthcare system simply find another administrative task to fill those minutes?
Efficiency in isolation is not automatically better healthcare.
The value lies in what we do with the efficiency we create.
This matters enormously for countries such as India
In a healthcare system serving a population as large and diverse as India’s, AI has obvious potential.
It could help prioritize imaging.
Support clinicians where specialists are scarce.
Flag deterioration.
Assist with documentation.
Interpret large streams of diagnostic information.
Extend expertise beyond major metropolitan centres.
But scale magnifies both benefit and error.
A system that is slightly wrong but used a few hundred times creates one kind of problem.
A system that is slightly wrong but used millions of times creates another.
This is why validation across different populations, hospitals, devices and clinical environments cannot be treated as an afterthought.
The technology needs to work not merely in the laboratory where it was created, but in the messy reality of healthcare.
We should resist two extremes
The AI conversation in medicine tends to drift toward opposite poles.
One says:
“AI will replace doctors.”
The other says:
“AI cannot possibly understand medicine.”
Neither position is particularly interesting.
AI will undoubtedly become deeply embedded in healthcare.
It already is.
Predictive algorithms can identify patients at risk of deterioration, computer-vision systems can help interpret images and ambient systems can assist with clinical documentation. But researchers have rightly pointed out that despite accelerating adoption, we still often lack strong evidence connecting an AI intervention directly to better clinical outcomes.
So the task ahead is not to prove that AI is impressive.
We already know it can be.
The task is to determine where it genuinely makes medicine better—and where it does not.
The future of medicine should not be technology-first
It should be patient-first.
If an AI system helps a doctor catch something earlier, wonderful.
If it allows a radiologist to concentrate attention on the studies that need it most, valuable.
If it reduces administrative burden so clinicians can spend more time with patients, important.
If it expands access to expertise in places where specialists are limited, potentially transformative.
But if it merely adds another dashboard, another alert and another layer of complexity without improving care, then sophistication alone is not progress.
Medicine has always adopted technology.
The stethoscope was technology.
Cardiopulmonary bypass was technology.
CT, MRI, robotic surgery and minimally invasive procedures were technologies.
The question was never whether the technology was exciting.
The question was whether it enabled us to care for people better.
AI deserves exactly the same standard.
Accuracy tells us whether the machine can perform.
Outcomes tell us whether medicine actually improved.
And that, ultimately, is the test that matters.
References
Kristina Lång. From algorithms to patient outcomes — lessons from one of the first randomized trials of AI in medicine. Nature Medicine, September 7, 2026.
Gommers J, Hernström V, Josefsson V, et al. Interval cancer, sensitivity, and specificity comparing AI-supported mammography screening with standard double reading without AI in the MASAI study. The Lancet, 2026.
Goldenberg A, Wiens J. Is AI actually improving healthcare? Nature Medicine, April 2026.


