- Claims about AI in healthcare are not equally supported. The gap between a randomised screening trial and a vendor slide about a retrospective dataset is the whole story.
- Breast screening has the strongest evidence: the MASAI trial randomised 105,934 women and reported 80.5 percent sensitivity with AI support against 73.8 percent for standard double reading, at identical specificity.
- Deterioration prediction has the weakest. An external validation of a widely deployed proprietary sepsis model found 33 percent sensitivity and a 12 percent positive predictive value, with alerts on 18 percent of hospitalisations.
- The arithmetic that explains most disappointment: a test with 95 percent sensitivity and 90 percent specificity, run on a condition with 1 percent prevalence, is wrong about nine times out of ten when it fires.
- Ambient documentation is real but modest. A study of 1,800 clinicians measured 16 minutes saved per eight hours of care, while burnout improvements in smaller studies look much larger than the time savings alone would explain.
- FDA authorisation says a device met premarket requirements for its stated intended use. It does not say the device was tested in your population, and the agency says its own list is not comprehensive.
Most published claims about AI in healthcare are true and misleading at the same time. They are true because the number was measured. They are misleading because the reader assumes the number came from a trial when it usually came from a dataset, and assumes the dataset resembles their patients when it usually does not.
This piece grades the evidence rather than the technology. For each major application, the question is not whether it works. It is what kind of study supports it, what went wrong when somebody deployed it, and what the regulator has actually said. Those three answers put a claim in its place faster than any accuracy figure.
How should you grade a claim about clinical AI?
By its design, not its headline. There is a ladder, and almost every disagreement about medical AI is really a disagreement about which rung a particular result sits on.
At the top sits the prospective randomised trial: patients are allocated, the algorithm runs in real time, and outcomes are measured on people. One rung down is prospective observational work, where the model runs live but nobody randomised anything. Below that, retrospective validation on data the developer did not curate. At the bottom, retrospective validation on data the developer did curate, which is where most marketing numbers come from.
The reason the ladder matters is that performance falls as you climb it. A model tuned on a curated retrospective set is being asked an easier question than the same model watching a live ward, where the data arrives late, the labels are ambiguous and the population is not the one in the paper.
What does the evidence look like application by application?
The table below is the core of this piece. Every row states the strongest published evidence type I could find for that application, the failure mode observed once it was deployed, and where regulation sits. It is assembled from the sources linked throughout, none of which publish it in this form.
| Application | Strongest published evidence | Failure mode seen in deployment | Regulatory position |
|---|---|---|---|
| Breast screening triage | Prospective randomised, population based (MASAI, 105,934 women) | Still needs at least one radiologist, so the saving is workload rather than headcount | Cleared devices exist; imaging dominates the authorised list |
| Sepsis and deterioration prediction | Retrospective external validation, and it was unflattering | 33 percent sensitivity, 12 percent positive predictive value, alerts on 18 percent of hospitalisations | Often deployed as a health system tool rather than a cleared device |
| Ambient documentation | Prospective observational and quality improvement, no randomised control in the largest study | Time saved is real but small, and unevenly distributed across specialties | Generally outside the device pathway, because it does not diagnose |
| Assistive diagnosis in general practice | Retrospective, plus survey evidence on adoption rather than outcomes | Skill erosion is the concern clinicians raise most, at 88 percent in the AMA survey | Varies by claim; a diagnostic claim triggers device rules |
| Coding and billing support | Operational, largely unpublished | Documented in vendor case studies rather than peer review | Not a device question, an audit question |
| Triage chatbots for patients | Thin. Mostly simulated vignettes rather than patients | Safe advice on average, with a long tail nobody has measured properly | Depends on whether the output is framed as advice |
Why is a 95 percent accurate model wrong most of the time?
Because accuracy is measured on a population and applied to a person, and the two are not the same arithmetic. This is the single calculation that explains more disappointed deployments than any other, so it is worth doing slowly with round numbers.
Take a model quoted at 95 percent sensitivity and 90 percent specificity. Run it on 10,000 patients for a condition with 1 percent prevalence. That means 100 patients have it and 9,900 do not.
- Of the 100 who have it, the model catches 95 and misses 5.
- Of the 9,900 who do not, it wrongly flags 10 percent, which is 990 people.
- So the model raises 1,085 alerts, of which 95 are correct.
The positive predictive value is 95 divided by 1,085, or roughly 8.8 percent. Nine out of ten alerts are false, from a model whose published numbers were excellent. Nothing was misrepresented. The prevalence did the damage.
This is why a pilot in a high acuity unit can look brilliant and the same model can be unusable on a general ward. The model did not change. The base rate did. Ask any vendor for the prevalence in the validation population before you ask for the accuracy.
Does that arithmetic show up in real deployments?
It does, with published numbers. The external validation of a widely implemented proprietary sepsis prediction model in JAMA Internal Medicine studied 27,697 patients across 38,455 hospitalisations at Michigan Medicine. Sepsis occurred in about 7 percent of them. At the recommended alerting threshold the model reached 33 percent sensitivity, 83 percent specificity, a positive predictive value of 12 percent and an area under the curve of 0.63.
Read those together and the operational picture is grim: the model missed two thirds of cases, generated alerts on 18 percent of all hospitalisations, and identified only 7 percent of the patients who missed timely antibiotics. The authors described poor discrimination and calibration, and named alert fatigue directly. This was not a fringe product. It was one of the most widely installed prediction models in American hospitals.
Where is the evidence genuinely strong?
Breast screening, and it is not close. The MASAI trial is the reference point because it did the expensive thing: it randomised.
As the ASCO Post reported on the final results, 105,934 women in Sweden were randomised to AI supported screen reading or to standard double reading. Sensitivity was 80.5 percent with AI support against 73.8 percent without, and specificity was identical at 98.5 percent in both arms. The interval cancer rate was 1.55 per 1,000 in the intervention group against 1.76 per 1,000 in the control group, and the cancers found in the AI arm were less often of the unfavourable kind.
Two details deserve as much attention as the headline. Sensitivity rose without specificity falling, which is the rare result that is not simply a threshold moved to a friendlier place. And the investigators stated plainly that AI supported screening still requires at least one human radiologist. The gain is workload, not replacement.
Does ambient documentation give time back, or move it?
It gives some back, less than the pitch implies, and the burnout effect looks larger than the clock effect. Those two facts sit uncomfortably together and both appear to be true.
On the clock side, STAT reported on a study of roughly 1,800 clinicians across five academic medical centres between 2023 and 2025. It measured 16 minutes of documentation time saved and 13 fewer minutes in the record for every eight hours of patient care, which the authors translated into roughly one extra patient every two weeks. There was no significant effect on time in the record outside working hours. Benefits were uneven: primary care physicians and female clinicians gained more than others. An earlier review had found savings under a minute per note.
On the wellbeing side the numbers look different. A quality improvement study across six health systems surveyed 263 ambulatory clinicians before and 30 days after adopting the same ambient scribe. Burnout fell from 51.9 percent to 38.8 percent. After hours documentation fell by 0.90 hours per week, about 11 minutes a workday, and self reported cognitive task load improved by 2.64 points on a ten point scale.
Notice the mismatch. Eleven minutes a day does not obviously explain a 13 point drop in burnout. The most plausible reading is that the burden being shifted is cognitive rather than temporal: not typing the note during the visit changes the experience of the visit more than it changes the length of the day. The authors of that study list their own limitations honestly, including no control group and self selection among early adopters, so treat the size of the effect as provisional and the direction as credible.
What does that mean for anyone buying one?
Buy it for attention, not for throughput. A business case built on seeing more patients will disappoint against a 16 minute figure. A business case built on retention of clinicians who were about to leave is harder to measure and probably closer to where the value sits.
What does regulatory clearance actually certify?
That a device met the applicable premarket requirements for a stated intended use, reviewed for safety and effectiveness. It does not certify performance in your population, your scanner, your workflow or your prevalence.
The FDA maintains a public list of AI enabled medical devices authorised for marketing in the United States. Two caveats on that page are worth quoting to anyone waving a clearance letter. First, the agency states that the list is not a comprehensive resource, because entries are identified largely by looking for AI related terms in the summary descriptions of authorisation documents. Second, the agency says it will explore methods to identify devices incorporating foundation models and large language models, which is an unusually clear signal that the current list is not built to answer the question everybody now asks.
The same evaluation discipline applies here as in any other domain where a model is scored on a benchmark and then deployed somewhere else. We wrote about the general version of this problem in a method for telling whether a change actually helped, and the medical case is the same logic with worse consequences: a metric measured on a fixed set tells you about the set.
How much of this is already happening anyway?
More than the evidence base would suggest, which is the uncomfortable part. The AMA's 2026 physician survey of nearly 1,700 doctors found that 81 percent use AI professionally, more than double the 2023 figure. The most common uses are summarising research and standards of care at 39 percent, drafting discharge instructions and care plans at 30 percent, and documentation of codes and notes at 28 percent. Assistive diagnosis sits far lower, at 17 percent.
That ordering is telling. Adoption is concentrated exactly where the evidence requirement is lowest and the reversibility is highest. A badly summarised paper is caught by the person reading it. A badly calibrated deterioration score is not.
The same survey records the reservations: 88 percent worry about skill loss, 88 percent want robust safety and efficacy validation, 86 percent put data privacy first, and 85 percent want to be consulted on adoption decisions. Clear liability rules ranked highest among the things that would build trust in oversight. Those are not the answers of a profession being swept along. They are the answers of one negotiating.
How do you read a vendor claim in one pass?
Five questions, in this order, and you can ask them in a meeting without any statistics background.
What was the design? If the answer contains the word retrospective, everything after it is a hypothesis. That is not an insult. Retrospective work is how you decide what to trial. It is simply not how you decide what to deploy.
Who assembled the data? A dataset curated by the developer answers a question the developer chose. An external set answers a question somebody else chose, which is much closer to the question your ward will ask.
What is the prevalence? Not the accuracy. The prevalence. If the validation population had a condition rate five times yours, the alert volume you inherit will not resemble the one in the paper, and the arithmetic above tells you exactly how it will differ.
What is the endpoint? Detection is not treatment and treatment is not outcome. A screening tool that finds more cancers has proven it finds more cancers. Whether that changes mortality is a separate study, usually an unfinished one.
Who owns the alert? Every deployed model creates work for somebody. If nobody can name the person who acts on a positive result and the time budget they have for it, the deployment is already failing and the model has not been switched on yet.
This is the same discipline that applies whenever a scored number has to survive contact with a real workload. We laid out the general form of it when arguing for a decision grid rather than a leaderboard when picking a model, and the medical version only differs in what a wrong answer costs. The grid wins because it forces the question of fit, and fit is where published accuracy quietly disappears.
What about the numbers vendors quote from published trials?
Check that the trial is theirs. A vendor citing a randomised trial of a different product in the same category is telling you about the category, not about the product. Imaging triage is a broad field with wide variation between systems, and the MASAI result belongs to the configuration MASAI tested, not to every algorithm that reads a mammogram.
Where is the evidence thin right now?
Three areas, stated plainly, because a piece that grades evidence has to grade the gaps too.
Patient facing triage. Most published work uses written vignettes rather than patients with real symptoms and real anxiety. Vignettes remove the two hardest parts of triage, which are incomplete history and the person who understates their pain.
Long horizon outcomes. Screening trials measure detection and interval cancers. Almost nothing measures whether a decision support tool changes mortality years later, because that study takes years and nobody funds it before the product ships.
Performance over time after deployment. A model validated in 2024 is running in a hospital whose coding practices, staffing and case mix have all moved. Drift monitoring is discussed far more often than it is published, and the FDA's own note about postmarket monitoring reflects that it is unfinished business.
What to re-check when this changes
Treat the following as the updatable part of this page. If you are returning to it later, these are the four things that will have moved.
- The authorised device list. Whether any device incorporating a foundation model or a large language model has been tagged as such. As of the FDA page cited here, the agency was still working out how to identify them.
- Randomised evidence outside imaging. Screening has its trial. Deterioration prediction, documentation and triage do not. The first credible randomised result in any of those changes the table above.
- Ambient scribe effect sizes with controls. The burnout numbers come mostly from before and after surveys. A controlled study either confirms the size or shrinks it.
- Liability rules. Physicians named clear liability frameworks as the top trust builder. Any movement there changes adoption faster than any accuracy improvement will.
None of this argues against using AI in healthcare. It argues for reading the claim before believing the number. The strongest result in this piece came from a randomised trial of more than a hundred thousand women and produced a seven point sensitivity gain with no specificity cost, which is a genuinely good outcome and an unglamorous one. The worst came from a widely installed model that alerted on nearly a fifth of hospital stays and was right about one alert in eight. Both were sold as AI in healthcare. Only one of them was graded before it was deployed, and the difference in how they turned out is not a coincidence.
If you are on the buying side of this, the practical protection is procedural rather than technical: ask for the prevalence in the validation population, ask whether the study was prospective, ask who is accountable when the alert is wrong, and hold the vendor to a documented answer. Those questions cost nothing. Our own view on how to handle data and access when tools like these enter an organisation sits on the MaShop security page, and the principle carries over: the controls you can describe are the only ones you actually have.