metir
metir
Docs
Download on App StoreGet it on Google PlayLog inSign up
Back to Blog
Healthcare AI
FDA
Medical Devices
Regulation
Radiology

FDA-Cleared AI Medical Devices: The Patient Outcomes Gap

A PLOS Digital Health audit found only 3 of 1,357 FDA-cleared AI medical devices were evaluated for patient outcomes. Here is what it means for hospitals.

Metir AI TeamOctober 5, 20266 min read
FDA-Cleared AI Medical Devices: The Patient Outcomes Gap

FDA-cleared AI medical devices are now a routine part of hospital technology, but a new audit asks a simple question that the clearance process does not require anyone to answer: do patients actually do better? A study published in PLOS Digital Health on August 19, 2026 examined all 1,357 AI and machine learning enabled devices that had received U.S. FDA clearance or approval through December 5, 2025. Only 3 of them, 0.2%, had been evaluated for patient-centered outcomes such as mortality, morbidity or readmissions.

1,357AI/ML devices analyzedCleared or approved through Dec 5, 2025
34 (2.5%)Linked to registered prospective trials
12 (0.9%)Posted results, and 12 with peer-reviewed publications
3 (0.2%)Evaluated patient-centered outcomes

What the FDA-cleared AI medical devices audit found

The paper, by Rawan Abulibdeh and colleagues from the University of Toronto, MIT Critical Data and several other institutions, combined the FDA device database with the American College of Radiology Data Science Institute catalogue. The authors then searched ClinicalTrials.gov and PubMed for linked trials and publications. Their count of the funnel is stark.

From 1,357 cleared devices to 3

Each step is a separate lens on the same device population, not a strict subset of the one above it.

1
Cleared or approved AI/ML devices1,357(100%)
Every AI/ML-enabled device the authors analyzed through December 5, 2025.
2
Linked to a registered prospective trial34(2.5%)
Found through ClinicalTrials.gov identifiers on FDA 510(k) summary pages.
3
Posted trial results12(0.9%)
Results posted to the registry.
4
Peer-reviewed publication12(0.9%)
Linked to a published, peer-reviewed manuscript.
5
Evaluated patient-centered outcomes3(0.2%)
Mortality, morbidity or readmissions as the primary endpoint.

Source: Abulibdeh et al., PLOS Digital Health, 2026. Bar width is proportional to device count, with a minimum width so small values stay visible.

Several details matter for interpretation:

  • Radiology dominates. Radiology accounts for 1,059 of the 1,357 devices (78%), yet the authors report prospective trials for fewer than 1% of those tools. Cardiovascular and neurology devices had trial rates of roughly 9.5% and 9.7%.
  • The trials that exist are small and narrow. Across the studies analyzed, 62% used observational designs. Among the 34 registered trials, about 73.5% enrolled fewer than 500 participants, and 68% ran only in the United States. Only 9 of 34 reported any subgroup analysis.
  • Industry runs most of them. 32 of the 34 trials (94%) were industry-led, and only 3 (9%) tested a therapeutic rather than diagnostic or screening function.
  • Some groups are routinely excluded. The authors report frequent exclusion of pregnant patients, non-English speakers and pediatric populations.

The authors also list limitations. Public registries can undercount proprietary manufacturer studies, not every 510(k) summary page was accessible, and the work did not look for unregistered validation studies. The count is a floor for registered evidence, not proof that nothing else exists.

“

AI tools must be life-tested before they can be called life-saving.

Abulibdeh et al., PLOS Digital Health, 2026

How FDA 510(k) clearance and substantial equivalence work

Most AI devices reach the U.S. market through the 510(k) pathway, with some going through De Novo classification or premarket approval. Under 510(k), a manufacturer shows that its device is substantially equivalent to a legally marketed predicate. FDA defines equivalence as the same intended use plus either the same technological characteristics, or different characteristics that do not raise different questions of safety and effectiveness.

Clinical data are not always required. FDA says performance data can include clinical data as well as non-clinical bench testing such as software validation and engineering performance testing. The audit's press release notes that most AI devices need only substantial equivalence, so developers are not required to demonstrate improved health outcomes.

That design made sense for devices whose function a predicate could easily frame. For software that interprets images or predicts risk, equivalence to a predicate can establish that a tool does a similar job, which is a different matter from showing that using it changes what happens to patients.

Accuracy evidence versus outcome evidence

The distinction at the center of the paper is between two kinds of proof:

  • Accuracy evidence asks whether the algorithm gets the answer right: sensitivity, specificity or agreement with expert readers on a test set.
  • Outcome evidence asks whether care improved: fewer deaths, fewer complications, fewer readmissions, better function or quality of life, measured in a prospective study.

A tool can be accurate and still leave outcomes unchanged, for example if clinicians do not act on its output, if it finds disease that would never have caused harm, or if it works worse in a hospital population unlike its test data. The authors classified an outcome as patient-centered only when the primary endpoint directly reflected mortality, major morbidity, hospitalization or readmission, validated quality of life, functional status or daily symptom burden.

FDA's AI-enabled device list and PCCP guidance

FDA maintains a public AI-Enabled Medical Device List, which it describes as a resource to identify AI-enabled devices authorized for marketing in the United States. The agency states the list is not comprehensive: devices were identified primarily from AI-related terms in their summary descriptions. Researchers and hospitals using it as a census should read it with that caveat in mind.

Exterior of FDA Life Sciences Laboratory I, Building 64, on the White Oak campus in Silver Spring, Maryland
FDA Life Sciences Laboratory I (Building 64) on the White Oak campus in Silver Spring, Maryland, which houses the Center for Devices and Radiological Health. Photo: U.S. Food and Drug Administration, public domain.

In December 2024, FDA finalized guidance on predetermined change control plans (PCCPs) for AI-enabled device software functions. A PCCP lets a manufacturer get certain future modifications pre-authorized in the original submission, so that qualifying updates do not each need a new marketing submission. Law firm summaries describe three components: a description of planned modifications, a modification protocol covering development, validation and implementation, and an impact assessment of benefits and risks.

PCCPs address how a cleared model can be updated safely. They are a lifecycle tool, and the audit's question is a different one: what evidence exists at the starting point and whether it includes patient outcomes. Both matter, and one does not substitute for the other.

Implications for hospitals

The paper's own conclusion is that readiness should not be defined by FDA clearance alone. Translating that into practice, hospitals and health systems may want to consider several steps. These are our analysis rather than recommendations from the study.

  • Ask vendors the outcome question directly. Request any prospective study, registered trial or publication, and whether its primary endpoint was a patient outcome or a diagnostic accuracy metric.
  • Validate locally. The audit found most trials were small, mostly U.S.-based and light on subgroup analysis, so performance on your own population and workflows is an open question.
  • Plan monitoring after go-live. If a vendor operates under a PCCP, ask what changes are pre-authorized and how updates will be communicated and re-tested locally.
  • Match the claim to the evidence. A tool cleared for triage or detection supports a workflow argument, not automatically a clinical-benefit argument.
  • Watch the global angle. The authors note that FDA clearance often acts as a gateway for deployment abroad, which raises the stakes for validation in lower-resource settings.

The authors call for a staged evidence framework, beginning with a pre-clearance phase of retrospective validation on diverse, representative datasets with mandatory demographic reporting. Whether regulators adopt anything similar is an open policy question. For now, the practical takeaway is narrow and measurable: clearance tells you a device may be marketed, and outcome evidence, where it exists, tells you whether it helps.

For related reading on imaging AI evidence, see our coverage of an open-source radiology AI model compared with radiologists.

Sources:

  • Abulibdeh R. et al., PLOS Digital Health, August 19, 2026: 1,357 AI medical devices cleared, 3 actually tested on patient outcomes (DOI 10.1371/journal.pdig.0001597)
  • EurekAlert press release: Most AI medical devices cleared for use were not tested on patient outcomes
  • FDA: Premarket Notification 510(k)
  • FDA: List of Artificial Intelligence-Enabled Medical Devices
  • Jones Day: FDA's Final Guidance Provides Practical Approach for AI-Enabled Devices Implementing Post-Market Modifications
  • Ropes & Gray: FDA Finalizes Guidance on Predetermined Change Control Plans for AI-Enabled Medical Device Software

Image credits

  • Hero: Main entrance of FDA Building 1, White Oak campus, Silver Spring, Maryland. Author: U.S. Food and Drug Administration. Source: Wikimedia Commons, File:FDA Entrance (16792957331).jpg. Licence: public domain (U.S. government work).
  • In-body: FDA Life Sciences Laboratory I, Building 64. Author: U.S. Food and Drug Administration. Source: Wikimedia Commons, File:FDA Bldg 64 - Exterior (5161375466).jpg. Licence: public domain (U.S. government work).

Ready to experience AI that adapts to you?

metir brings together the world's best AI models in one seamless experience. Start for free today.

Get Started Free
metir

Agentic Operating System for Professionals buried in meetings, emails and docs.

© 2026 metir. All rights reserved.

Product

  • Features
  • Pricing
  • Research
  • Docs
  • Blog
  • Enterprise

Company

  • Docs
  • Support
  • Careers

Legal

  • Terms of service
  • Privacy policy

Personalisation is powerful. Privacy is non-negotiable.

Status: All systems operational