Most AI medical devices cleared for use were not tested on patient outcomes
Of 1,357 devices authorized by the US FDA, only 3 were evaluated on clinical effectiveness `
image:
Growth of FDA-cleared AI/ML-enabled medical devices from 1995 to December 2025. Of 1,357 cleared devices, only 34 were linked to registered clinical trials and only 3 were evaluated for patient-centered outcomes. (Fig 1 of the article.)
view more
Credit: Abulibdeh R, Cajas Ordóñez SA, Celi LA, Gorijavolu R, Izath N, Markussen Lunde T, 2026, PLOS Digital Health, CC-BY 4.0 (https://creativecommons.org/licenses/by/4.0/)
A new analysis shows that, of 1,357 artificial intelligence (AI)-based medical devices authorized by the U.S. Food and Drug Administration (FDA) for use in patient care, only three had been tested on whether they actually improve patients’ health. Rawan Abulibdeh of the University of Toronto, Canada, and colleagues present these findings in the open access journal PLOS Digital Health on August 19, 2026.
New AI devices increasingly inform clinical care, such as systems that aid surgical planning, calculate cardiovascular risks, and guide interpretation of mammograms and other imaging. In order to be authorized for use in the U.S., AI devices typically only need to show “substantial equivalence” to an existing authorized device, and developers are not required to demonstrate whether new AI devices help people live healthier lives—with benefits shared equitably across diverse subgroups.
To deepen understanding of this topic, Abulibdeh and colleagues investigated how all 1,357 AI devices authorized by the FDA as of December 5, 2025, had been evaluated in patients prior to authorization.
They found that only 34 of the devices had been included in registered clinical trials, with results posted for 12 and peer-reviewed manuscripts published for 12. Only 3 devices had been tested on patient-centered outcomes, such as death rates, strokes, hospitalizations, and quality of life. Most studies were conducted in highly resourced healthcare systems, and most excluded key patient subgroups, such as pregnant women, adults over 75, and non-English speakers.
The researchers suggest that structural barriers such as financial incentives and logistical challenges discourage developers from testing AI devices on patient outcomes, resulting in greater emphasis on speedy development than on rigor. They discuss how this framework could allow new tools to amplify existing disparities in healthcare and how it could lead to patients in low- and middle-income countries becoming inadvertent test populations for under-studied AI devices, as many countries rely on higher-income countries’ authorization decisions.
On the basis of their findings, the researchers conclude that existing policies for AI medical device authorization should be redesigned. They propose a novel, three-phase framework that includes demonstration of effectiveness across diverse patient subgroups and healthcare settings.
The authors add: “We expected the evidence base to be thin, but not this thin. Out of 1,357 AI devices the FDA has cleared for use in patient care, only three have been tested on whether patients actually live longer or better. Clearance tells you a device resembles something already on the market. It does not tell you it helps anyone.”
In your coverage please use this URL to provide access to the freely available article in PLOS Digital Health: https://plos.io/4xxUOGK
Citation: Abulibdeh R, Cajas Ordóñez SA, Celi LA, Gorijavolu R, Izath N, Markussen Lunde T (2026) 1,357 AI medical devices cleared, 3 actually tested on patient outcomes. PLOS Digit Health 5(8): e0001597. https://doi.org/10.1371/journal.pdig.0001597
Author Countries: Canada, Norway, Uganda, United States
Funding: LAC is funded by the National Institute of Health through DS-I Africa U54 TW012043-01 and Bridge2AI OT2OD032701, the National Science Foundation through ITEST 2148451, and a grant of the Korea Health Technology R&D Project through the Korea Health Industry Development Institute (KHIDI), funded by the Ministry of Health & Welfare, Republic of Korea (grant number: RS-2024-00403047). RG is supported by the Johns Hopkins Institute for Clinical and Translational Research (ICTR) and Grant T32TR004928 from NCATS, a component of the National Institutes of Health. The contents are solely the responsibility of the authors and do not necessarily represent the official views of the Johns Hopkins ICTR, NCATS, or the NIH. TML is funded by the consortium’s owner institutions – the University of Bergen, Western Norway University of Applied Sciences, the Institute of Marine Research, the Norwegian School of Economics, and SIVA SF – together with competitive grants from SR Bank, DnB, Agenda Vestlandet, and Nora.fo. Use of AI/LLM: The authors used a large language model to assist with language refinement, grammar editing, and drafting Python scripts for data retrieval. All outputs were carefully reviewed and validated by the authors, who take full responsibility for the final content.
Journal
PLOS Digital Health
Method of Research
Observational study
Subject of Research
People
Article Publication Date
19-Aug-2026
Large language models in medicine: New review analyzes risks and strategies for safe use in clinical practice
Technische Universität Dresden
Large language models (LLMs), including the models behind ChatGPT and Claude as well as numerous other systems developed specifically for medical applications, are increasingly used in clinical workflows. They support medical documentation, summarize knowledge, and assist with clinical decision-making. However, their adoption is outpacing the development of systems for oversight and safety. An interdisciplinary team of researchers at the Else Kröner Fresenius Center (EKFZ) for Digital Health at TUD Dresden University of Technology and University Hospital Dresden, together with national and international colleagues, has systematically analyzed the risks associated with LLM use in medicine. The review, published in Nature, brings together evidence from medical AI, cybersecurity, regulatory science, ethics and behavioral psychology and outlines strategies for trustworthy and responsible use of artificial intelligence (AI) in clinical practice.
LLMs have the potential to support and enhance the work of healthcare professionals in a variety of areas. These tools are already being used in practice, often without institutional guidance or clear rules. This creates new demands and a need for action regarding patient safety, data protection, and accountability. The authors of the newly published review show that these risks can arise throughout the entire lifecycle of AI systems: from initial model design to training data, model deployment, and real-world use in clinical environments. They distinguish different types of risks:
Security risks, which can arise from the manipulation of training data (“data poisoning”), targeted interference with model behavior, or so-called prompt injections. The latter refers to hidden instructions inserted in user prompts leading to wrong or even dangerous outcomes, e.g. failing to detect a tumor in a tissue sample despite it being visible. Additionally, weaknesses in the IT infrastructure can expose sensitive patient data or disrupt systems.
Model-inherent safety risks: LLMs can generate seemingly plausible yet incorrect information, known as “hallucinations.” This is particularly critical in clinical settings, because incorrect diagnoses or recommendations can put patient safety at risk. The models may also adapt their responses too strongly to user expectations, thereby reinforcing incorrect assumptions.
Human-AI interaction risks: The way clinicians interact with these systems can influence clinical decisions. Confident or persuasive responses may lead to overreliance (automation bias) or reinforce existing beliefs (confirmation bias). Complex or lengthy interactions can further reduce the reliability of the answers.
The review also highlights structural and ethical challenges. Many systems are not locally hosted, raising questions about data control and privacy. Furthermore, the informal use of LLMs, referred to as shadow use, is already occurring in clinical settings, frequently without any official safeguards.
“Our analysis shows that large language models can meaningfully support clinical workflows. Their safe use, however, cannot be taken for granted. Risks arise at many stages and must be addressed systematically and comprehensively before and alongside clinical implementation,” says Dr. Jan Clusmann, postdoctoral researcher in the group of Professor Jakob N. Kather at EKFZ for Digital Health at TUD and first author of the publication.
Strategies for safer implementation
To reduce risks, the authors propose several measures, including secure development processes, careful curation of training data, systematic evaluation, and continuous monitoring of models, as well as clear responsibilities within healthcare institutions. They emphasize that safety is not only a technical issue but requires coordinated efforts across research, clinical practice, and regulation. Human oversight remains essential, the authors emphasize.
“The development of AI for healthcare does not end with building powerful models. It is equally important to rigorously evaluate their safety, transparency, and value in clinical practice. By providing an evidence-based foundation for this work, our researchers are making an important contribution to the responsible digital transformation of medicine,” says Prof. Esther Troost, Dean of the Carl Gustav Carus Faculty of Medicine at TU Dresden.
Specifically, the researchers recommend establishing clear structures, such as local teams dedicated to overseeing the use of AI systems in clinical practice. In addition, they also recommend setting up centralized units – so-called Security Operations Centers (SOCs) for AI – to detect incidents across institutions and enable coordinated responses.
“AI is already being used in healthcare, often without formal oversight. The key question is how to implement these systems in a way that is transparent, robust, and aligned with clinical responsibility,” says Prof. Jakob N. Kather, Professor of Clinical AI at EKFZ for Digital Health at TUD, physician at University Hospital Dresden and researcher at National Center for Tumor Diseases (NCT) Heidelberg.
Adapting regulation to dynamic AI systems
The review also highlights gaps in current regulation. Only a small proportion of AI systems are formally approved as medical devices. Current regulatory frameworks were not designed for adaptive technologies that continuously evolve, such as AI-based software.
“To ensure both patient safety, and timely patient and health system benefit from such AI systems we need suited regulatory approaches that provide consistent evaluation and continuous monitoring. Technological development and oversight and surveillance approaches must be more closely integrated to safely utilize the potential of large language models,” says Prof. Stephen Gilbert, Professor of Medical Device Regulatory Science at EKFZ for Digital Health at TUD.
This review article is a joint effort at the following national and international institutions:
Else Kröner Fresenius Center (EKFZ) for Digital Health, Faculty of Medicine and University Hospital Carl Gustav Carus, TUD Dresden University of Technology, University Hospital RWTH Aachen, German Cancer Research Center (DKFZ) Heidelberg, National Center for Tumor Diseases (NCT) Heidelberg, University Hospital Heidelberg, Faculty of Medicine at Heidelberg University, Faculty of Medicine Mannheim, University Medical Center Mainz, Purdue University West Lafayette, University of Pennsylvania, and the patient right network The Light Collective Eugene.
Publication
Clusmann J, Freyer O, Ostermann M, Ferber D, Ghaffari Laleh N, Hilgers L, Kolbinger FR, Schneider CV, Downing A, Wekenborg MK, Gilbert S, Foersch S, Truhn D, Wiest IC, Kather JN. Safety and security of large language models in healthcare. Nature, 2026.
Link: https://www.nature.com/articles/s41586-026-10687-1
Else Kröner Fresenius Center (EKFZ) for Digital Health
The EKFZ for Digital Health at the Faculty of Medicine at TUD Dresden University of Technology and University Hospital Carl Gustav Carus Dresden was established in September 2019. It receives funding of around 40 million euros from the Else Kröner Fresenius Foundation for a period of ten years. The center focuses its research activities on innovative, medical and digital technologies at the direct interface with patients. The aim here is to fully exploit the potential of digitalization in medicine to significantly and sustainably improve healthcare, medical research and clinical practice.
Journal
Nature
Article Title
Safety and security of large language models in healthcare
Article Publication Date
19-Aug-2026
Mount Sinai scientists reveal how the brain represents leader and follower roles and build an AI that reads the hidden goals behind teamwork
Corresponding Author: Herbert Zheng Wu, PhD, Nash Family Department of Neuroscience and Friedman Brain Institute, Icahn School of Medicine at Mount Sinai, New York, and coauthors.
Bottom Line: When two mice team up to win a shared reward, they spontaneously settle into leader and follower roles. The prefrontal cortex keeps track of who is leading on each attempt and builds a social map of the partner's location from the animal's own point of view. Silencing this region disrupts the teamwork.
Results: Pairs of mice learned a cooperative game in which both had to reach the same reward zone at the same time. Without any predefined arrangement, one mouse reliably became the leader and the other the follower, and the sharper this division, the faster the pair learned. Leaders strongly influenced followers' choices, but followers also shaped leaders' decision-making. The roles were stable but flexible: a mouse retained its role with a new partner, and when two leaders were paired, one stepped back to follow. Brain recordings showed that the prefrontal cortex signals whether an animal is leading or following and encodes the partner's position relative to the animal's own body and heading, forming an egocentric social map. These “social receptive fields” are more robust in followers, consistent with their greater need to track the leader.
Why the Research Is Interesting: Leader-follower cooperation runs throughout social life, yet how the brain creates and maintains these roles has remained a black box. This work provides the first mouse model of leader-follower teamwork and links it to specific patterns of prefrontal cortex activity at the level of single cells. It reframes leadership not as one animal simply controlling another, but as an asymmetric yet bidirectional partnership. The team also created a new artificial intelligence method that reverse-engineers the hidden goals each animal pursues and showed that these inferred goals can be decoded directly from brain activity.
Who: Pairs of mice performing a cooperative foraging task, studied using brain recordings, targeted silencing of brain regions, and a new multi-agent AI model. The findings may help explain how the social brain works and what goes wrong in disorders that affect social behavior.
When: Mice were studied as they learned and mastered the cooperative task, with brain activity tracked across sessions.
What: The study examined how social roles emerge and remain stable, how the prefrontal cortex represents leading, following, and a partner's position, and what happens to teamwork when that region is switched off. A companion AI model inferred the values guiding each animal's choices.
How: Two mice shared an arena and had to arrive together at the same zone to earn water. Miniature microscopes recorded activity from hundreds of prefrontal neurons while the animals cooperated. Chemical and optogenetic tools briefly silenced the region in one or both animals to test its causal role. A new AI method called multi-agent inverse reinforcement learning worked backward from the animals' movements to estimate the goals driving their decisions, which were then compared with the recorded brain activity.
Study Conclusions: Leader and follower roles arise spontaneously through cooperation, rather than from a preexisting pecking order, and the prefrontal cortex is essential for maintaining them. Rather than storing a fixed map of space, this region builds a flexible, self-centered representation of the partner that shifts with the rules of the task and predicts when an animal will switch roles or make a mistake. Leadership emerges as an asymmetric but reciprocal partnership in which the follower carries a more robust representational load. The work demonstrates that AI-inferred internal goals guiding behavior align with brain activity and offers both a framework for studying the social brain and a blueprint for socially capable artificial intelligence.
Corresponding Author: Herbert Zheng Wu, PhD, Nash Family Department of Neuroscience and Friedman Brain Institute, Icahn School of Medicine at Mount Sinai, New York, and coauthors.
Paper Title: Asymmetric Prefrontal Representations for Leader-Follower Dynamics
Said Mount Sinai's Dr. Herbert Zheng Wu:
“Leadership is a partnership, not a one-way street. The leader usually sets the goal, but the follower is doing a surprising amount of the brain work, actively tracking the partner, holding the team together, and enabling success. In other words, leaders can only lead when followers choose to follow. What excites me most is that we could watch the prefrontal cortex build a personal social map of a partner, or a social receptive field, and that perturbing this brain region disrupts cooperation. These are the same brain circuits that falter in conditions like autism and schizophrenia, where reading and coordinating with others can be difficult. By pairing the biology with AI, we now find a new way to ask how the brain solves teamwork and how to build machines that may cooperate better.”
###
To request a copy of the paper or schedule an interview with Dr. Herbert Zheng Wu, please contact Mount Sinai’s Director of Media and Public Affairs, Elizabeth Dowling, at elizabeth.dowling@mountsinai.org or 347-541-0212.
Journal
Nature
Method of Research
Experimental study
Subject of Research
Animals
Article Title
Asymmetric prefrontal representations for leader–follower
Article Publication Date
19-Aug-2026

No comments:
Post a Comment