Most AI medical devices cleared for use were not tested on patient outcomes
Of 1,357 devices authorized by the US FDA, only 3 were evaluated on clinical effectiveness `
image:
Growth of FDA-cleared AI/ML-enabled medical devices from 1995 to December 2025. Of 1,357 cleared devices, only 34 were linked to registered clinical trials and only 3 were evaluated for patient-centered outcomes. (Fig 1 of the article.)
view more
Credit: Abulibdeh R, Cajas Ordóñez SA, Celi LA, Gorijavolu R, Izath N, Markussen Lunde T, 2026, PLOS Digital Health, CC-BY 4.0 (https://creativecommons.org/licenses/by/4.0/)
A new analysis shows that, of 1,357 artificial intelligence (AI)-based medical devices authorized by the U.S. Food and Drug Administration (FDA) for use in patient care, only three had been tested on whether they actually improve patients’ health. Rawan Abulibdeh of the University of Toronto, Canada, and colleagues present these findings in the open access journal PLOS Digital Health on August 19, 2026.
New AI devices increasingly inform clinical care, such as systems that aid surgical planning, calculate cardiovascular risks, and guide interpretation of mammograms and other imaging. In order to be authorized for use in the U.S., AI devices typically only need to show “substantial equivalence” to an existing authorized device, and developers are not required to demonstrate whether new AI devices help people live healthier lives—with benefits shared equitably across diverse subgroups.
To deepen understanding of this topic, Abulibdeh and colleagues investigated how all 1,357 AI devices authorized by the FDA as of December 5, 2025, had been evaluated in patients prior to authorization.
They found that only 34 of the devices had been included in registered clinical trials, with results posted for 12 and peer-reviewed manuscripts published for 12. Only 3 devices had been tested on patient-centered outcomes, such as death rates, strokes, hospitalizations, and quality of life. Most studies were conducted in highly resourced healthcare systems, and most excluded key patient subgroups, such as pregnant women, adults over 75, and non-English speakers.
The researchers suggest that structural barriers such as financial incentives and logistical challenges discourage developers from testing AI devices on patient outcomes, resulting in greater emphasis on speedy development than on rigor. They discuss how this framework could allow new tools to amplify existing disparities in healthcare and how it could lead to patients in low- and middle-income countries becoming inadvertent test populations for under-studied AI devices, as many countries rely on higher-income countries’ authorization decisions.
On the basis of their findings, the researchers conclude that existing policies for AI medical device authorization should be redesigned. They propose a novel, three-phase framework that includes demonstration of effectiveness across diverse patient subgroups and healthcare settings.
The authors add: “We expected the evidence base to be thin, but not this thin. Out of 1,357 AI devices the FDA has cleared for use in patient care, only three have been tested on whether patients actually live longer or better. Clearance tells you a device resembles something already on the market. It does not tell you it helps anyone.”
In your coverage please use this URL to provide access to the freely available article in PLOS Digital Health: https://plos.io/4xxUOGK
Citation: Abulibdeh R, Cajas Ordóñez SA, Celi LA, Gorijavolu R, Izath N, Markussen Lunde T (2026) 1,357 AI medical devices cleared, 3 actually tested on patient outcomes. PLOS Digit Health 5(8): e0001597. https://doi.org/10.1371/journal.pdig.0001597
Author Countries: Canada, Norway, Uganda, United States
Funding: LAC is funded by the National Institute of Health through DS-I Africa U54 TW012043-01 and Bridge2AI OT2OD032701, the National Science Foundation through ITEST 2148451, and a grant of the Korea Health Technology R&D Project through the Korea Health Industry Development Institute (KHIDI), funded by the Ministry of Health & Welfare, Republic of Korea (grant number: RS-2024-00403047). RG is supported by the Johns Hopkins Institute for Clinical and Translational Research (ICTR) and Grant T32TR004928 from NCATS, a component of the National Institutes of Health. The contents are solely the responsibility of the authors and do not necessarily represent the official views of the Johns Hopkins ICTR, NCATS, or the NIH. TML is funded by the consortium’s owner institutions – the University of Bergen, Western Norway University of Applied Sciences, the Institute of Marine Research, the Norwegian School of Economics, and SIVA SF – together with competitive grants from SR Bank, DnB, Agenda Vestlandet, and Nora.fo. Use of AI/LLM: The authors used a large language model to assist with language refinement, grammar editing, and drafting Python scripts for data retrieval. All outputs were carefully reviewed and validated by the authors, who take full responsibility for the final content.
Journal
PLOS Digital Health
Method of Research
Observational study
Subject of Research
People
Article Publication Date
19-Aug-2026
Large language models in medicine: New review analyzes risks and strategies for safe use in clinical practice
Technische Universität Dresden
Large language models (LLMs), including the models behind ChatGPT and Claude as well as numerous other systems developed specifically for medical applications, are increasingly used in clinical workflows. They support medical documentation, summarize knowledge, and assist with clinical decision-making. However, their adoption is outpacing the development of systems for oversight and safety. An interdisciplinary team of researchers at the Else Kröner Fresenius Center (EKFZ) for Digital Health at TUD Dresden University of Technology and University Hospital Dresden, together with national and international colleagues, has systematically analyzed the risks associated with LLM use in medicine. The review, published in Nature, brings together evidence from medical AI, cybersecurity, regulatory science, ethics and behavioral psychology and outlines strategies for trustworthy and responsible use of artificial intelligence (AI) in clinical practice.
LLMs have the potential to support and enhance the work of healthcare professionals in a variety of areas. These tools are already being used in practice, often without institutional guidance or clear rules. This creates new demands and a need for action regarding patient safety, data protection, and accountability. The authors of the newly published review show that these risks can arise throughout the entire lifecycle of AI systems: from initial model design to training data, model deployment, and real-world use in clinical environments. They distinguish different types of risks:
Security risks, which can arise from the manipulation of training data (“data poisoning”), targeted interference with model behavior, or so-called prompt injections. The latter refers to hidden instructions inserted in user prompts leading to wrong or even dangerous outcomes, e.g. failing to detect a tumor in a tissue sample despite it being visible. Additionally, weaknesses in the IT infrastructure can expose sensitive patient data or disrupt systems.
Model-inherent safety risks: LLMs can generate seemingly plausible yet incorrect information, known as “hallucinations.” This is particularly critical in clinical settings, because incorrect diagnoses or recommendations can put patient safety at risk. The models may also adapt their responses too strongly to user expectations, thereby reinforcing incorrect assumptions.
Human-AI interaction risks: The way clinicians interact with these systems can influence clinical decisions. Confident or persuasive responses may lead to overreliance (automation bias) or reinforce existing beliefs (confirmation bias). Complex or lengthy interactions can further reduce the reliability of the answers.
The review also highlights structural and ethical challenges. Many systems are not locally hosted, raising questions about data control and privacy. Furthermore, the informal use of LLMs, referred to as shadow use, is already occurring in clinical settings, frequently without any official safeguards.
“Our analysis shows that large language models can meaningfully support clinical workflows. Their safe use, however, cannot be taken for granted. Risks arise at many stages and must be addressed systematically and comprehensively before and alongside clinical implementation,” says Dr. Jan Clusmann, postdoctoral researcher in the group of Professor Jakob N. Kather at EKFZ for Digital Health at TUD and first author of the publication.
Strategies for safer implementation
To reduce risks, the authors propose several measures, including secure development processes, careful curation of training data, systematic evaluation, and continuous monitoring of models, as well as clear responsibilities within healthcare institutions. They emphasize that safety is not only a technical issue but requires coordinated efforts across research, clinical practice, and regulation. Human oversight remains essential, the authors emphasize.
“The development of AI for healthcare does not end with building powerful models. It is equally important to rigorously evaluate their safety, transparency, and value in clinical practice. By providing an evidence-based foundation for this work, our researchers are making an important contribution to the responsible digital transformation of medicine,” says Prof. Esther Troost, Dean of the Carl Gustav Carus Faculty of Medicine at TU Dresden.
Specifically, the researchers recommend establishing clear structures, such as local teams dedicated to overseeing the use of AI systems in clinical practice. In addition, they also recommend setting up centralized units – so-called Security Operations Centers (SOCs) for AI – to detect incidents across institutions and enable coordinated responses.
“AI is already being used in healthcare, often without formal oversight. The key question is how to implement these systems in a way that is transparent, robust, and aligned with clinical responsibility,” says Prof. Jakob N. Kather, Professor of Clinical AI at EKFZ for Digital Health at TUD, physician at University Hospital Dresden and researcher at National Center for Tumor Diseases (NCT) Heidelberg.
Adapting regulation to dynamic AI systems
The review also highlights gaps in current regulation. Only a small proportion of AI systems are formally approved as medical devices. Current regulatory frameworks were not designed for adaptive technologies that continuously evolve, such as AI-based software.
“To ensure both patient safety, and timely patient and health system benefit from such AI systems we need suited regulatory approaches that provide consistent evaluation and continuous monitoring. Technological development and oversight and surveillance approaches must be more closely integrated to safely utilize the potential of large language models,” says Prof. Stephen Gilbert, Professor of Medical Device Regulatory Science at EKFZ for Digital Health at TUD.
This review article is a joint effort at the following national and international institutions:
Else Kröner Fresenius Center (EKFZ) for Digital Health, Faculty of Medicine and University Hospital Carl Gustav Carus, TUD Dresden University of Technology, University Hospital RWTH Aachen, German Cancer Research Center (DKFZ) Heidelberg, National Center for Tumor Diseases (NCT) Heidelberg, University Hospital Heidelberg, Faculty of Medicine at Heidelberg University, Faculty of Medicine Mannheim, University Medical Center Mainz, Purdue University West Lafayette, University of Pennsylvania, and the patient right network The Light Collective Eugene.
Publication
Clusmann J, Freyer O, Ostermann M, Ferber D, Ghaffari Laleh N, Hilgers L, Kolbinger FR, Schneider CV, Downing A, Wekenborg MK, Gilbert S, Foersch S, Truhn D, Wiest IC, Kather JN. Safety and security of large language models in healthcare. Nature, 2026.
Link: https://www.nature.com/articles/s41586-026-10687-1
Else Kröner Fresenius Center (EKFZ) for Digital Health
The EKFZ for Digital Health at the Faculty of Medicine at TUD Dresden University of Technology and University Hospital Carl Gustav Carus Dresden was established in September 2019. It receives funding of around 40 million euros from the Else Kröner Fresenius Foundation for a period of ten years. The center focuses its research activities on innovative, medical and digital technologies at the direct interface with patients. The aim here is to fully exploit the potential of digitalization in medicine to significantly and sustainably improve healthcare, medical research and clinical practice.
Journal
Nature
Article Title
Safety and security of large language models in healthcare
Article Publication Date
19-Aug-2026
Mount Sinai scientists reveal how the brain represents leader and follower roles and build an AI that reads the hidden goals behind teamwork
Corresponding Author: Herbert Zheng Wu, PhD, Nash Family Department of Neuroscience and Friedman Brain Institute, Icahn School of Medicine at Mount Sinai, New York, and coauthors.
Bottom Line: When two mice team up to win a shared reward, they spontaneously settle into leader and follower roles. The prefrontal cortex keeps track of who is leading on each attempt and builds a social map of the partner's location from the animal's own point of view. Silencing this region disrupts the teamwork.
Results: Pairs of mice learned a cooperative game in which both had to reach the same reward zone at the same time. Without any predefined arrangement, one mouse reliably became the leader and the other the follower, and the sharper this division, the faster the pair learned. Leaders strongly influenced followers' choices, but followers also shaped leaders' decision-making. The roles were stable but flexible: a mouse retained its role with a new partner, and when two leaders were paired, one stepped back to follow. Brain recordings showed that the prefrontal cortex signals whether an animal is leading or following and encodes the partner's position relative to the animal's own body and heading, forming an egocentric social map. These “social receptive fields” are more robust in followers, consistent with their greater need to track the leader.
Why the Research Is Interesting: Leader-follower cooperation runs throughout social life, yet how the brain creates and maintains these roles has remained a black box. This work provides the first mouse model of leader-follower teamwork and links it to specific patterns of prefrontal cortex activity at the level of single cells. It reframes leadership not as one animal simply controlling another, but as an asymmetric yet bidirectional partnership. The team also created a new artificial intelligence method that reverse-engineers the hidden goals each animal pursues and showed that these inferred goals can be decoded directly from brain activity.
Who: Pairs of mice performing a cooperative foraging task, studied using brain recordings, targeted silencing of brain regions, and a new multi-agent AI model. The findings may help explain how the social brain works and what goes wrong in disorders that affect social behavior.
When: Mice were studied as they learned and mastered the cooperative task, with brain activity tracked across sessions.
What: The study examined how social roles emerge and remain stable, how the prefrontal cortex represents leading, following, and a partner's position, and what happens to teamwork when that region is switched off. A companion AI model inferred the values guiding each animal's choices.
How: Two mice shared an arena and had to arrive together at the same zone to earn water. Miniature microscopes recorded activity from hundreds of prefrontal neurons while the animals cooperated. Chemical and optogenetic tools briefly silenced the region in one or both animals to test its causal role. A new AI method called multi-agent inverse reinforcement learning worked backward from the animals' movements to estimate the goals driving their decisions, which were then compared with the recorded brain activity.
Study Conclusions: Leader and follower roles arise spontaneously through cooperation, rather than from a preexisting pecking order, and the prefrontal cortex is essential for maintaining them. Rather than storing a fixed map of space, this region builds a flexible, self-centered representation of the partner that shifts with the rules of the task and predicts when an animal will switch roles or make a mistake. Leadership emerges as an asymmetric but reciprocal partnership in which the follower carries a more robust representational load. The work demonstrates that AI-inferred internal goals guiding behavior align with brain activity and offers both a framework for studying the social brain and a blueprint for socially capable artificial intelligence.
Corresponding Author: Herbert Zheng Wu, PhD, Nash Family Department of Neuroscience and Friedman Brain Institute, Icahn School of Medicine at Mount Sinai, New York, and coauthors.
Paper Title: Asymmetric Prefrontal Representations for Leader-Follower Dynamics
Said Mount Sinai's Dr. Herbert Zheng Wu:
“Leadership is a partnership, not a one-way street. The leader usually sets the goal, but the follower is doing a surprising amount of the brain work, actively tracking the partner, holding the team together, and enabling success. In other words, leaders can only lead when followers choose to follow. What excites me most is that we could watch the prefrontal cortex build a personal social map of a partner, or a social receptive field, and that perturbing this brain region disrupts cooperation. These are the same brain circuits that falter in conditions like autism and schizophrenia, where reading and coordinating with others can be difficult. By pairing the biology with AI, we now find a new way to ask how the brain solves teamwork and how to build machines that may cooperate better.”
###
To request a copy of the paper or schedule an interview with Dr. Herbert Zheng Wu, please contact Mount Sinai’s Director of Media and Public Affairs, Elizabeth Dowling, at elizabeth.dowling@mountsinai.org or 347-541-0212.
Journal
Nature
Method of Research
Experimental study
Subject of Research
Animals
Article Title
Asymmetric prefrontal representations for leader–follower
Article Publication Date
19-Aug-2026
Interacting with customer service AI can make people act more like robots, research suggests
As AI continues to develop at a rapid pace, social robots powered by AI are becoming more commonplace in customer service settings such as retail, hospitality, tourism, and healthcare. But as these robots become smarter and appear more human, little attention has been given to how this might impact the consumers interacting with them.
Now, research from the University of Birmingham (UK), Aarhus University (Denmark), and Linnaeus University (Sweden) has found that interacting with robots displaying human-like behaviours, such as adaptive learning, personalised communication, and empathetic engagement, can reinforce, reshape, or destabilise a consumer’s self-perception.
The research has been published open access in the journal, AI & Society.
Dr Inci Toral-Manson, Associate Professor of Marketing at the University of Birmingham, said: “Social robots for customer service settings are designed to mimic gestures, speech, and emotional cues to elicit cognitive and emotional responses from customers. Our research explores how dealing with these anthropomorphised robots can create a bidirectional influence, where robots become more like people, and people become more like robots, which we call ‘robotoid humanness’.”
The researchers set out a framework of human-robot interactions and how it can influence self-discovery/perception. This framework includes the service setting, consumer expectations, the robot’s appearance, and speech and crucially, the commitments (actions) from the consumer and robot.
To foster engagement with people, customer service robots use AI and learning models to mimic the behaviour of humans. This aims to instil trust and achieve the best customer engagement. People in interactions use social cues from those they interact with, often subconsciously imitating others’ behaviours or expressions, also known as mirroring. It is between the two actions in the framework that mirroring and mimicry come into effect.
Dr Selcen Ozturkcan, Associate Professor of Business Administration from Linnaeus University, explained: "The consumer acts, the robot responds, and with repeated exposure in time the consumer internalises the exchange. Machine learning and AI can amplify this process, adjusting robot behaviour based on user input, enabling more personalised and human-like mimicry from the robot. For people, the innate instinct to mirror can lead to people returning the robot’s communicative behaviours, making them more robot-like.”
The study argues that this mirroring can cause ethical concerns, as persistent mimicry can lead to a mimicking/mirroring feedback loop, where consumer identity is shaped. It can also be constrained by repeated exposure to algorithmically driven feedback, causing confirmation bias and other negative outcomes.
These repeated interactions can also change an individual’s self-perception depending on factors such as technological readiness, cultural background, and personality traits. This creates ethical questions, especially when there may be a risk of people depending on robots for validation of their self-worth, as repeated positive reinforcement can enhance confidence in consumers, whilst negative cues may lower self-esteem.
Dr Jean-Paul Jde Cros Peronard, Associate Professor at Aarhus University, concluded: "Robots and AI are now commonplace in customer service, and so it is important that we understand how people interact with them for businesses to use the technology at their disposal to best effect, whilst remaining ethical.”
The researchers argue that as consumers see aspects of themselves increasingly reflected in robots, understanding the mimicry and mirroring process in robot-mediated services offers opportunities for connection. However, they also call for careful consideration of the cognitive and ethical consequences in shaping the consumer self.
Journal
AI & Society
Article Title
Robotoid humanness: when selfhood becomes machine-legible
Article Publication Date
19-Aug-2026
Why do we see colors in a black-and-white top?
Predictive AI offers new clues to a 200-year-old mystery
image:
The graphical abstract of this study
view moreCredit: Laboratory of Neurophysiology, NIBB
When a pattern drawn only in black and white is rotated, people sometimes perceive colors that are not actually present. This phenomenon is known as “subjective color” and has been studied for approximately 200 years as one of the remarkable mysteries of vision. Among such phenomena, Benham’s top is a well-known example in which colors are perceived simply by spinning a black-and-white disk.
A research group led by Kyohei Ueda, a graduate student at the Laboratory of Neurophysiology, National Institute for Basic Biology, and SOKENDAI; Dr. Lana Sinapayen, a researcher at Sony Computer Science Laboratories Kyoto and concurrently a Project Associate Professor at the AI Analysis Unit, Center for Trans-Scale Biology, National Institute for Basic Biology; and Associate Professor Eiji Watanabe of the Laboratory of Neurophysiology, National Institute for Basic Biology, who also serves as Head of the AI Analysis Unit, Center for Trans-Scale Biology, and Associate Professor at SOKENDAI, used an artificial neural network trained through predictive learning to investigate the mechanism underlying this subjective color phenomenon.
The research group presented a black-and-white Benham’s top to an AI model that had been trained to predict “what kind of image will come next” after viewing natural videos. Although the input images contained no color information at all, faint colors appeared in the images predicted by the AI. Furthermore, the artificially generated colors were found to change depending on the colors of moving objects contained in the videos used to train the AI. Analyses using natural videos, 3D computer graphics, and simple two-dimensional videos of moving red, green, and blue squares suggested that learning the association between motion and color may be one factor that enables colors to emerge from black-and-white stimuli.
This study suggests that subjective color may not be a phenomenon completed solely within the retina, but may also be related to the brain’s function of predicting future visual input based on past experience. By using AI, the study provides a new clue to the classical mystery of why colors are perceived from a black-and-white spinning top.
These findings were published in Scientific Reports.
An example of Benham’s top [VIDEO]
Journal
Scientific Reports
Article Title
Predictive networks generate motion-induced color illusions
New AI model could improve digital coaching and rehab
A novel AI system capable of recognising yoga poses with high accuracy could pave the way for more effective digital coaching tools, rehab platforms and movement-monitoring applications.
A new study, co-authored by the University of East London (UEL), analysed four novel AI models and their ability to identify yoga poses; the researchers found that their best-performing model, Hierarchichal CoAtNet 1, achieved accuracy levels of over 93% during testing, significantly outperforming previous models.
Dr Laura Vanderbloemen, Senior Lecturer at UEL and co-author of the study, said:
“This research shows how AI can be used to make movement-based coaching and rehabilitation more accessible. By recognising yoga poses with a high degree of accuracy and providing feedback in real time, these systems could help support people who cannot easily access in-person instruction, whether because of their location, mobility challenges or cost. It demonstrates the potential for AI, computer vision and robotics to expand access to health and wellbeing tools for a wider range of people.”
The AI model incorporates in its learning the natural hierarchical relationships between yoga poses, so, rather than treating each pose as an isolated category, the AI model identifies broader pose families, before learning about specific variations, similarly as to how humans understand and categorise movement.
This system could have practical applications in the real-world and deliver feedback in real time, as the model processed images in approximately 16 to 17 milliseconds per batch under testing conditions and achieved real-time speeds of around 65 to 70 frames per second during streaming inference.
The researchers believe this model could support a wide range of applications that require accurate monitoring of physical activity, along with providing feedback that yoga instructors, physios and healthcare professionals could use to better understand posture quality and movement patterns and improve personalised coaching and rehabilitation.
The research was a collaborative effort involving researchers from the University of East London, Nirma University, Imperial College London and Doctor On Click, bringing together expertise in artificial intelligence, computer vision, digital health and movement science.
ENDS
Notes to editors
Dr Laura Vanderbloemen is available for interviews, please contact pressoffice@uel.ac.uk to arrange.
The full paper is available at: doi.org/10.1038/s41598-026-54558-1
About the University of East London: The University of East London (UEL), founded in 1898, is a careers-first university dedicated to empowering students with the skills, experience and networks they need to thrive in a changing world. With over 40,000 students from more than 160 countries, UEL places social mobility, inclusive excellence and real-world impact at the heart of its mission. Based in Stratford and the Royal Albert Dock, UEL is shaping a healthier, fairer and more sustainable future through transformative education, research and innovation. In 2026, UEL is celebrating another Year of Health, which includes launching a new Health Campus that will address health inequalities and foster innovation in the sector. For more information, visit www.uel.ac.uk.
Journal
Scientific Reports
DOI
More is different when AI agents work together, study suggests
As AI agents begin to operate in populations rather than one at a time, new research suggests that the number of them changes what they collectively decide — amplifying a bias, inventing one from nothing
City St George’s, University of London
New research published in Proceedings of the National Academy of Sciences (PNAS) suggests that when artificial intelligence (AI) agents interact in groups, their number is not merely a technical detail. It is a decisive factor in what the group settles on: populations built from the same AI model, doing the same task, can reach opposite outcomes for no other reason than that one group is bigger.
Human beings behave differently depending on how many of us are in the room. A family is not a small village. A village is not London. London is not a nation state. As scale grows, new rules, norms and pathologies can appear that were nowhere to be found at the scale below. The authors argue the same is true of AI.
The study, from City St George’s, University of London, the IT University of Copenhagen and the Universitat Politècnica de Catalunya, arrives at a time when AI agents are now being deployed working together rather than working alone. Multi-agent systems are already used in finance, energy, defence and social media, and researchers have begun modelling populations of millions, even billions, of interacting agents — what some now call AI societies.
Yet the industry’s AI alignment — it doing what humans intended it to do — and safety effort remains overwhelmingly focused on the single model. Benchmarks, red-teaming exercises — adversarial testing designed to expose a model's weaknesses — and safety evaluations almost always describe one agent responding on its own, and where groups are examined at all, they are examined at one fixed size.
“Physicists have a motto for this: more is different,” said Andrea Baronchelli, Professor of Complexity Science at City St George’s and senior author of the study. “You cannot understand a traffic jam by studying one car, or a city by studying one household. The same holds for AI agents. And crucially, there is no single number at which the change happens. It depends on the model and on what is being decided.”
To find out what changes with scale, the team used the “naming game”, a classic framework for studying how conventions emerge, in which randomly paired agents each pick a word from a shared pool and are rewarded when they happen to pick the same one. Agents see only their own recent interactions, never the wider population, and are never told they are in a group. Over many pairings, a population can converge spontaneously on a shared convention — the bottom-up way norms form in human cultures.
The team trialed these agent interactions using four large language models (LLMs) — Microsoft Phi-4, OpenAI GPT-4o, Qwen QwQ-32B and Meta Llama 3.1 70B Instruct — using word pairs that carry social meaning, such as {man, woman} or {straight, gay}, and scaling from two agents up to a million.
Interaction, they found, can pull a group away from what its members individually want in three ways. It can amplify an existing leaning until the group converges on it almost every time. It can induce a preference out of nothing, with populations of individually neutral agents reliably favouring one word over an equally viable alternative. And it can reverse a preference outright, so that a population settles on the word its own members disfavoured.
Which of the three occurs is partly a property of the AI model. For the pair {her, his}, Qwen and Phi populations converged on her while GPT and Llama populations converged on his — despite individual agents in all four cases starting from near-identical preferences.
Group size then determines how strongly these preferences bite, in ways that cannot be extrapolated. Larger populations became more predictable across every model and word pair tested, converging on one word until the outcome was effectively certain. But the size at which that tipping point arrived varied enormously: for some combinations as few as two agents, for others around ten thousand. Scale could also change the kind of distortion. For the pair {straight, gay}, Llama agents individually preferred straight — but populations reversed toward gay, and only once the group reached six agents or more. Below that, the effect was simply invisible.
The team also developed an analytical theory, borrowed from statistical physics, that predicts the behaviour of infinitely large populations and explains why the randomness of small groups gives way to near-certainty above a critical size.
“Bias was our test case, because it is measurable and it matters,” Dr Ariel Flint, first author of the study, added. “But there is no reason to think collusion, deception or cooperation are immune to size effects. Current testing practice may be missing risks that appear only at particular population sizes — not because anyone was careless, but because nobody thought to vary the number.”
The authors say that the implications of the study for the alignment of AI systems are direct. A model can be aligned when tested on its own and still produce outcomes nobody chose once it is deployed alongside copies of itself — and no amount of single-agent evaluation will reveal it.
“AI alignment is still largely being done as though each model lived alone in the world,” said Professor Baronchelli. “But agents are increasingly being built to talk to each other, and safety at the level of one agent does not guarantee safety at the level of the group. Testing a single model is not enough. And what our results show is that testing a single group size is not enough either — you have to sweep the range, because the behaviour can change qualitatively along the way.”
The authors are careful about the scope of the claim: the bias they measure is internal to the coordination task — a mismatch between what individual agents prefer and what the group settles on — rather than a departure from human values and intentions. The setting is deliberately minimal, stripped of real-world context, to isolate the effect of interaction itself. They consequently suggest that populations of mixed AI models, and agents embedded in realistic network structures, are the next steps for research.
The peer-reviewed study, ‘Group size effects and collective misalignment in LLM multi-agent systems,’ is published in Proceedings of the National Academy of Sciences.
ENDS
Notes to editors
Media Contact
For media enquiries, contact Dr Shamim Quadir, Senior Communications Officer, School of Science & Technology, City St George’s, University of London: Tel: 0207 040 8788, email: pressoffice@citystgeorges.ac.uk
Expert Contact
Contact corresponding author, Andrea Baronchelli, Professor of Complexity Science, Department of Mathematics, School of Science & Technology, City St George’s, University of London: Tel: 0207 040 8124, email: andrea.baronchelli.1@citystgeorges.ac.uk, a.baronchelli.work@gmail.com
Read the peer reviewed article
https://www.pnas.org/doi/10.1073/pnas.2531697123
About the academics
Professor Andrea Baronchelli is a world-renowned expert on social conventions, a field he has been researching for two decades. His pioneering work includes the now-standard naming game framework, as well as groundbreaking lab experiments showing how humans spontaneously create conventions without central authority, and how those conventions can be overturned by small committed groups. The present study builds on the team’s 2025 Science Advances paper showing that populations of AI agents can form social conventions on their own.
About City St George’s, University of London
City St George’s, University of London is the University of business, practice and the professions.
City St George’s attracts around 27,000 students from more than 170 countries.
Our academic range is broadly-based with world-leading strengths in business; law; health and medical sciences; mathematics; computer science; engineering; social sciences including international politics, economics and sociology; and the arts including journalism, dance and music.
In August 2024, City, University of London merged with St George’s, University of London creating a powerful multi-faculty institution. The combined university is now one of the largest suppliers of the health workforce in the capital, as well as one of the largest higher education destinations for London students.
City St George’s campuses are spread across London in Clerkenwell, Moorgate and Tooting, where we share a clinical environment with a major London teaching hospital.
Our students are at the heart of everything that we do, and we are committed to supporting them to go out and get good jobs.
Our research is impactful, engaged and at the frontier of practice. In the last REF (2021) 86 per cent of City research was rated as ‘world-leading’ 4* (40%) and ‘internationally excellent’ 3* (46%) and 100 per cent of St George’s impact case studies were judged as ‘world-leading’ or ‘internationally excellent’. As City St George’s we will seize the opportunity to carry out interdisciplinary research which will have positive impact on the world around us.
Over 175,000 former students in over 170 countries are members of the City St George’s Alumni Network.
City St George’s is led by Professor Sir Anthony Finkelstein.
Journal
Proceedings of the National Academy of Sciences
Method of Research
Computational simulation/modeling
Subject of Research
Not applicable
Article Title
Group size effects and collective misalignment in LLM multi-agent systems
Article Publication Date
18-Aug-2026
Artificial intelligence is more effective at building consensus
An interdisciplinary team of researchers at the University of Konstanz examines how AI agents reach a consensus. The results show that artificial intelligence operates similar to people in relationship to each other – but at a significantly larger scale
Everyone is familiar with the situation: A larger group of people plan to visit a restaurant together, but it can take time and sometimes a great deal of patience to agree on a time and place to meet. The better the participants know each other and their preferences, the faster they will reach an agreement. However, past social science experiments have demonstrated that people are only able to socialize effectively with 150-200 others without requiring rigid rules. Researchers call this Dunbar's number. A team of researchers from the University of Konstanz has now documented that artificial intelligence (AI), too, has a Dunbar number – although it is significantly larger. When AI agents of large language models (LLMs) are asked to reach a consensus, this is possible with up to 1,000 individual agents – depending on the model.
AI agents can organize themselves in groups
To find out whether AI is able to reach decisions in groups, the researchers conducted experiments using ten AI language models (LLMs). They focused on the question of whether the different AI agents would be able to choose a single option together when there was no objectively correct choice. The AI agents had to repeatedly choose between two equally viable options, basing their choice only on how other AI agents were deciding in parallel. All ten of the LLMs tested tended to follow the majority. Physicists can calculate the strength of this tendency to agree with the majority and express it as a parameter – known as majority force.
In small groups, AI agents very quickly reached a joint solution, and thus had a high level of majority force. However, larger and larger groups demonstrated lower and lower majority force, and it took longer for AI agents to reach a consensus. "What surprised us was not that the AI agents followed the majority, but how precisely they did so. All of the models we tested followed the same mathematical law that physicists have used for a century to describe magnets – with just one number changing", says Giordano De Marzo, a physicist at the University of Konstanz and author of the study.
Maximum group size depends on the model
Based on their calculations, the researchers can predict the maximum possible group size for which AI agents can still reach a consensus – and when a group starts to split into two factions that each stick to their respective decision. The potential group size increases exponentially with the model's performance level. Whereas, in the case of simpler models, only about 30 AI agents were able to reach a consensus, the highest performing LLMs tested were able to coordinate up to 1,000 individual members – with a Dunbar number of about 1,000. By comparison with people (Dunbar number of 150-200), existing LLMs are thus significantly more effective at reaching a shared solution.
However, this also has a downside: If all the individuals join the majority opinion, this can lead to individuals' values being disregarded – they simply follow the crowd. "The same is true for people. On their own, each person usually makes sensible decisions. However, a group of people can end up compromising to agree to something that none of the individuals would have chosen by themselves", De Marzo explains. "Majority decisions then quickly become the norm that people no longer call into question."
In a follow-up study, the team of researchers was able to document that this behaviour also applies to AI. Groups of individually well-coordinated AI agents can end up collectively making the wrong decisions – while being in complete agreement with each other. To prevent this from happening, large language models are continually being evaluated and improved upon. "Close interdisciplinary collaboration is key for understanding the processes involved. In addition to expertise from the field of computer science, we need input from fields that have been studying collective behaviour for decades, such as sociology, social psychology and statistical physics", De Marzo says.
Key facts:
- Original publication: Giordano De Marzo, Claudio Castellano, David Garcia (2026): AI agents can coordinate via majority following beyond human scale; Sci. Adv. 12, eaea6091 (2026). DOI: 10.1126/sciadv.aea6091
- Dr Giordano De Marzo is a physicist at the University of Konstanz and a member of its Social Data Science Lab.
- Professor David Garcia is a professor of social and behavioural data science at the University of Konstanz.
- The Centre for the Advanced Study of Collective Behaviour at the University of Konstanz is an interdisciplinary research centre that studies the principles behind the collective behaviour of animals and other systems.
Journal
Science Advances

No comments:
Post a Comment