When AI chatbots give bad advice, no one can see the damage
New paper explores how chatbot health guidance leaves no trace for clinicians or regulators, and calls for transparency safeguards
As more and more people turn to AI chatbots for quick, 24/7 medical advice, a new perspective paper co-authored by faculty at Binghamton University, State University of New York explores the risks this practice carries that are largely invisible to users – and healthcare systems.
The paper, written by researchers at Binghamton University, Stanford University, Texas A&M University, and Indiana University, and published in Nature Health, examines what happens when a chatbot gives someone bad health advice. The authors argue that the harm is hard to catch because the conversation stays on the AI company's platform, the person has no place to report the problem, and outside researchers cannot measure the damage.
“We’re getting medical advice from these chatbots, but nobody—literally no one—is looking at this. If there’s an error or an issue, there’s just no way for people to know,” said Kaicheng Yang, an assistant professor in the School of Computing at Binghamton University’s Thomas J. Watson College of Engineering and Applied Science and one of the paper’s authors.
The researchers gave the example of a 60-year-old man who asked ChatGPT how to cut chloride from his diet and swapped table salt with sodium bromide. The man ended up being hospitalized with bromide toxicity after experiencing hallucinations and paranoia. The issue this paper highlights is that because the man received this advice from a chatbot and was unable to retrieve the conversation, his doctors had no way of knowing exactly what the man was told, making it nearly impossible to piece together exactly what happened.
Because there is no system of transparency in place for AI companies as it relates to the visibility of health advice it generates for patients, this paper asserts that this built-in, systemic lack of visibility is a key aspect of AI-generated health information because it actively prevents oversight. “Harms are not only possible, but structurally hidden from the clinicians, researchers, and regulators who could otherwise detect and correct them,” the authors note.
This paper also emphasizes a key facet of chatbots generating answers for people with health questions – they can be wrong. Chatbots can generate erroneous answers confidently, oversimplify information, or draw from outdated information/false claims. However, the authors also note that those risks can trickle down to those who are not actively seeking answers. With the prevalence of AI summaries in search engines and the integration of AI in social media platforms like X and Meta, you do not have to be actively searching for an answer to be given an incorrect bit of information.
“Sometimes I’m not even looking for health information, but just by browsing my social media feeds, it’s there. It just shows up, and we believe that could have undesirable outcomes, especially if there’s medical misinformation or state actors trying to manipulate the online discussion,” Yang said.
Regardless of whether a person is actively looking for answers or comes across them incidentally, the authors of this paper argue that the harm rarely leaves traces that clinicians, regulators, or researchers can verify.
To combat this, the authors propose several methods to increase transparency and responsibility for AI companies. They note that AI companies should give users access to their own health conversations to provide them the opportunity to share this information with clinicians so they can better investigate the pathway that led to a harmful event. They also recommend that those companies develop a disclosure system to allow health guidance to be flagged, reported, and investigated.
They further recommend that social media companies take more responsibility by enhancing clear labeling of AI-generated health content, withholding that content until it is vetted by medical governing bodies, and similarly, that search engines should only utilize vetted information in summaries. Policymakers, the authors suggest, should also extend physician malpractice liability to AI chatbot companies.
“We do not think this is something we should rely on the companies to do, because their incentive is always to make more money,” Yang said. “Building such a system goes against that incentive. So we have to have some kind of third-party monitoring system, an independent evaluation.”
Journal
Nature Health
Method of Research
Commentary/editorial
Subject of Research
Not applicable
Article Title
The Invisible Risks of AI-Generated Health Information
Article Publication Date
9-Oct-2026
Introducing AlphaProof Nexus: An AI tool for formal mathematical proof discovery
Summary author: Walter Beckwith
A new large language model (LLM)-based artificial intelligence framework called AlphaProof Nexus can autonomously tackle select math problems, seeking proofs and using verification to ensure the resulting solutions are logically sound, researchers report. The system solved dozens of previously open problems across several mathematical fields, suggesting that artificial intelligence (AI) could become a useful tool for automated mathematical discovery and research. “[The authors report] that even unsuccessful proof attempts by AI could help them understand the problems better and make progress on solving them,” write Jeremy Avigad and Matthew Ballard in a related Perspective. “This underscores that an essential goal of developing AI for mathematics is to support mathematicians in the search for knowledge and understanding that lead to further advances.” Large language models (LLMs) have demonstrated growing ability to solve difficult mathematical problems, but their tendency to produce subtle logical errors or “hallucinations” makes them unreliable for research without extensive human expert review. One promising way to mitigate these issues is to have AI agents generate mathematical proofs in formal programming languages such as Lean, which automatically verifies each logical step and prevents errors from going unnoticed. While this approach has been successfully applied to competition mathematics and in the human-aided formalization of natural language arguments, its potential in solving open research-level mathematical problems remains unknown.
To address this gap, George Tsoukalas and colleagues developed AlphaProof Nexus, a framework that uses multiple AI agents to search for mathematical proofs with feedback from the Lean compiler. Tsoukalas et al. also developed a more advanced, full-featured agent that coordinated subagents through an evolutionary algorithm to use AlphaProof as a specialized proof tool. In tests, the system was able to solve nine of the 353 attempted Erdős problems, including two that had remained unsolved for more than 50 years. It was also able to solve 44 of 492 open On-Line Encyclopedia of Integer Sequences (OEIS) conjectures as well as several other research-level problems in fields such as algebraic geometry, optimization, quantum optics, and graph theory.
Journal
Science
Article Title
Advancing mathematics research with AI-driven formal proof search
Article Publication Date
8-Oct-2026

No comments:
Post a Comment