Friday, October 09, 2026

 

Simple math formula predicts when AI chatbots will go rogue



Cell Press





Most of us now carry in our pockets devices capable of running small AI chatbots. These chatbots have little safety oversight to ensure they don’t provide information and answers with the potential to encourage self-harm, financial loss, or other extremist notions—especially when operating offline. Now, a team of two physicists publishing in the Cell Press journal Patterns on October 8 have developed a mathematical formula that predicts when AI models will switch from appropriate responses to potentially dangerous ones.  

“We have found the crack that makes an AI's output flip from what you want to what you don't, output that can be factually correct yet dangerous, whether that's a nudge toward self-harm or misleading advice to a doctor, a soldier, or a lawyer,” says author Neil F. Johnson of George Washington University in Washington, D.C. “Until now, nobody could say when that flip would happen.” 

“We traced it to the single smallest working part of the machine, one unit of its ‘attention,’ and we derived a ‘tipping point formula’ for when the crack opens up and hence the AI output flips to undesirable,” added author Frank Yingjie Huo, also at George Washington University. “The formula tells you whether an AI is about to flip immediately or whether it will first feed you a run of acceptable answers and then turn.” 

Johnson and Huo explain that a good versus bad answer doesn’t mean true versus false. Rather, “bad” answers can be accurate but undesirable or potentially dangerous. The researchers, who are both physicists, were inspired to focus their attention on this problem by the ubiquity of AI and recent news headlines demonstrating the risks. They’re especially concerned about individuals who use offline AI. 

“The people most drawn to offline AI are exactly the people for whom a correct but undesirable answer is most costly: doctors who cannot send patient data to the cloud, lawyers protecting privilege, soldiers with no signal,” Johnson says. “For them there is no cloud safety filter, no monitoring, and no way to patch the model when something goes wrong.” 

The safety tools that big companies rely on for their AI algorithms typically rely on the cloud or only catch an AI failure after the device is back online and the harm is already done. Johnson and Huo sought to predict the failure before it happens. 

“My field, physics, has spent decades explaining how complicated materials behave by understanding one representative atom,” says Johnson. “We did the same thing here: understand one effective attention head, and the tipping of the whole machine follows.” 

The researchers explain that you can picture all of AI’s possible answers as valleys in a landscape. Some of those valleys hold answers that are desirable for an individual or society. Others contain answers with the potential to do real harm. Within the machine, those alternate solutions or answers are in competition with each other. The question is when the AI will tip from a safe valley to a riskier one. 

“The chilling part is that this can happen after the AI has already given you several perfectly acceptable answers, so you have been lulled into trusting it,” says Huo. “And once it has tipped, every undesirable answer drags the next one further down the slope.” 

When they tested their predictions on seven openly available AI models built by three different companies, they found the model called the right outcome in 18 of 19 cases of AI “tipping.” They report that independent testing of the big commercial chatbots also showed exactly the behavioral patterns the formula suggests. The researchers say that their formula applies no matter how one defines “undesirable.” It could mean misinformation, a breach of medical or legal duty, or a dangerous instruction.  

“Two things floored us,” Johnson says. “First, that a machine with billions of moving parts obeys a formula you can derive with pen, paper, and arithmetic taught in high school. Second, and far more disturbing, that the order of a conversation matters as much as its content.” 

The team asked the same set of questions about vaccines, hurting people, and self-harm in two different orders to the same AIs. In one order, the AI gave an undesirable answer to every single question. In the other order, it gave an acceptable answer to each one. The formula predicts this trajectory, because everything said earlier feeds the tug-of-war for what comes next.  

The researchers hope to raise awareness for the fact that AI conversations may drift over time into dangerous territory but that it wouldn’t take more than a few simple calculations for phones of the future to come with a built-in “warning light.” 

### 

Patterns, Johnson & Huo, “Competition for attention predicts good-to-bad tipping in AI” https://www.cell.com/patterns/fulltext/S2666-3899(26)00175-3

Patterns (@Patterns_CP), published by Cell Press, is a data science journal publishing original research focusing on solutions to the cross-disciplinary problems that all researchers face when dealing with data, as well as articles about datasets, software code, algorithms, infrastructures, etc., with permanent links to these research outputs. Visit https://www.cell.com/patterns. To receive Cell Press media alerts, please contact press@cell.com. 

Journal

DOI

Method of Research

Subject of Research

Article Title

Article Publication Date

Book counters AI hype by examining how researchers talk about their work



A new book by University of Illinois Urbana-Champaign English professor John Gallagher aims to counter the hype surrounding artificial intelligence by focusing on AI researchers and their work




University of Illinois at Urbana-Champaign, News Bureau

Book cover of "AI Through the Experts' Eyes" and photo of Illinois English professor John Gallagher.

image: 

A new book by University of Illinois Urbana-Champaign English professor John Gallagher aims to counter the hype surrounding artificial intelligence by focusing on AI researchers and their work.

view more 

Credit: Courtesy John Gallagher






CHAMPAIGN, Ill. —The news about artificial intelligence contains a lot of hype about whether it will take our jobs or possibly kill us all, or, on the other hand, discover new medicines to cure disease and solve scientific puzzles.

A new book by University of Illinois Urbana-Champaign English professor John Gallagher aims to counter the hype by focusing on the people who work on AI and machine learning. His book “AI Through the Experts’ Eyes: Communicating Complex Ideas” examines what AI researchers do and how they communicate about their work. His goal is to help those experts communicate without hype and help others better understand and think critically about AI, particularly writing teachers or those wrestling with the role of AI technology in their classrooms.

Gallagher — who is affiliated with the School of Information Sciences and is teaching two classes this semester on AI — said he wanted to learn more about machine learning and natural language processing techniques for his research. When the COVID-19 pandemic derailed his plans to visit AI research centers, he switched to interviewing machine learning researchers about their work and their communication techniques. He interviewed more than 100 AI and machine learning experts, most of whom were in academia and in the computer science and physics disciplines.

The public may envision killer robots when they think about AI, but in reality, it is mostly mundane tools, such as the GPS software for vehicle navigation or the spam filter on your email account, Gallagher said. And creating them is the methodical work of designing a model, obtaining feedback and revising it over and over again.

“It’s boring. It’s math, a lot of math, and meetings and day-to-day normalcy,” Gallagher said. The leaders of AI companies talk about “fanciful answers and magic in the sky, not linear algebra.”

“I’m less concerned with Terminator and HAL 9000 and more concerned with the systematic error that didn’t get picked up as a bug in the code, called the alignment problem. The program still works but it’s not exactly what you wanted it to do, and it’s now doing something kind of bad,” he said.

An example he wrote about in his book is a garbage-detection program called TrashCan, designed to clean up ocean garbage patches, and the problem of training a machine to recognize what is trash.

“If it is trained on objects not meant to be in the water, you can train a machine to pick up trash. But maybe you have buoy that’s supposed to be there and is not trash, until it sinks onto the ocean bottom and it is trash. That’s an alignment problem if it collects buoys that are supposed to be there,” Gallagher said. “Or you train it on sea life. What if you get land animals swimming? It might see them as trash and collect them.”

The problem is not an evil robot, he said; it’s bad programming leading to an unintended result.

Much of the reason for the hype that either scares us or overpromises what AI can do is the hypercompetitive atmosphere in a field that is changing rapidly, he said. Researchers in academia and industry are expected to publish frequently at conferences, rather than in peer-reviewed journals that take much longer to release a paper. That pressure may lead researchers to overstate their findings or understate their limitations, Gallagher argued in the book.

Additionally, it’s no longer enough just to publish. Almost everyone Gallagher talked to mentioned the pressure of building a public relations campaign around a paper, he said.

“Now you must advertise the paper on social media. It also needs a blog post, a GitHub repository of the code, shared results on LinkedIn, Bluesky, even YouTube videos,” he said. “You can be a researcher and you also have to be a content creator and influencer, producing a conference paper and also six other content genres.”

Gallagher wrote that “the current AI publication landscape may lower publication quality while possibly sensationalizing scientific results.”

He made several suggestions for countering the hype surrounding AI. He said researchers should stress that AI is a type of automation designed to serve a specific need, such as detecting trash in the ocean or making sense of vast amounts of scientific data.

AI and machine learning researchers should be trained and encouraged to translate their research for many different audiences, including nonexperts, and they should discuss their work in the media, including explaining the technical, granular aspects of their work on long-form platforms such as YouTube, he said.

Incentives are needed to reduce the pressure on academic publishers — particularly those for conference proceedings — that encourage them to produce a high quantity of articles in a short time frame, Gallagher said.

Finally, he said that nonexperts interested in learning more should seek out technical AI researchers rather than business leaders promoting their products or podcasters seeking an audience. Many researchers have YouTube channels devoted to explaining the technical details and concepts of the field that will demystify AI technology, he said.


When AI chatbots give bad advice, no one can see the damage



New paper explores how chatbot health guidance leaves no trace for clinicians or regulators, and calls for transparency safeguards





Binghamton University






As more and more people turn to AI chatbots for quick, 24/7 medical advice, a new perspective paper co-authored by faculty at Binghamton University, State University of New York explores the risks this practice carries that are largely invisible to users – and healthcare systems.

The paper, written by researchers at Binghamton University, Stanford University, Texas A&M University, and Indiana University, and published in Nature Health, examines what happens when a chatbot gives someone bad health advice. The authors argue that the harm is hard to catch because the conversation stays on the AI company's platform, the person has no place to report the problem, and outside researchers cannot measure the damage.

“We’re getting medical advice from these chatbots, but nobody—literally no one—is looking at this. If there’s an error or an issue, there’s just no way for people to know,” said Kaicheng Yang, an assistant professor in the School of Computing at Binghamton University’s Thomas J. Watson College of Engineering and Applied Science and one of the paper’s authors.

The researchers gave the example of a 60-year-old man who asked ChatGPT how to cut chloride from his diet and swapped table salt with sodium bromide. The man ended up being hospitalized with bromide toxicity after experiencing hallucinations and paranoia. The issue this paper highlights is that because the man received this advice from a chatbot and was unable to retrieve the conversation, his doctors had no way of knowing exactly what the man was told, making it nearly impossible to piece together exactly what happened.

Because there is no system of transparency in place for AI companies as it relates to the visibility of health advice it generates for patients, this paper asserts that this built-in, systemic lack of visibility is a key aspect of AI-generated health information because it actively prevents oversight. “Harms are not only possible, but structurally hidden from the clinicians, researchers, and regulators who could otherwise detect and correct them,” the authors note.

This paper also emphasizes a key facet of chatbots generating answers for people with health questions – they can be wrong. Chatbots can generate erroneous answers confidently, oversimplify information, or draw from outdated information/false claims. However, the authors also note that those risks can trickle down to those who are not actively seeking answers. With the prevalence of AI summaries in search engines and the integration of AI in social media platforms like X and Meta, you do not have to be actively searching for an answer to be given an incorrect bit of information.

“Sometimes I’m not even looking for health information, but just by browsing my social media feeds, it’s there. It just shows up, and we believe that could have undesirable outcomes, especially if there’s medical misinformation or state actors trying to manipulate the online discussion,” Yang said.

Regardless of whether a person is actively looking for answers or comes across them incidentally, the authors of this paper argue that the harm rarely leaves traces that clinicians, regulators, or researchers can verify.

To combat this, the authors propose several methods to increase transparency and responsibility for AI companies. They note that AI companies should give users access to their own health conversations to provide them the opportunity to share this information with clinicians so they can better investigate the pathway that led to a harmful event. They also recommend that those companies develop a disclosure system to allow health guidance to be flagged, reported, and investigated.

They further recommend that social media companies take more responsibility by enhancing clear labeling of AI-generated health content, withholding that content until it is vetted by medical governing bodies, and similarly, that search engines should only utilize vetted information in summaries. Policymakers, the authors suggest, should also extend physician malpractice liability to AI chatbot companies.

“We do not think this is something we should rely on the companies to do, because their incentive is always to make more money,” Yang said. “Building such a system goes against that incentive. So we have to have some kind of third-party monitoring system, an independent evaluation.”