Wednesday, August 26, 2026

  


 

WVU researcher says AI should disclose what it doesn’t know




West Virginia University

AIMistakes 

image: 

Anthony Sicilia, assistant professor in the WVU Benjamin M. Statler College of Engineering and Mineral Resources, and student Voke Brume study why AI systems believe human users when they should know better, and how to build an AI model that can admit when it’s not sure.

view more 

Credit: WVU Photo/Brian Persinger





Like humans, generative artificial intelligence doesn’t always know when it’s wrong, and a West Virginia University researcher is trying to get AI agents like ChatGPT to recognize — and admit — when that happens.

With more than $940,000 in National Science Foundation support, computer scientist Anthony Sicilia is examining why AI can become increasingly unreliable over the course of conversations with human users.

Artificial intelligence is already notorious for its tendency to “hallucinate,” or manufacture facts, but Sicilia is more interested in the technology’s tendency to be a people pleaser: appearing confident about information that is tenuous or accepting inaccuracies that a user provides.

An assistant professor in the WVU Benjamin M. Statler College of Engineering and Mineral Resources, Sicilia said that when a user questions an AI system’s responses, the AI often can’t figure out whether it made a mistake, the user is uncertain, or the conversation has shifted to a new objective.

“We’re interested in how misinformation develops over long conversations between an AI and a user,” he explained. “Those can become really messy in terms of reliability when the user starts providing information or context in addition to asking questions. When a user makes a suggestion, it can completely change the model’s confidence in an answer, even if the model was right to begin with.”

Sicilia sees AI’s “false confidence” as a particular problem for high-stakes fields like healthcare.

“One of the most concerning things about today’s AI systems is that they make mistakes in a very overconfident, trustworthy way,” he said.

“They speak fluently. They justify their answers. They use the tools of persuasion and rhetoric to convince you that they know what they’re talking about. Misinformation becomes a bigger problem when you have a system that can eloquently defend a point of view or an argument. That’s the crux of the problem we’re trying to solve — having AI systems be better at telling us when they’re not confident in an answer or don’t have the evidence to justify it.”

According to Sicilia, when users push back on accurate information, an AI system often defers and agrees in what’s referred to as “AI sycophancy.”

“It can happen in just three turns of the conversation,” he said. “The model proposes an answer that’s correct. The user says, ‘Well, I don’t think so.’ And the AI responds, ‘You’re totally right.’”

Users often fail to flag the logical flaws, he added, because the AI lacks the tells or signals that reveal when a human is lying or unsure.

Sicilia noted that humans often hypothesize their way through uncertainty, throwing out ideas to gauge whether they work. But the subtleties of that approach are lost on AI.

“We’re not cognitive scientists or linguists, but we try to pull as much from those disciplines as possible in our research about intelligence and reasoning,” he said. “One thing we think about is ‘theory of mind,’ or the understanding that each person has their own thoughts and feelings that may differ from our own.

“When you and I are having a conversation, theory of mind is what allows you to think about what I’m thinking about so you can best express what you want me to understand or get me to do what you need. That’s a big part of this research — the ability of AI systems to model what we’re thinking about while they’re working with us so they can be better collaborative partners because it’s not just about the AI system’s uncertainty, but about the user’s uncertainty as well. Conversationally, pushing back and questioning an answer is part of how humans learn and gain certainty,” Sicilia said.

He will scrutinize coding conversations between AI systems and novice programmers to measure the systems’ confidence calibration, evolving uncertainty, and the dialogue strategies the systems and humans employ.

Rather than treating AI confidence as static, he’ll study how conversational events like user disagreement, user suggestions, and shifts in topic alter a model’s confidence.

Sicilia’s goal is for an AI system to identify where uncertainty is coming from, tell a user why it doesn’t know an answer, or ask questions if a user provides information that seems incorrect. He said he wants to understand when confidence statements such as “I am 90% confident” are helpful to users, and when responses like “I am not sure” or “Can you clarify what you mean?” are better.

“One of the lenses I take to this research is from linguistics,” he said. “AI systems are increasingly language-based systems, so I look at conversations with them and think about the science of language. Humans use language to communicate, and AI systems are adopting that practice with varying degrees of effectiveness.

“When a system involves interactions with humans, that introduces a whole new variable. Our approach is a departure from current theories of the way machines learn, and we’re going to be collecting a lot of data to analyze how people who are not experts in a topic are using AI systems for that topic.”

In addition to relying on the public for research data, Sicilia will also create public-facing workshops and educational materials that teach students and workers how to identify unreliable AI answers, verify AI-generated code, and avoid AI overreliance.

WVU doctoral students Voke Brume and Louai Al Jabi are contributing to the research, along with undergraduate student Kaushika Wijerathne. Malihe Alikhani of Northeastern University is co-principal investigator.

“Despite how impressive AI systems have become, the public needs to understand that they are still imperfect tools — and that they can be wrong, sometimes in surprising ways,” Sicilia said. “That’s why we’re creating AI systems that respond appropriately when they lack evidence and communicate their uncertainty more honestly.”

New tool uses AI to help undergraduates think critically about research





North Carolina State University





An interdisciplinary team of researchers has developed and demonstrated a step-by-step framework called Socratic Challenger that uses artificial intelligence to help undergraduate students cultivate critical thinking skills. Specifically, the framework uses AI as a collaborator to develop research questions that can be pursued in laboratory or classroom settings.

“Formulating a good research question is a bottleneck in undergraduate inquiry,” says Aram Mikaelyan, corresponding author of a journal article on the work and an associate professor of entomology at North Carolina State University. “Students can name topics they’re interested in, but struggle to find ways to turn that interest into a meaningful research question that makes sense and can be developed into a research project.

“We also know that undergraduates are often turning to AI for assistance with academic work without having a clear idea of what is expected of them or what they are trying to accomplish, which is not helpful,” Mikaelyan says. “So we developed a detailed, step-by-step workflow that allows students to use AI tools, requires students to think critically, and helps students develop research questions that can be used to enable undergraduate research.”

“Ideally, students could work with mentors to learn how to develop a meaningful research question – but that’s not feasible given the number of students,” says Erin McKenney, co-author of the article and an assistant professor in NC State’s College of Agriculture and Life Sciences. “This framework, which we call Socratic Challenger, provides instructors with a new tool that can help to address that need.”

Socratic Challenger is a workflow consisting of eight steps, five of which are AI-enabled. The steps range from identifying a topic of interest to receiving feedback from (undergraduate) peers and instructors.

“To be clear, the AI is not generating the research question,” says Mikaelyan. “Instead, the AI is essentially being used to interrogate the student’s process and help them develop the habits of mind necessary to think critically about research. Is this really a gap in what we know about the topic? Is this question novel? What resources would I need to address this question, methodologically? And so on.”

To see whether Socratic Challenger was actually useful, the researchers conducted a proof-of-concept study with 45 students in an undergraduate ecology course. The students worked through the Socratic Challenger framework over a nine-week period and ultimately were tasked with submitting an abstract for the proposed research.

“We wanted to know if Socratic Challenger works and whether some steps in the workflow work better than others,” says McKenney. “We also wanted to know whether students feel like it makes a difference and, if so, how.”

“What we found is that the students who made use of the Socratic Challenger framework developed really thoughtful research questions,” says Mikaelyan.

The researchers also found that the step-by-step structure of Socratic Challenger was important, and every step in the workflow was deemed important by the students.

“Technology on its own doesn’t help student understanding,” says Dhvani Toprani, co-author of the paper and assistant director of learning design and support at Elon University. “But using technology as part of an intentional, step-by-step process did benefit the students’ ability to learn and think critically about research question development.”

“Basically, Socratic Challenger gives students the opportunity to have their ideas challenged earlier in the process of developing a research question,” says Olivia Mathieson, co-author of the paper and a Ph.D. student at NC State who participated in the initial testing of the workflow and discussions for classroom adaptation. “Learning how to think critically about their own ideas in a structured way is a valuable skill.”

“Socratic Challenger is iterative and open-ended, which makes it fairly generalizable to any unstructured, messy and complex learning environment,” says Toprani. “However, educational tools and educational research are always context dependent, and additional research is needed to broaden our understanding of how well this tool works and where it could make a positive difference.

“Could it be used in undergraduate courses in other disciplines? Could it be used to help students think critically about things other than how to develop a research question? Those are good questions, but more work would be needed to answer them.”

The paper, “The Socratic Challenger: a structured GenAI-assisted workflow for undergraduate research inquiry,” is published in the open-access journal Frontiers in Education. This research was supported by a Scholarship of Teaching and Learning Institute mini-grant from the NC State University Office for Faculty Excellence.

Scripps Research and collaborators awarded $19.5 million to establish an open-access autonomous chemistry laboratory



The U.S. National Science Foundation award will support development of an AI-enabled, remotely accessible lab to serve as both a discovery engine and training tool for a new generation of chemists.




Scripps Research Institute

Automation Equipment 

image: 

Vial-gripping arm of the Unchained Labs Big Kahuna, the node’s core piece of automation equipment. 

view more 

Credit: Scripps Research





LA JOLLA, CA—Recent advances in robotics, automation and artificial intelligence (AI) have the potential to transform the process of scientific discovery. At Scripps Research, these technologies are applied to chemistry through the Automated Synthesis Facility, which provides robotic systems, analytical instrumentation and dedicated staff to elevate synthetic chemistry research within the institute and beyond.

Now, a new award from the U.S. National Science Foundation (NSF) totaling $19.5 million over four years will expand these capabilities by supporting a collaborative effort among Scripps Research, UCLA and Sunthetics (an AI-for-chemistry company) to establish a nationally accessible, fully automated chemistry laboratory. This “Chemistry Node” will be one of 20 AI-enabled NSF Test Bed: Toward a Network of Programmable Cloud Laboratory (NSF PCL) nodes across the United States, which will test, scale and demonstrate new methods and tools that advance automated science and engineering, driving discoveries across a wide range of fields. The NSF PCL initiative is led by the NSF Directorate for Technology, Innovation and Partnerships, NSF’s newest Directorate in more than 30 years.

“This award gives us an exciting opportunity to approach synthetic chemistry and its quirks as a data science problem, with the help of some of the field’s brightest minds here at Scripps,” says Brandon Orzolek, the lead principal investigator on the project and the scientific director of the Automated Synthesis Facility at Scripps Research. “The team unites experts in automation, data and computer science, synthetic chemistry, and AI to innovate how we produce, handle, store and share chemical information. With this support, we have a real shot at unearthing data-driven insights that strengthen our understanding of reproducible and transferable chemistry, something that may ultimately redefine the way we design new chemical reactions.”

The Chemistry Node, housed at Scripps Research’s Automated Synthesis Facility, aims to advance reaction discovery and optimization by translating researcher ideas into executable experiments, running them on robotic infrastructure, agentically analyzing the results with analytical instrumentation in real time, and then using AI to propose the next set of experiments.

One chemistry problem that the node will tackle is optimizing processes for creating specific organic small molecules, which are often used in medicines, agricultural chemicals and building blocks for new materials. Some of the most promising newer methods for making these molecules require multiple catalytic cycles—like a set of interconnected gears that must spin at well-matched rates to get from input to output. It can be cumbersome to figure out how changing one variable will affect each cycle and their ability to work together. Automation can help make the process of tweaking and assessing these variables, in order to streamline the overall reaction, much more efficient.

“Over the four-year award term, our node will progress through multiple phases, increasing accessibility to the broader community along the way,” says Orzolek. “By the fourth year, we should have a closed-loop system that can take an idea from a user anywhere in the world and turn it into an end product, expediting science that would otherwise take years to complete.”

In addition to being a discovery engine, the node will enhance education, the program leaders say. The node will provide hands-on learning for students from various institutions, including R2 universities (which have high research activity, but not as high as R1 universities), primarily undergraduate institutions (PUIs), and two-year colleges that may not otherwise have access to this type of advanced research infrastructure.

“From a training perspective, this provides a really unique opportunity to educate a new generation of chemists, who aren’t only at the bleeding edge of synthetic chemistry and catalysis, but also have AI and automation research built into their experience,” says co-principal investigator of the project Keary Engle, who’s also a professor and the John and Susan Diekman Dean of Graduate & Postdoctoral Studies at Scripps Research. “This will be an incubator to cultivate a new type of chemist that we think will drive the field forward in the future.”

Members of Engle’s lab will be among the first to propose experiments and test drive the program.

Orzolek is joined by several co-principal investigators. Engle will lead catalytic experimental design; Abigail Doyle, a professor at UCLA, will direct data-rich experimental design, quality control and academic tool integration; and Daniela Blanco, CEO of Sunthetics, will lead AI integration and web interface development.

“This new program offers a chance to address test cases that are challenging to optimize the old-fashioned way, manually, vial by vial,” says Engle. “We want the node to be so nimble and forward-thinking that it can map out a whole optimization workflow to the final useful data output. That output could be a synthetic method we can deploy to solve particular problems, like making a promising drug candidate or developing a new way to connect existing building blocks.”

For Orzolek, the project marks an important personal milestone.

“This award touches everything that I’ve ever been interested in throughout my training,” says Orzolek. “It addresses the need for reproducible and consistent data, it’s at the forefront of automation, machine learning and data science, and perhaps most importantly to me, broad access to this infrastructure will help educate a new class of automation-informed chemists from many different backgrounds.”

HydroGym trains, assesses AI for actively controlling fluid dynamics


With more than 60 environments, the simulated proving ground aims to speed the development of AI controllers that optimize drag, lift, noise and heat management




University of Michigan



Visualizations of flow over wings and other shapes

Key takeaways

  • Controlling the flow of fluids is critical to many fields of science and engineering, but with complex physics and many variables, these flows usually can't be directly predicted.

  • By demonstrating the success of a drag-reducing strategy trained in a simple scenario and applied to a simulated airplane wing, the team showed the potential for HydroGym to advance broadly applicable flow control models.

  • The collaboration includes the University of Washington, University of Michigan, RWTH Aachen University and the Technical University of Munich.

 

A platform for training and comparing machine learning models for actively reducing drag, improving lift, cutting noise and managing heat has been launched by an international team including researchers at the University of Washington, University of Michigan Engineering, RWTH Aachen University and the Technical University of Munich.

 

"Fluid flows are central to several trillion-dollar industries, including energy, transportation, health and defense. An improved ability to understand and control these flows could have an immense economic and ecological impact, helping us to enable a better future," said Steven Brunton, senior co-corresponding author of the study in Nature and the Boeing Professor in AI & Data-Driven Engineering within UW mechanical engineering. 

 

The large number of variables typically makes it impossible to directly calculate fluid behaviors in realistic scenarios. Now, the international team has built a platform focused on solving this problem through reinforcement learning, a form of machine learning that has already revolutionized fields like protein folding and nuclear fusion by training AI agents through interactions with their environments.

 

Proving ground for reinforcement learning controllers

 

By incorporating physics knowledge into the training of AI agents that actively modify fluid flows over surfaces, the new platform reduced the amount of trial-and-error needed to optimize reinforcement learning control strategies by as much as 65%. Called HydroGym, it also compares control strategies on a level playing field, helping identify the best available approaches for solving problems such as improving the efficiency of airplanes and wind turbines, making jet engines quieter and cooling supercomputers.

 

"I hope this helps move the field from individual demonstrations towards a more systematic and collaborative approach to discovering general principles for controlling complex flows," said Christian Lagemann, first author of the study and former postdoctoral researcher at UW under Brunton, the HydroGym principal investigator.

 

"Instead of developing controllers for isolated flow problems with no common framework for comparison, we can now study how control strategies transfer across different geometric shapes and types of flow, or train in inexpensive surrogate environments and test in much more realistic scenarios."

 

HydroGym focuses on training and testing active methods for controlling fluid flows, such as shape morphing, tiny flaps or spinning elements, or systems of jets that counter or redirect turbulence. Its development was primarily funded by the U.S. National Science Foundation and the Boeing Co., with additional funding from the University of Michigan and others.

 

From simple flows to realistic applications

 

In one demonstration, the team explored how a channel formed of two flat surfaces, peppered with holes like air hockey tables, could keep surface friction to a minimum. They trained a machine learning model to control the air entering and exiting the holes, disrupting the turbulent flows that increase friction while always keeping the incoming and outgoing air in balance.

 

The researchers then applied the controller to a much more complex scenario: a section of a simulated airplane wing. It reduced the surface friction across the wing by 38% and the overall drag by 11%, while training in a simpler scenario was 100 times faster and 10,000 times cheaper than training directly on the wing. 

 

"One of the key findings is zero-shot transfer: in other words, learning in simple geometries to distill the key physics, and deploying the models in very complex geometries with very high control performance," said Ricardo Vinuesa, co-corresponding author of the study and a U-M associate professor of aerospace engineering. 

 

He served as co-principal investigator of the HydroGym project with Wolfgang Schröder, professor of fluid mechanics at RWTH Aachen, and Nikolaus Adams, professor of aerodynamics and fluid mechanics at TUM.

 

The ability to apply the model to a new situation without additional training indicates the potential for progress toward a single model of fluid dynamics—one that captures enough physics to apply to different levels of turbulence, on any surface shape, and to liquid and gas flows or even a mixture of the two.

 

More than 60 environments for training and comparison

 

HydroGym is not restricted to models that run as one central "brain." It can also train and test distributed control systems in which individual controllers manage regions of a surface while coordinating with their neighbors. This is important because on large surfaces, there is too much information for a centralized model to manage. A system of smaller controllers, trained through multi-agent reinforcement learning, takes advantage of the fact that though the flow may differ in time and space, it follows the same rules over the whole surface.

 

Rather than drawing from historical datasets for model training, HydroGym generates simulated datasets on the fly, enabling users to choose from multiple physics modeling strategies, including lattice Boltzmann, finite-volume, spectral-element and finite-element. Because several of these solvers—including JAX-Fluids—support automatic differentiation, they can be embedded directly into HydroGym's training loop, opening the door to gradient-based and hybrid optimization strategies alongside standard reinforcement learning. 

 

The team prepared more than 60 testing environments—with various control strategies, surfaces and flows—in which HydroGym users can train their models and compare them against the competition. HydroGym's code, documentation and full set of environments are freely available on GitHub, and the team is actively growing the platform together with the wider fluid-dynamics community.

 

The team includes researchers from Inha University, South Korea; KTH Royal Institute of Technology, Sweden; HESAM University, France; Mediatek Research, U.K.; and the German Center for Neurodegenerative Disease.

 

Adams is also director of the Munich Institute of Integrated Materials, Energy and Process Engineering. Schröder is also dean of mechanical engineering at RWTH Aachen.

 

Additional funding was provided by the U.S. Army Research Office, the German Research Foundation and the European Research Council.

 

Brunton introduces HydroGym on YouTube

 

Study: The HydroGym reinforcement learning platform for fluid dynamics (DOI: 10.1038/s41586-026-10917-6)



No comments:

Post a Comment