It’s possible that I shall make an ass of myself. But in that case one can always get out of it with a little dialectic. I have, of course, so worded my proposition as to be right either way (K.Marx, Letter to F.Engels on the Indian Mutiny)
Extreme temperature events are becoming more frequent and intense as climate change accelerates. Yet predicting these abrupt, non-stationary phenomena remains a fundamental challenge for conventional time series models, which often fail to capture rapid transitions and anomalous patterns that deviate significantly from historical behaviors.
Now, researchers from the Hangzhou Institute for Advanced Study at the University of Chinese Academy of Sciences have developed Hankelformer, a novel deep learning architecture that dramatically improves the forecasting of extreme weather events. The findings are published in National Science Review.
Hankelformer introduces two key innovations. First, it employs a structured augmentation module that constructs Hankel matrices to capture local spatiotemporal dynamics without disrupting temporal coherence, generating delay-embedding-inspired views that are topologically equivalent to the original input sequences. Second, it uses a dual-stream contrastive learning framework in which both the original and augmented sequences are processed through shared-weight Transformer encoders, maximizing agreement between the two representations. This approach significantly enhances feature invariance and robustness against distribution shifts.
The team evaluated Hankelformer on nine benchmark datasets spanning energy, transportation, and extreme weather domains. On three custom datasets capturing real-world extreme events—the 2021 Texas winter storm (TexasFreeze), the 2021 Pacific Northwest heat dome (Heatwave), and the 2020 Antarctic Peninsula heat event (Antarctic Heat)—Hankelformer consistently achieved state-of-the-art performance. Compared to leading baselines, the model delivered up to 34% improvement in Mean Squared Error (MSE).
In validation tests on a 90-dimensional chaotic Lorenz system, Hankelformer demonstrated exceptional noise robustness, maintaining low prediction error even under strong Gaussian noise interference. Ablation studies confirmed that both the Hankel augmentation and the contrastive learning components are essential: using augmentation without the contrastive loss actually degraded performance, highlighting the critical role of contrastive learning in aligning heterogeneous representations and stabilizing optimization.
The success of Hankelformer provides a promising tool for reliable extreme temperature forecasting and underscores the value of constructing topologically equivalent sequences for spatiotemporal representation learning in handling real-world non-stationary time series. Beyond climate monitoring, the framework shows potential for applications in energy system management, traffic flow prediction, and other safety-critical domains where prediction failures can have severe consequences.
Predictions are made using the KIST-Ocean model, which has been trained on big data from atmospheric and oceanic observations and simulations spanning several decades. Three-dimensional ocean state data and atmospheric boundary conditions are input as initial conditions for the forecast, which are then used to predict the three-dimensional ocean state five days later. By repeating this process-where the predicted three-dimensional ocean state is re-inputted-up to 40 times, a global three-dimensional ocean forecast covering up to 200 days at five-day intervals is generated.
With the development of this year's Super El Niño, extreme weather events are occurring frequently around the world, such as the heatwave in Western Europe in June that saw temperatures exceed 40°C. As climate change increases the frequency and intensity of extreme weather events-including heatwaves, torrential rains, droughts, and typhoons-the importance of more accurate climate prediction technologies is growing. In particular, the ocean covers approximately 70% of the Earth's surface and serves as a key factor in determining seasonal and long-term climate variations-including El Ni?o and La Ni?a-by storing and circulating vast amounts of heat and carbon while continuously interacting with the atmosphere. However, existing ocean prediction models require massive supercomputing resources and lengthy computation times to solve complex physical equations, which has limited their ability to provide rapid forecasts and analyze various climate scenarios.
A research team led by Dr. Kang Daehyun at the Center for Climate and Carbon Cycle Research of the Korea Institute of Science and Technology (KIST; President Oh Sang-rok) announced that it has developed KIST-Ocean, an AI-based global ocean prediction model, to overcome these limitations. This achievement is the result of the KIST Climate and Environment Research Institute, which was established to predict and respond to climate change through science and technology. It is significant in that it has successfully integrated AI technology into the field of long-term ocean and climate forecasting-which requires vast amounts of observational data and physics-based analysis capabilities-thereby achieving both high accuracy and computational efficiency.
KIST-Ocean, developed by the research team, is an AI model that predicts future ocean conditions by learning from global 3D ocean data accumulated over the past several decades. Based on various physical parameters of the global ocean-such as sea surface temperature, ocean currents, and salinity-it forecasts the three-dimensional state of the ocean at five-day intervals and can generate changes down to a depth of 600 meters. In particular, by applying the latest AI technology, it can generate approximately 200 days' worth of global ocean forecast results in just a few seconds using a single GPU, significantly reducing computation time and costs compared to existing numerical models. This high computational efficiency is expected to be effectively utilized in research involving the iterative analysis of various climate scenarios or the performance of large-scale ensemble forecasts.
The research team conducted various experiments to verify how accurately KIST-Ocean replicates actual oceanic physical phenomena. In a virtual wind-generation experiment, ocean waves and upwelling and downwelling phenomena resulting from atmospheric changes were observed to align with existing ocean physics theories, demonstrating that AI can effectively reflect the complex physical interactions between the atmosphere and the ocean-going beyond simply learning from historical data. Furthermore, using the 2015 Super El Niño-a representative climate phenomenon-as a case study, the model successfully reproduced key developmental processes, such as the rise in sea surface temperatures in the equatorial Pacific and changes in the internal heat distribution of the ocean, thereby demonstrating the predictive accuracy and reliability of AI-based ocean models.
The research team expects KIST-Ocean to serve as a foundational technology that expands the scope of AI applications beyond existing short-term weather forecasts to include seasonal and annual climate forecasts. It is expected not only to lay the groundwork for the development of AI-based Earth system models that integrate the atmosphere, ocean, and land surface, but also to lower the barriers to entry for ocean and climate research and accelerate various studies on climate change adaptation, thanks to its speed and cost-effectiveness.
Dr. Kang Daehyun of KIST stated, "Through KIST-Ocean, we have demonstrated that AI predictions not only exhibit outstanding efficiency and accuracy but can also realistically reproduce the complex physical relationships between the atmosphere and the ocean," adding, "By actively utilizing these AI models, we will be able to significantly enhance our capacity to respond to future climate crises." Responding to climate disasters is a national public research area that requires a long-term accumulation of research and continuous investment. KIST plans to further refine this technology and develop it into a proactive forecasting tool that contributes to ensuring public safety and minimizing socio-economic damage in the era of the climate crisis.
###
KIST was established in 1966 as the first government-funded research institute in Korea. KIST now strives to solve national and social challenges and secure growth engines through leading and innovative research. For more information, please visit KIST’s website at https://kist.re.kr/eng/index.do
This study was conducted with support from the Ministry of Science and ICT (Minister Bae Kyung-hoon) as part of KIST's major projects, including the Project for the Advancement of Extreme-Performance Computing (NRF-2022M3K3A1094114). The findings of this study were published in the latest issue of the international academic journal *Science Advances* (IF 13.9, 8.2% in the JCR field).
Data-driven global ocean model resolving atmospherically forced ocean dynamics
An evaluation of El Niño development responses based on KIST-Ocean simulations using initial conditions from May 3, 2015, when the past Super El Niño began to develop.
(Left) When wind stress from 2015 was input, the simulation successfully reproduced the development of the past Super El Ni?o, with rising sea surface temperatures in the central to eastern Pacific.
(Right) When prescribing normal-year wind stress unrelated to El Niño development, no El Niño development occurred; instead, a La Niña response was observed, characterized by a decline in sea surface temperatures.
The results of the KIST-Ocean experiments are consistent with previous studies indicating that wind stress in the tropical Pacific plays a major role in the development of a super El Niño, demonstrating that KIST-Ocean possesses excellent physical fidelity.
Credit
Korea Institute of Science and Technology
(Left) Phase speed of oceanic Rossby waves observed in KIST-Ocean when artificial wind stress is applied to the equatorial Pacific. The x-axis denotes the latitude at which the wind stress was imposed. The y-axis shows the phase speed predicted by KIST-Ocean together with the corresponding theoretical values (shown in red), which vary with latitude. Each data point represents the results of multiple predictions performed at that latitude to ensure statistical significance.
(Right) Temperature response in the ocean at a depth of 105 m generated by KIST-Ocean following the injection of counterclockwise and clockwise wind stresses in the subtropical ocean. The model successfully predicted ocean temperature changes associated with upwelling and downwelling due to Ekman transport in a realistic manner.
Credit
Korea Institute of Science and Technology
What do people really think about generative AI?
Longitudinal study of Reddit posts since 2022 shows persistent tension between trust and distrust in generative AI
Is generative artificial intelligence technology a helpful tool or a big problem? Nearly four years after ChatGPT became a household name, most people are either using AI applications or generally aware of what the technology can do. But understanding how much they actually trust the work it produces, and the answers it provides, remains an important question. New research from Drexel University, based on a longitudinal analysis of hundreds of thousands of Reddit posts since 2022, sheds light on the general perception of AI and suggests that the public remains largely divided on how much it can be trusted.
Recently published in the journalTransactions of the Association for Computational Linguistics, the study examined how trust and distrust toward generative AI were expressed over time, the dimensions and reasons underlying these attitudes and how these patterns differed across groups represented in the posts.
Its primary finding is that trust, which was expressed in about 31% of posts, modestly outpaced distrust (26%) in the technology, while 41% of posts expressed neither and 1% expressed both.
Across the study period from 2022 to 2025, trust generally maintained this modest lead, although the balance fluctuated and distrust briefly surpassed trust during some periods. Despite a steady flow of new models and applications, the overall pattern suggests that attitudes expressed in these Reddit communities toward generative AI have remained divided, rather than moving steadily toward either trust or distrust. The findings stand in contrast to recent reports that suggest that while more people are using the technology, they do not trust it and trust has been declining.
“These findings give us an important starting point and help to establish a baseline understanding which can help inform responsible AI design, governance and literacy efforts,” said Shadi Rezapour, PhD, an assistant professor in the Nick Howley College of Engineering and Computing, who led the research. “It will be important to see how these attitudes around trust and distrust evolve as the technology becomes more widely used.”
In what is believed to be the first large-scale, longitudinal study of how attitudes of trust and distrust in AI have evolved over the last four years, Rezapour’s research group in the School of Computer and Information Sciences, joined by researchers from Drexel’s College of Arts and Sciences and LeBow College of Business, gathered and analyzed more than 230,000 posts from 39 AI-related subreddits between November 2022 and June 2025. They looked at posts expressing views on systems such as ChatGPT, LLaMA, Claude and other widely discussed GenAI tools.
What does it mean to ‘trust’ AI?
“We defined trust as a belief that Generative AI is reliable, competent or acts with integrity, leading people to have positive expectations about its performance or behavior,” said Aria Pessianzadeh, a doctoral candidate in the School of Computer and Information Sciences, and lead author of the study. “Distrust is more than simply the absence of trust. It reflects active skepticism or concern about the technology’s reliability, competence or ethical implications, which can lead to negative expectations or more cautious behavior.”
In addition to trust and distrust, and their dimensions, to better understand how these attitudes vary between different groups of people, the team broke down the type of commenter, or “trustor,” represented in each post into 10 categories using self-identifying information from their posts: generative AI users, software developers, researchers or academics, tech industry professionals, general public, media and journalists, business leaders or executives, AI ethicists or advocacy groups, artists or creatives and educators or knowledge workers.
Using computational models, the researchers analyzed the large volume of Reddit posts and categorized them based on the trust or distrust expressed and the reasons behind those views.
Views of trust tended to outweigh distrust among business leaders, academics, software developers and tech professionals. Distrust was more frequently expressed among posts categorized as representing the general public, AI ethicists and media and journalists. Generative AI users, which were the largest group, along with educators and knowledge workers, showed a relatively balanced level of trust and distrust, the study reported.
How do people form opinions about AI?
One theme that persisted across user groups was that personal experience with AI was the most frequently expressed reason underlying both trust and distrust. People also tended to express their views in terms of technological efficacy and reliability of the output, rather than ethical or normative aspects of the technology, such as judgments associated with a program’s relative transparency or integrity.
“What stood out was how practical people’s judgments of AI tended to be,” Pessianzadeh said. “They were primarily asking: ‘Does it work? Is it accurate? Can I rely on it?’ Competence, reliability and familiarity were central to expressions of trust, while unreliability and incompetence were major sources of distrust. Ethical and value-oriented concerns were certainly present, but they were much less prominent in everyday discussions. For many users, trust in AI is expressed primarily in relation to their direct experience with what these systems can actually do, suggesting that people are more focused on whether AI programs ‘work’ than whether the technology is ‘good.’”
Some of the clearest fluctuations in trust and distrust occurred around periods when the researchers observed the public release of new AI products, such as GPT-4 and LLaMA 2, as well as major announcements, such as OpenAI’s Dev Day. Around several major releases in 2023, the monthly share of posts expressing trust increased modestly. By contrast, distrust increased around OpenAI’s Dev Day in late 2023.
“Despite these fluctuations, trust maintained its modest lead over distrust. However, neither category dominated the discourse, underscoring a persistent tension in attitudes expressed in GenAI discourse,” the researchers reported in the study.
How to Move Forward
Acknowledging this co-existence of trust and distrust will be important, the researchers suggest, as regulators set up governance and design standards to ensure safe use of the technology.
“A persistent duality like this means that responsible governance of the technology should consider that users will include both those who are confident in AI as well as those who are skeptical,” said Rezapour. “Our findings also suggest that governance cannot focus only on whether people find AI useful. It should address questions of reliability, transparency, bias and accountability in ways that connect these concerns to people's everyday experiences with the technology.”
The researchers note that although this is a large-scale study, the fact that all of its data was drawn from Reddit posts could be a limiting factor. The relatively low prominence of ethical dimensions simply may be a reflection of how people communicate on social media – emphasizing immediate experiences of usefulness and performance, while broader ethical or normative concerns may be expressed less explicitly or require more contextual interpretation, according to the researchers.
They recommend that future research broaden the dataset to include languages other than English, additional social media platforms and populations beyond the online communities represented in the current study.
“These findings are a useful starting point, but trust in AI is not static and people’s attitudes change based on their experiences with these systems, new capabilities and major developments in the technology,” Rezapour said. “Therefore, researchers should continue to examine how attitudes about AI technology evolve in the coming years. With the rapidity of its adoption, it’s understandable that perceptions will be divided for some time. As these systems become increasingly embedded in everyday life, it will be important to understand whether trust and distrust continue to fluctuate, and what experiences, capabilities and broader developments drive those changes.”
Journal
Transactions of the Association for Computational Linguistics
Artificial intelligence (AI) is moving beyond screens into cars, drones, service robots and collaborative machines that can perceive, reason and act. A new survey maps the security and ethical risks created when vision-language models guide these embodied systems, where a mistaken description or manipulated command can become a physical action. The review connects failures across perception, planning, instruction following and human-robot interaction, covering hallucinations, synthetic forgeries, adversarial attacks, privacy leakage and unsafe execution. It also shows how the same multimodal capabilities can support defense through contextual checking, forgery detection, privacy protection and risk-aware reasoning, offering a roadmap toward embodied intelligence that is not only capable, but dependable in real-world settings.
Vision-language models (VLMs) connect images with text, while vision-language-action models (VLAs) extend that connection to robot plans and control signals. This makes natural-language instruction, scene understanding and flexible task execution possible, but it also creates a chain of dependency: flawed data can distort perception, weak visual-language alignment can produce hallucinations, and malicious inputs can redirect decisions. In a chatbot, such errors may generate misinformation; in an autonomous vehicle or industrial robot, they may lead to collisions, damaged equipment or failed missions. Existing safeguards are often benchmark-specific, fragmented across system layers or too computationally costly for real-time use. Because of these challenges, deeper research is needed into unified, adaptive safeguards for multimodal agents operating under uncertain physical conditions.
Published (DOI: 10.1007/s11633-025-1626-x) online on July 13, 2026, in Machine Intelligence Research, the review was conducted by researchers from the Institute of Automation, Chinese Academy of Sciences; University College London (UCL); Minzu University of China; and the China Academy of Electronics and Information Technology. The team examined how VLMs and VLAs are used in embodied intelligence (EI), organized the major security threats and defensive approaches, and connected technical safety with accountability, fairness, privacy, environmental sustainability and human oversight. The article appears in the journal’s special issue on the security and ethics of generative AI.
The survey first tracks VLM and VLA use across four linked functions: perception, planning, instruction following and human-robot interaction (HRI). It then shows how failures can cascade. Biased training data, weak visual encoders or poor cross-modal alignment can make a model describe objects that are not present. Forged traffic signs, altered labels, cloned voices or deceptive captions can misguide perception and planning. Tiny adversarial perturbations, hidden backdoor triggers and multimodal jailbreak prompts may bypass safety controls, while persistent sensing can expose identity, location, possessions and social behavior. The authors organize countermeasures into equally connected layers. These include hallucination filtering and vision-grounded alignment; cross-modal forgery detection, watermarking and provenance tracing; defenses against perturbations, backdoors and jailbreaks; differential privacy (DP), secure multi-party computation (SMPC) and homomorphic encryption (HE); and safeguards for navigation, communications and physical control. A further strand uses causal explanations, intent alignment and risk assessment so robots can interpret ambiguous instructions, anticipate hazards and correct actions. The review's central insight is that no single filter can secure an embodied agent: protection must follow the entire path from sensor input to model reasoning, system architecture and physical execution.
The authors said the central challenge is not simply making models more accurate, but ensuring that a system remains safe when its sensors, language inputs and operating conditions are imperfect. They said defenses should be combined rather than deployed as isolated patches, with transparent risk metrics, continuous monitoring and human oversight for critical decisions. A trustworthy robot must also explain what it is doing, recognize when it is uncertain and fall back safely instead of acting with false confidence. The authors added that technical progress must move alongside privacy protection, fairness, accountability and responsible governance.
For developers and regulators, the survey provides a practical checklist for evaluating embodied systems before large-scale deployment. Future platforms could combine interpretable reasoning, attack detection, privacy-preserving computation and dynamic safety controls under reproducible, open evaluation protocols. The authors call for designs that address four dimensions together: technical robustness, regulatory alignment, social equity and environmental sustainability. Such an approach could support safer autonomous transport, healthcare assistance, warehouse automation, industrial inspection and collaborative robotics, while making responsibility easier to trace when failures occur. The review also warns that strong laboratory results may not transfer cleanly to noisy, culturally diverse and resource-constrained environments. Progress will therefore depend on cross-disciplinary cooperation and testing that measures not only task success, but safe behavior under stress.
This work was partially supported by the National Natural Science Foundation of China (Nos. 62506362, 62302539 and U21B2045), the Strategic Priority Research Program of Chinese Academy of Sciences, China (No. XDA0480302), and Engineering and Physical Sciences Research Council (EPSRC) Funded Grant, UK (No. EP/Y028805/1).
Machine Intelligence Research (original title: International Journal of Automation and Computing) is published by Springer and sponsored by the Institute of Automation, Chinese Academy of Sciences. The journal publishes high-quality papers on original theoretical and experimental research, targets special issues on emerging topics, and strives to bridge the gap between theoretical research and practical applications.
Designing an effective antibody drug is like searching for the right key in a warehouse of locks. Scientists may begin with millions—or even billions—of antibody candidates, but only a tiny fraction will recognize and bind tightly to the disease target. Identifying those rare candidates has long been one of the biggest challenges in developing antibody medicines.
Boston University researchers have now developed an antibody-specific AI framework that dramatically narrows that search. Rather than building a larger AI model, the team redesigned how AI learns, focusing it on the small regions of antibodies that recognize disease targets.
"Instead of treating antibodies like generic proteins, we designed an antibody-specific language model that learns the fundamental patterns in the regions responsible for antigen binding," says Diane Joseph-McCarthy, PhD, the study's principal investigator and executive director of Boston University’s Bioengineering Technology & Entrepreneurship Center. "That focused approach helps researchers identify the most promising therapeutic candidates before they ever enter the laboratory."
Artificial intelligence has transformed biology by identifying patterns across millions of protein sequences. Like ChatGPT predicts missing words, protein language models predict masked amino acids to learn the "language" of proteins.
For most proteins, randomly hiding amino acids throughout a sequence is an effective training strategy because biologically important information is distributed across the molecule. Reconstructing the missing pieces helps the model discover the patterns that determine protein structure and function.
Antibodies, however, are different.
The regions responsible for recognizing viruses, bacteria, and cancer cells continually evolve so the immune system can adapt to new threats. That diversity makes antibodies extraordinarily powerful—and more difficult for AI to model.
Most of an antibody serves as a structural scaffold. The information that determines what an antibody recognizes and how tightly it binds is concentrated within six tiny loops called complementarity-determining regions, or CDRs.
"Think of an antibody like a screwdriver," says John Misasi, MD, a study co-author and assistant professor of virology, immunology, and microbiology at Boston University’s Chobanian and Avedisian School of Medicine. "It doesn't matter whether it's long or short—the shape of the tip determines what kind of screw it fits. The CDRs are like that tip: they determine which target the antibody recognizes."
Teaching AI the Biology That Matters
Recognizing that antibodies break many of the assumptions behind general protein language models, the researchers redesigned the training process around antibody biology instead of treating every amino acid as equally important.
The model focused its learning on the CDRs—the regions directly responsible for recognizing disease targets—and was trained using more than 1.6 million naturally paired antibody heavy and light chains that together form the binding site. During training, the researchers deliberately masked up to half of the amino acids within the CDRs while leaving most of the surrounding antibody structure intact, repeatedly challenging the AI to reconstruct the regions most critical for binding.
"A lot of AI research has focused on building larger models," says Ioannis (Yannis) Paschalidis, PhD, a co-author of the study and director of Boston University’s Hariri Institute for Computing. "We asked a different question: How can we teach the model the biology that matters most? That turned out to be a much more effective strategy."
The result was a smaller, more focused model containing about 600 million parameters that matched or outperformed much larger antibody language models on multiple benchmark tests. It improved binding affinity prediction by as much as 27 percent across datasets containing more than 90,000 engineered antibody variants targeting six different antigens. By training on millions rather than billions of antibody sequences, it required substantially less computational effort.
"One of the exciting findings is that we didn't need a larger model or vastly more data,” says Paschalidis. "That’s similar to what's been observed with human-language AI models, where smaller domain-specific models trained on high-quality data can often outperform much larger, more general ones."
The approach emerged from a convergent research effort that brought together expertise in artificial intelligence, immunology, structural biology, and experimental science—not simply to apply AI to biology, but to redesign how AI learns using biological knowledge.
From Prediction to Prioritization
The study addresses one of the biggest bottlenecks in antibody discovery: deciding which candidates to test. Even small changes to an antibody's sequence can create trillions of possible variants—far more than laboratories can realistically evaluate experimentally. The model helps narrow those possibilities before laboratory testing, reducing unnecessary experiments and accelerating antibody optimization.
"Knowing not just whether an antibody binds, but how strongly it binds, gives researchers a much better starting point for deciding which candidates to move forward," says Misasi, core faculty at BU’s National Emerging Infectious Diseases Laboratories (NEIDL). "If a computer can narrow millions of possibilities down to the few hundred most promising candidates, that saves an enormous amount of time, labor, and cost in the laboratory."
Beyond selecting antibody candidates for testing, the approach could improve antibody engineering. Researchers could use the model to predict which sequence changes are most likely to strengthen existing antibodies, helping optimize therapies against evolving viruses or other disease targets before moving those designs into experimental testing.
"If another infectious disease outbreak occurs, we'd like to identify promising antibody candidates as quickly as possible," says Misasi. "Computational tools like this could help us find those candidates sooner and even suggest how existing antibodies might be adapted as viruses change over time."
Looking Ahead
The findings suggest that biologically informed AI may offer a more effective path for antibody discovery. Beyond this study, the researchers believe biologically informed AI could ultimately transform therapeutic antibody development by helping scientists better understand antibody-antigen recognition, prioritize candidates for laboratory testing, and design more effective therapies.
"Understanding how antibodies recognize their targets is fundamental to developing better antibody therapies, diagnostic tests, and vaccines," says Joseph-McCarthy. "By teaching AI the biology that matters most, we hope to give researchers better tools to discover, optimize, and ultimately design the next generation of antibody therapies."
Pipeline for predicting antibody–antigen binding affinity. A protein language model encodes the antibody sequence, and a downstream model uses this representation to predict binding affinity to the target antigen.
The Y-shaped antibody contains heavy (VH) and light (VL) chains that form an antigen-binding site, including six complementarity-determining regions (CDRs) that recognize and bind the antigen.
Physics-guided, validity-gated multi-objective Bayesian optimization for CMOS LNA design. The method improves valid simulation yield from 47.5% to 95.7% and achieves 62% lower power and 10.7 dB better linearity than manual design within 190 simulations.
Credit: Jithish Jayarajan/Singapore University of Technology and Design; Kiat Seng Yeo/Singapore University of Technology and Design, Tianjin University; Shuoyu Ji, Bharatha Kumar Thangarasu, Nagarajan Mahalingam/Tianjin University; Baoji Miao/Henan University of Technology
A physics-guided AI framework automatically optimizes 2.4 GHz CMOS low-noise amplifiers, reducing power consumption by 62% and improving linearity by 10.7 dB compared with an expert-designed baseline, while nearly doubling the rate of successful circuit simulations under foundry design constraints.
Every smartphone, Wi-Fi router and wearable device relies on a tiny circuit called a low-noise amplifier (LNA) to extract faint radio signals out of the air without drowning them in noise. Designing one is a delicate balancing act: power consumption, noise and signal fidelity pull against each other, and a skilled engineer typically spends days of iterative simulation and hand-tuning to get it right.
Automating that job has proven surprisingly hard for artificial intelligence. Chip designs must obey the strict rules of a foundry's process design kit (PDK). For example, on-chip inductors can only be chosen from a fixed library of pre-characterized components and most randomly generated candidate circuits simply fail. In the team's preliminary experiments, more than half of the simulated designs were invalid, violating transistor operating regions or impedance-matching requirements before the optimization could even get started.
Researchers from the Singapore University of Technology and Design, Tianjin University and Henan University of Technology now report a way around this bottleneck: teach the algorithm some physics before letting it search. Much like handing a student the textbook before the exam, their framework first uses simple resonance and impedance-matching relationships to generate “warm-start” designs that are physically sensible and buildable from the foundry's component library. A validity gate then screens every simulated candidate, so that only physically meaningful results are used to train the artificial intelligence (AI) model that steers optimization toward the best power–noise–linearity trade-offs.
The payoff is striking. Optimizing a 2.4 GHz LNA in a commercial 40 nm CMOS process with a budget of just 190 circuit simulations, roughly 16–19 hours of computing, versus days of manual iteration. The framework cut power consumption from 14.1 mW to 5.4 mW, a 62% reduction, while improving linearity (IIP3) by 10.7 dB compared with a handcrafted expert design, at the cost of a modest 0.21 dB noise-figure penalty. The share of valid simulations jumped from 47.5% during unguided sampling to 95.7% under guided optimization, roughly twice the rate achieved by popular open-source optimizers under the same budget.
To show the recipe is not a one-off, the team applied the identical pipeline to a structurally different amplifier. Without retuning the algorithm, it again outperformed the expert baseline, reducing power by 25%, lowering the noise figure by 0.36 dB and improving linearity by 2.59 dB.
The authors note that the study is a schematic-level proof of concept; future work will extend the framework to layout-extracted designs, process–voltage–temperature corners and other RF blocks such as power amplifiers, mixers and oscillators. Ultimately, the approach points toward practical, physics-aware AI design assistants for the analog circuits that connect our devices to the world.
This paper “Physics-guided multi-objective Bayesian optimization for process-aware CMOS LNA design” was published in Interdiscipline. Jayarajan J, Ji S, Thangarasu B, Mahalingam N, Miao B, et al. Physics-guided multi-objective Bayesian optimization for process-aware CMOS LNA design. Interdiscipline 2026(1):0002, https://doi.org/10.55092/interdiscipline20260002.
This Viewpoint discusses the role of artificial intelligence (AI) in ensuring that patients eligible for highly effective interventions are identified and receive them. Authors propose a public health agenda for artificial intelligence that focuses on improving delivery of proven health interventions through patient identification, outreach, and care coordination. They argue that realizing this potential will require government leadership, cross-sector partnerships, and sustained investment.
Visit JAMA+ AI for editor-curated JAMA Network research, commentary, multimedia, and educational tools.
Corresponding Author: Adam L. Beckman, MD, MBA, Health and Opportunity Leadership Institute, City College of New York, 160 Convent Ave, New York, NY 10031 (abeckman1@ccny.cuny.edu).
Link to the article in your story We encourage you to link out to this article in your story using the link below. It includes an access token that will give free access to the article for your readers up to one year after publication. (The link will be live after the article publishes and embargo is lifted.)
Editor's Note: Please see the article for additional information, including full author list, author contributions and affiliations, conflict of interest and financial disclosures, and funding and support.
Journal
JAMA
Variations between labs can misinform scientific AI models, team reports
A study conducted at four laboratories across the country demonstrated how experimental protocol and equipment standardization govern result variability. Recognizing this variability is essential when incorporating real-world data into artificial intelligence/machine learning models.
Credit: Provided by Greg Stewart/SLAC National Accelerator Laboratory
UNIVERSITY PARK, Pa. — To turn abundant carbon dioxide into valuable fuel, researchers need a fast and efficient way of determining which catalysts work best over the longest time. Artificial intelligence (AI) models have the potential to help guide catalyst selection, but as with internet chatbots, AI models are only as good as the data put into them.
By convening four laboratories from across the nation to test an experimental carbon monoxide-producing catalyst, a key first step in turning carbon dioxide into fuels, scientists at SLAC National Accelerator Laboratory, alongside two researchers from Penn State, have demonstrated the importance of generating highly reproducible experimental data when building AI models for investigations in science. They published the results inNature Catalysis.
“An AI model is only as good as the data used to create it, which will undoubtedly come from multiple sources,” said Robert Rioux, Friedrich G. Helfferich Professor of Chemical Engineering at Penn State and co-author on the study. “This study represents the first round-robin study focused on heterogeneous catalysis that quantifies the uncertainty from studies done between labs, and the impact this uncertainty will have on future data-driven AI modeling.”
Speeding up catalyst development – with AI
With a good AI model, researchers can enter conditions such as temperature, length of time of the reaction, and catalyst formulation, then run the simulation and see a prediction of how well the catalyst performs. They can then confirm the predictions with a few well-designed experiments, ultimately speeding up catalyst discovery and implementation at a global scale.
In addition to saving time and money, such models can also explore conditions that are difficult to achieve in the lab. Most lab catalysis studies can only look at short time periods — days — but catalyst deactivation occurs over the course of months or years due to buildup of impurities and repeated exposure to high temperatures.
AI models need large amounts of high-quality data for training. To generate the data, the four labs performed a set of round-robin experiments, in which multiple laboratories conduct the same tests to evaluate reproducibility using previously agreed upon protocols and the same rhodium-based catalyst.
Squaring the data from round-robin experiments
To the researchers’ surprise, achieving the same results from four labs working independently was harder than anticipated. When they got together to share their results, they realized that they had a problem.
Each of the four research teams produced results that varied in the amounts of carbon monoxide and methane, an undesirable side product, produced. The computer would not be able to learn from four sets of data that contain different outcomes.
“It was a bit of an eye-opener,” said SLAC staff scientist Adam Hoffman, senior author of the study. “This experience shines light on the practical challenges of including real-world data into machine learning models.”
Painstakingly, the teams evaluated their methods. Through rigorous testing, they found a handful of sources of mismatch, with one of the biggest contributors to the variability coming down to how hard the mixture was shaken or stirred.
“Our findings are a reminder to exercise caution about what information we feed into a machine-learning model, and how the consistency of experimental data can influence the reliability of the outcomes,” said Selin Bac, a postdoctoral researcher at the University of California, Santa Barbara, and first author on the study.
With further standardization across the four labs — which in addition to SLAC included groups at Penn State, Stanford University and University of California, Santa Barbara — the results began to look more consistent. The team outlined several recommendations to strengthen experimental reproducibility, including enhancing the consistency of reactor design, operating protocols and experimental conditions.
“We contributed to the characterization of the catalytic materials used in the round-robin study,” Rioux said. “Our findings reinforce the idea that high-quality experimental data dictates the success of any AI-based model by determining model accuracy, reliability and scope.”
Hoffman said he hopes the study will help experimentalists and data scientists who are designing AI models to consider how small variations in experimental design across labs can lead to problems with reproducibility and impact on long-term predictions for AI modes.
“We see this work as a guide for the community as to how to think about designing experiments for inclusion in machine learning models,” Hoffman said.
Greg Barber, assistant professor of chemistry at Penn State Altoona and affiliate researcher in the Institute of Energy and the Environment and co-author on the paper, also contributed to this research.
This work was supported in part by the U.S. Department of Energy's Office of Science under award number FWP 101064. Testing equipment was supplied in part by Co-ACCESS, part of the SUNCAT Center for Interface Science and Catalysis, a joint research center supported by SLAC National Accelerator Laboratory and Stanford University. This content is solely the responsibility of the authors and does not necessarily represent the views of the Department of Energy.
Large language models (LLMs) like ChatGPT are widely used for public health information retrieval. Prior urological studies have mainly validated ChatGPT performance based on AUA and EAU guidelines, while evidence aligned with CUA guidelines remains scarce. This study by Wyatt MacNevin’s team from Dalhousie University evaluated the appropriateness and reliability of ChatGPT-4.0 responses to common patient-oriented urological questions using CUA guidelines as the benchmark.
The research team selected 10 common urological questions covering nephrolithiasis, prostate cancer, benign prostatic hyperplasia (BPH), erectile dysfunction (ED), overactive bladder (OAB), urinary tract infections (UTI), andrology, hypogonadism, pediatric urology, and kidney cancer. Each question was phrased in plain language and submitted to ChatGPT‑4.0 (March 2025 version) in three independent sessions, generating 30 responses. Three evaluators—two practicing Canadian‑licensed urologists and one urology resident—graded each response on a four‑point Likert scale (0 = Completely incorrect, 1 = Some correct and some incorrect, 2 = Correct but inadequate, and 3 = Comprehensive), with an “appropriate” response defined as a score ≥ 2.00.
Of the 30 responses generated by ChatGPT‑4.0, only 40% (12/30) met the appropriateness threshold, with an overall mean score of 1.64 ± 0.85 falling between "some correct and some incorrect" and "correct but inadequate." When stratified by question difficulty, easy questions scored significantly higher than medium‑difficulty questions (1.87 ± 1.01 vs. 1.31 ± 0.47, p < 0.05), confirming that ChatGPT performs best on fact‑based, definitive queries. In domain‑specific analyses, prostate cancer, ED, andrology, and kidney cancer received perfect median scores of 3.00, likely due to the abundance of standardized online information on these topics, whereas UTI, OAB, nephrolithiasis, and hypogonadism scored lower—even when rated as “easy”—possibly reflecting inconsistent online content or discordance between CUA and other international guidelines. The mean variance across repeated questions was 0.27, indicating strong model consistency and reliability.
ChatGPT has gained popularity as a healthcare information source, yet most existing studies have referenced American or European guidelines. This is one of the few studies to assess its performance against Canadian standards. The 40% appropriateness rate signals potential, but also highlights that current outputs are insufficient for independent patient use.
Urologists should be aware that patients using ChatGPT may receive incomplete or inaccurate information. Proactive patient education about the model’s limitations is essential. The authors recommend mandatory disclaimers on medical responses and future integration of authoritative guidelines into LLM training—alongside prospective studies evaluating real‑world clinical impact.
ChatGPT produced appropriate answers for 40% of general urology questions when benchmarked against CUA guidelines. While promising, further refinement and rigorous validation are required before widespread clinical adoption can be recommended.
The work titled “Assessing the utility of a natural language processing model in answering common urological questions” was published in UroPrecision (published on August 20, 2025).