Tuesday, August 18, 2026

  


 

Every Measurement Has A Physical Limit: Scientific AI Pretends Otherwise – OpEd



August 18, 2026

By Burak Oktenli

Key Takeaways

Fundamental physical and statistical limits (Abbe’s diffraction limit, Cramér-Rao bound, Fisher information) mean that no amount of sophisticated analysis or AI can extract more spatial or parametric information than the underlying measurement actually contains.
 
AI models can produce confident numerical outputs even when the experiment is weakly informative, ambiguous or nearly non-identifiable, because confidence scores primarily reflect the model’s internal assumptions rather than the information content of the data.
 
Scientific AI systems should therefore incorporate structured, fail-closed abstention mechanisms that refuse to issue a qualified estimate when the measurement geometry, noise level or sensitivity is insufficient, and should report both error and coverage so that limits on the claim are treated as valuable scientific information rather than failure.


AI can produce a precise estimate even when an experiment no longer contains enough information to identify the target. Science should treat abstention as a valid result, not a failure.



In 1873, Ernst Abbe formalized a hard truth about microscopes: optical resolution is constrained by the wavelength of light and the aperture of the instrument. Below that limit, the problem is not that the lens needs a cleverer analyst. The measurement itself does not carry arbitrarily fine spatial information.

Seven decades later, C. R. Rao and Harald Cramér expressed the same idea in statistical language. Under the assumptions of the Cramér-Rao framework, Fisher information places a lower bound on the variance attainable by an unbiased estimator. The floor is set by the measurement model, its sensitivity to the quantity being estimated, and the noise. Better mathematics can use available information more efficiently. It cannot manufacture information that the experiment never recorded.

Nature keeps accounts.

Artificial intelligence is now becoming part of the analytical machinery of science. Learned models solve inverse problems, infer parameters from sparse observations, separate overlapping signals, classify images, and extract structure from data too large for manual inspection. Physics-informed neural networks, for example, have been used for forward and inverse problems in differential equations, while later work has explicitly confronted the identifiability of parameters inferred by such models.

These systems can be extraordinarily useful. But they introduce an easy category error: confusing confidence in an answer with evidence that the measurement could identify the answer in the first place.

A confidence score can tell us something important about the model under its assumptions. It does not, by itself, prove that the current experiment contains enough independent information about the target quantity. Priors, regularization, learned correlations, and familiarity with the training distribution can all keep an output numerically stable even as the measurement geometry becomes weak, ambiguous, or nearly non-identifiable. Research on predictive uncertainty has already shown that uncertainty quality itself can deteriorate when data move away from the conditions on which the model was developed.

Put more simply: confidence is evidence about the model. It is not automatically evidence about the experiment.

Consider three ordinary scientific situations. Two astronomical sources become so blended that their signatures are nearly indistinguishable. A navigation receiver loses the geometric diversity needed to separate state variables. Two spectral components become so similar that many combinations explain the same measured curve. An algorithm can still return a number in each case. The harder question is whether the data uniquely support that number without leaning on assumptions that have silently become more important than the measurement.

That distinction matters most near the edge of a scientific claim. Strong signals are comparatively forgiving. At the boundary, where a source is faint, two mechanisms are difficult to separate, or an effect would be consequential if real, the information supplied by the measurement can collapse faster than a model’s confidence score does.

Scientific AI therefore needs a capability that sounds unimpressive but is foundational to measurement science: the ability to refuse a qualified estimate because the evidence is physically insufficient.

Machine learning already has a research tradition called selective prediction, in which a system can abstain rather than predict on every case. SelectiveNet, for example, explicitly optimized the tradeoff between prediction risk and coverage. That is useful, but scientific inference needs an additional question upstream of model confidence: can the instrument, geometry, noise level, and measurement model identify the requested quantity at all?

Structural biology offers a useful cultural example. AlphaFold does not merely publish a three-dimensional structure; it also reports per-residue confidence. Researchers have learned not to treat low-confidence regions as inconvenient blanks to be cosmetically filled. In the human-proteome analysis, AlphaFold confidence was even evaluated as a predictor of unresolved or disordered regions. The lesson is not that a model-confidence score equals physical identifiability. It does not. The lesson is that a scientific community can learn to treat limits on a model’s claim as information rather than embarrassment.

The next step is to apply the same discipline to the measurement itself.

Before accepting an AI-assisted scientific estimate, ask what the physics permits. Is the target identifiable under the current geometry? How does the relevant Fisher information or sensitivity change as the experiment degrades? What uncertainty floor follows from the measurement model? Does a different parameter combination produce nearly the same observation? If the requested precision lies below what the data can support, a confident model output should not be upgraded into a confident scientific claim.

Science does not need another branded AI confidence score. Established tools already exist: conditioning diagnostics, Fisher information, Cramér-Rao bounds, posterior-width analysis under declared priors, residual checks, and explicit tests for distribution shift. The important change is architectural. The model may produce a candidate estimate, but release of a qualified scientific estimate should depend on evidence about whether the measurement supports it.

That permission should also be fail-closed. If the experiment is insufficiently informative, the output should not be a mysterious null or an invisible software exception. It should be a structured abstention: what quantity was requested, which information condition failed, what threshold had been fixed, what data and configuration were used, and what new measurement would be needed to make the question answerable.

There is no universal numerical cutoff. A telescope, a seismic array, a navigation geometry, a medical imager, and a plasma diagnostic lose information in different ways. The principle can travel across fields. The threshold cannot simply be copied from one instrument to another.

Nor should the abstention rule be tuned after the embarrassing cases are known. The information diagnostic, threshold-selection procedure, and acceptable error criterion should be fixed on development or calibration data and then applied unchanged to untouched evaluation cases. Otherwise abstention becomes another form of post-selection: the system learns to be humble only where the analyst already knows it was wrong.

This creates a straightforward standard for journals, reviewers, laboratories, and funders. An AI-assisted method that makes physical estimates should state the information limits of the measurement, show how those limits are checked at runtime or analysis time, report both error and coverage, and disclose when the system refuses to make a qualified estimate. A method that answers fewer questions may be scientifically stronger if it can explain why the rejected questions were not answerable from the evidence.

The goal is not to slow scientific AI. It is to keep AI from erasing a distinction that experimental science spent centuries learning: a calculation can be sophisticated while the observation remains insufficient.

A paper that cannot show where its measurement stops supporting its inference is not ready to turn a model output into a discovery claim. Editors should view that omission with the same instinct they bring to a p-value of exactly 0.049 after a long chain of analytic choices: not automatic rejection, but an immediate request to see the full evidentiary path.

We grade scientific AI today on the questions it answers correctly. We should start grading it on the questions it correctly refuses to answer.



About Burak Oktenli
Burak Oktenli holds an MBA and a Master of Professional Studies in Applied Intelligence from Georgetown University. His research addresses the governance of authority in autonomous and AI-enabled systems, and his writing has appeared at the Modern War Institute at West Point, RUSI, RealClearDefense, RealClearMarkets, and Geopolitical Monitor. He is the author of Authority Architectures for Autonomous Systems, a ten-volume series on how authority in autonomous systems is delegated, monitored and recovered, at authority-architecture.me.
View all posts by Burak Oktenli →



Washington may force Kazakhstan to choose between competing US, China AI alliances

Washington may force Kazakhstan to choose between competing US, China AI alliances
Kazakhstan's President Kassym-Jomart Tokayev with US counterpart Donald Trump in September last year. / Kazakh presidency

By bne IntelliNews August 18, 2026

The Trump administration is preparing to tell Kazakhstan and 34 other countries that they may have to choose between competing US-backed and China-backed artificial intelligence frameworks as Washington intensifies efforts to shape global technology alliances.

The scenario is outlined in an exclusive report by Reuters. A US official told the news service that partners “can’t have it both ways”.

A draft US State Department letter obtained by the Reuters is addressed to 35 countries that signed the US AI Opportunity Statement in Washington, DC in June. It warns that countries joining the US-led Pax Silica framework would not be able to participate in competing initiatives where the requirements conflict with those of the American framework.

The draft reportedly urges countries to “choose deliberately” and says membership of Pax Silica represents a commitment rather than simply a diplomatic designation. It is reported to state: "To be part of everything is to be part of nothing. Signature of the Pax Silica Declaration is not merely a membership subscription, but a commitment."

The letter has reportedly not been sent and could still be amended.

First to join

Kazakhstan was the first Central Asian country to join the US-led AI framework. Washington has highlighted the country’s importance as a potential essential supplier of critical minerals used in semiconductors and advanced technologies. However, Astana has also joined China’s World Artificial Intelligence Cooperation Organisation (WAICO). It was launched in Shanghai by Chinese President Xi Jinping in July.

It is thought Kazakhstan is presently the only country signed up to both the US and Chinese frameworks.

The competing commitments present a particular challenge for Kazakhstan. It pursues a multi-vector foreign policy aimed at maintaining relationships with all the major powers while avoiding excessive dependence on any one partner.

Kazakhstan, for instance, maintains strong economic ties with both China and Russia while seeking deeper cooperation with the US and European countries in areas including critical minerals, energy and technology.

The US-China competition over AI has intensified as Chinese open-weight AI models have made advances against proprietary systems developed by US companies including OpenAI and Anthropic, the report noted.

Washington is seeking closer alignment among partner countries partly to limit China’s access to critical minerals, semiconductor technologies and other resources considered important to the development of advanced AI systems.

The draft State Department letter indicates that Washington could increasingly make participation in its technology initiatives conditional on restrictions on cooperation with competing Chinese frameworks. A final US position has yet to be communicated to Kazakhstan.

A report by Tech Times argued that the draft has a major weakness – it does not specify what would constitute a conflicting commitment or what consequences countries would face for refusing to choose. It does not name China, set a deadline or establish a mechanism for determining violations and enforcing compliance, the report observed.

Weak enforcement mechanisms

The potential enforcement mechanism lies in access to advanced chips, the report said. AI accelerators require replacement every few years, creating a recurring dependence on supply channels controlled by the US and its partners. Washington could potentially restrict future chip access through export controls, a mechanism demonstrated in the case of the UAE's G42, which removed Huawei equipment as part of its closer alignment with the US.

However, applying such pressure to Kazakhstan could undermine Washington's objectives, the report argued. Kazakhstan's critical-mineral reserves are a major reason for US interest in bringing the country into Pax Silica, meaning sanctions or an exclusion over its WAICO membership could weaken the supply-chain diversification the framework seeks to promote.

The report also cited China analyst Rui Ma as saying that the approach could ultimately cost Washington more than it gains by encouraging countries to view supply chains as instruments of geopolitical pressure rather than as stable partnerships.

For now, the draft's call to “choose deliberately” remains a political signal rather than an enforceable requirement, the report concluded. 

AI shares human tendency to infer character from facial features




PNAS Nexus

faces 

image: 

Computer-generated images used in the study with questions about trustworthiness. The face on the right is usually seen as more trustworthy by both humans and LLMs. 

view more 

Credit: Alex Todorov





Human beings have a tendency to infer personality or character traits from other people’s facial features, and these biases—ungrounded in any actual relationship between faces and behavior—lead to unfair outcomes.

Steven Lehr and colleagues explored whether AI models, which are trained primarily on text but have the ability to “see” images, share these biases. The authors asked GPT-4o to make over 4,500 forced-choice judgments between computer-generated faces, and presented thousands of additional forced choices to GPT-5, Gemini 3 Flash Preview, and Claude Sonnet 4.5. In some of the experiments, models were asked to choose which of two faces was more competent or more trustworthy. Other trials used related traits, including asking which face was more confident, smart, hardworking, lazy, inept, careless, warm, helpful, sincere, selfish, hypocritical, or aggressive. Some experiments asked LLMs to judge which computer-generated human face would be more likely to be a serial killer, to be arrested for human trafficking, or to defraud the public using a Ponzi scheme. Finally, LLMs were asked to choose between faces in the contexts of hiring a university president, investing in a tech startup, or selecting a financial manager. In all these cases, the models were willing to weigh in and in the majority of cases chose the same face that a human would typically see as more trustworthy or confident. Across studies, GPT-4o selected the face that would be expected based on human ratings 74.88% of the time. GPT-5 showed notably more bias than its predecessor. When GPT-5 advised on consequential decisions, the model recommended the more competent-looking individual fully 97.04% of the time, as compared to GPT-4o’s 75.19%. Models from other companies produced similar results.

If LLMs were free of human face-to-character bias, they could be used as a tool to help eliminate this form of bias in contexts such as job candidate selection or parole decisions. According to the authors, AI in its present form is instead likely to worsen unfairness if used in such contexts.

Journal

Article Title

Article Publication Date

COI Statement

Hanbat National University researchers reveal physics-informed AI for rapid optimization of thermal energy storage systems



The proposed physics-informed neural network framework enables rapid, autonomous design optimization of latent heat thermal energy storage systems




Hanbat National University Industry–University Cooperation Foundation

Proposed PINN framework 

image: 

The proposed framework is trained on 15 ground-truth datasets, enabling rapid, autonomous exploration of optimal latent heat thermal energy storage system designs.

view more 

Credit: Assistant Professor Joo Hyun Moon from Hanbat National University





With intensifying climate change, decarbonization of the global building sector has become a key priority. A substantial portion of a building’s energy demands consists of heating and cooling needs. Consequently, developing efficient thermal energy storage systems is a crucial part of this effort. Among available options, latent heat thermal energy storage (LHTES) systems that utilize phase change materials offer unique advantages. These include a high energy storage density and the ability to release large amounts of thermal energy at a near-constant temperature, which is crucial for stable thermal management. Indeed, some studies have shown that LHTES systems can reduce heating and cooling energy consumption by up to 45%.

Despite these advantages, accurately modelling and optimizing LHTES systems remains a major challenge. The coupled heat-transfer and fluid-flow processes involved are difficult to simulate accurately. Although physical experiments provide reliable ground-truth data, they are generally limited to laboratory-scale systems. Computational fluid dynamics (CFD) simulations, on the other hand, can capture these complex physical processes across different scales. However, they are computationally expensive and time-consuming, making large-scale design optimization impractical.

In a new study, a collaborative team of researchers from the Republic of Korea, led by Assistant Professor Joo Hyun Moon from the Department of Building Systems Engineering at Hanbat National University in South Korea, has developed a hybrid physics-informed neural network (PINN) framework for optimization of LHTES systems. “The LHTES system utilizes a special wax-type material, called a phase change material, that soaks up a huge amount of heat when it melts and gives it back when it hardens, acting like a battery for warmth,” explains Dr. Moon. “Testing every new design using conventional computer simulations is slow and computationally intensive. In our PINN framework, we teach the AI model the governing laws of physics, enabling it to accurately reproduce the system's behavior and rapidly explore tens of thousands of possible designs.” Their study was made available online on May 05, 2026, and published in Volume 167 of the Journal of Energy Storage on July 30, 2026.

To develop the data-driven PINN, the researchers first created a high-fidelity ground-truth dataset that captures the physics of the LHTES system. For this, they developed a laboratory-scale experimental LHTES setup, and a corresponding CFD model. Experimental measurements were then used to validate the CFD simulations, which showed excellent agreement with only minor deviations. Once validated, the CFD model was used to generate a sparse dataset of 15 high-fidelity simulations that served as training data for the PINN.

The proposed PINN framework employs a zero-dimensional physical model and embeds the system’s governing physical equations directly as loss functions. A key innovation is that the PINN learns case-specific effective heat-transfer coefficients, while geometric effects are captured through a response surface model. By focusing only on discovering the laws of energy conservation, the PINN avoids overfitting despite being trained on sparse data. The response surface model (RSM) then enables the framework to predict heat transfer properties for any arbitrary, unseen geometry and flow condition during the optimization process. Together, the PINN and RSM form a fast digital twin of the LHTES system.

The digital twin is then coupled with a Non-dominated Sorting Genetic Algorithm II (NSGA-II) to perform multi-objective design optimization. The optimization simultaneously searches for designs that maximize total discharged heat and average discharge power while minimizing pumping power.

In numerical experiments, the PINN reproduced the behavior predicted by the CFD simulations with excellent accuracy while enabling rapid, autonomous design optimization. The optimal design obtained using the multi-objective design optimization process performed similarly to the best-performing baseline design, while significantly reducing pumping power. Moreover, the process also shows that flatter pipes are more favorable.

Beyond buildings, the same physics-informed design approach could help improve thermal management in electric-vehicle batteries, data centers, cold-chain logistics, and solar thermal systems. Earlier studies also suggest that smarter control of latent heat storage can cut electricity costs by more than 70%, showing the broader potential of this technology for reducing energy use and emissions.

“Our approach moves the LHTES design process from manual evaluation of discrete cases to autonomous exploration of the entire design space,” concludes Dr. Moon. “It provides engineers with a practical tool for developing more efficient thermal energy storage systems, helping reduce the energy consumption and carbon footprint of buildings while supporting a more sustainable energy future.

 

Reference
DOI: https://doi.org/10.1016/j.est.2026.122514

 

About the institute
Established in 1927, Hanbat National University (HBNU) is a university in Daejeon, South Korea. As a leading national university in the region, HBNU strives to take lead in solving problems in the local community and solidify its cooperation with industries. The university’s vision is to become a ‘Global industry-university cooperation university creating future values’. HBNU has been chosen for a variety of nationwide-level projects such as An Autonomous Improvement University Project and Leaders in INdustry-University Cooperation+(LINC+), among others. With its focus on practical education and regional impact, HBNU continually advances technological solutions grounded in creative thinking and real-world relevance.

Website: https://www.hanbat.ac.kr/eng/

 

About the author
Dr. Joo Hyun Moon is an Assistant Professor of Building Systems Engineering at Hanbat National University in Daejeon, Republic of Korea. He received his Ph.D. in Mechanical Engineering from Chung-Ang University in 2017. Before joining Hanbat National University, he served as an Assistant Professor at Sejong University and as a postdoctoral researcher at the University of Texas at Dallas. His research spans thermal system optimization, phase-change heat transfer, and physics-informed, machine learning-based models for real-time energy-system optimization.

CHART: Mining vs AI – Nvidia sneezes, sheds five BHPs


A pit battle. Image: Definitely AI generated

A year ago this week Nvidia became the first company to touch $4 trillion, and MINING.COM ran the numbers against the combined worth of the world’s 50 most valuable mining companies. It was not even close.

The chip designer was worth 2.7 times the MINING.COM TOP 50 (and since valuations quickly shrink outside that, probably twice the entire global mining industry).

This year Nvidia has been having the kind of quarter miners unfortunately know all too well. 

After peaking at roughly $5.5 trillion in the middle of May, the stock shed about $1 trillion over the next eight weeks. At the bottom of the slide, Nvidia was briefly trading at 18 times forward earnings – below the S&P 500 average, which made the poster child of the AI age, at least technically, a value stock.

The chipmaker’s P/E is now back near 20 times forward earnings, but mining’s majors change hands for barely 13 – cheaper on next year’s profits than Nvidia managed to look even at its bargain-bin bottom and despite all the criticality clamour surrounding mining. 

The TOP 50 had a quarter of its own. A record $2.4 trillion at the end of March, before gold’s slip back below $4,000 an ounce took $228 billion off the ranking, or for the Mag 7 slash Lag 7, a bad Friday afternoon. 

The scoreboard now reads $5.11 trillion versus $2.19 trillion. Nvidia is worth 2.3 times the top 50 miners, down from 2.7 times last July, and for a few sessions this month the multiple flirted with two – territory last visited when cheap and cheerful DeepSeek gave the AI trade its first proper scare

Over the past twelve months mining even outgrew the machine – up 47% against Nvidia’s 27% – a first in the short history of this exercise. 

But to return to theme of mainstream investors not valuing the production of copper the same way as the production of hallucinations about the production of copper, here’s a sobering thought:

What Nvidia shed between mid-May and early July – nearly five BHPs, and BHP has never been worth more – was about what the whole Top 50 was worth for most of this decade.

Renaming everything critical minerals was a good start (and kudos to met coal and lead for making the list), but it’s time to launch a worldwide public/investor awareness campaign. What about this old nugget for a tagline: If it can’t be grown, it has to be mined. 

That includes, for the record, the silicon, copper, gold, silver, tungsten, tantalum, titanium, cobalt, aluminium, tin, nickel, hafnium, ruthenium, molybdenum, indium, palladium, gallium, germanium, arsenic, antimony, bismuth, boron, phosphorus, cerium, lanthanum, yttrium, gadolinium, europium and praseodymium that Nvidia’s chips are made from.


Nvidia Backs 8-GW Ohio AI Campus for OpenAI With $1.5 Billion Investment

Nvidia is backing a massive artificial intelligence infrastructure development in Ohio that could ultimately require at least 10 gigawatts of new power generation, underscoring the rapidly growing impact of AI data centers on U.S. electricity demand.

The chipmaker said it has partnered with SB Energy to secure land, power and data center shell capacity at the PORTS-Pike Technology Campus in Pike County, Ohio, where OpenAI will be the customer under a 20-year lease.

The first phase is designed to provide 4.25 gigawatts of IT capacity, with Nvidia holding an option covering another 3.75 GW. At full buildout, the project would provide 8 GW of AI computing capacity.

SB Energy, which is backed by Japan's SoftBank Group and OpenAI, will build, own and operate the data center infrastructure. Nvidia will exclusively supply the site's AI computing systems, including GPUs, CPUs and networking equipment through its DSX AI factory platform.

The planned scale of the development highlights the increasingly close relationship between AI infrastructure and the energy sector. SB Energy and SoftBank plan to develop at least 10 GW of new electricity generation to support the campus and invest at least $4.2 billion in regional grid infrastructure through a partnership with AEP Ohio.

The companies said the arrangement is structured to prevent the costs of serving the data center campus from being shifted onto existing electricity customers.

Development is expected to proceed in phases beginning in 2028.

The campus is being built around the former Portsmouth Gaseous Diffusion Plant, a major Cold War-era uranium enrichment site in southern Ohio. The developers are working with AEP Ohio as well as the U.S. Departments of Energy and Commerce to redevelop the area into a large-scale technology and power hub.

Nvidia will also invest $1.5 billion directly in SB Energy, joining SoftBank and OpenAI as investors in the infrastructure company. The funding will support SB Energy's broader expansion as demand grows for purpose-built electricity and data center infrastructure serving AI computing.

The project adds to a wave of multibillion-dollar data center developments that are transforming U.S. power demand forecasts. Hyperscalers and AI companies are increasingly seeking dedicated generation, transmission capacity and long-term power arrangements as access to electricity becomes one of the principal constraints on new computing capacity.

The PORTS-Pike development is particularly notable for its scale: 10 GW of generation capacity would be comparable to the output of several large nuclear power stations and represents a major new source of electricity demand concentrated at a single industrial site.

The developers also plan significant local investment. OpenAI has added $40 million to an existing $40 million community benefits fund established by SB Energy, bringing the initial fund to $80 million. The money is intended for programs including energy affordability, workforce development and regional economic development.

The companies said the project could ultimately support tens of thousands of jobs in Ohio.

By Charles Kennedy for Oilprice.com

No comments: