
Image: ChatGPT
August 27, 2026
By Burak Oktenli
Key Takeaways:
A computational prediction of stability is not the same as a usable technology; real materials must clear successive hurdles of synthesizability, reproducible manufacturing, and performance under operating conditions.
The author argues the next bottleneck is evidence infrastructure—such as a standardized “materials AI evidence passport” that tracks model uncertainty, synthesis, independent measurement, and qualification—so scarce experimental resources go to the candidates most worth proving.
Generative models are rapidly expanding the search for new batteries, semiconductors, catalysts, aerospace materials and defense technologies. The next bottleneck is no longer finding candidates. It is proving that a predicted material can actually be synthesized, manufactured and trusted in the real world.
Materials science has traditionally suffered from a scarcity problem. Researchers could explore only a small fraction of the almost unimaginable number of possible compounds, structures and compositions that might possess useful properties.
Artificial intelligence is beginning to reverse that problem.
The emerging challenge may be abundance.
Machine-learning systems can now screen enormous chemical spaces, predict properties and increasingly generate candidate materials designed around specified characteristics. What once required researchers to select a relatively small number of hypotheses can increasingly become a computational search across thousands, millions or even more possible structures.
That is a remarkable scientific advance.
It also creates a new question: What happens when artificial intelligence can propose promising materials much faster than laboratories can determine whether those materials are real, manufacturable and useful?
The answer matters far beyond materials science.
Advanced materials sit beneath many of the technologies governments now consider strategically important: batteries, semiconductors, solar cells, carbon capture, aerospace systems, nuclear technologies, medical devices and defense equipment. The countries and companies that learn to convert AI-generated candidates into reliable physical materials will possess an advantage that cannot be measured simply by the number of structures their algorithms produce.
The new race is therefore not only to discover materials faster.
It is to validate them faster without confusing prediction with proof.
From Scarcity to Abundance
The scale of the shift became visible with Google DeepMind’s GNoME project.
In a 2023 Nature paper, the researchers reported more than 2.2 million crystal structures stable relative to previously known materials, with approximately 381,000 appearing on an updated stability frontier. The study represented an extraordinary expansion of the computationally accessible materials landscape.
But the same work also illustrates the distinction between discovering a computational candidate and possessing a usable material.
The paper reported 736 structures that had been independently experimentally verified. Its authors also identified synthesizability, dynamic stability and phase behavior among the remaining challenges between computational discovery and real-world application.
That gap is not a weakness of the research. It is the next scientific problem.
The transition is already visible in practical materials research. AI-guided, high-throughput experiments have been used to search enormous molecular spaces for improved photovoltaic materials, combining computational selection with automated synthesis and direct measurement in working solar cells. The important part of that workflow is not AI alone. It is the closed loop between prediction, synthesis and experiment.
Microsoft’s MatterGen provides another indication of where the field is heading. Rather than merely screening existing candidates, MatterGen can generate inorganic materials conditioned on desired characteristics, including mechanical, electronic and magnetic properties.
Its researchers went an important step further: they experimentally synthesized one AI-designed material and found its measured property to be within roughly 20 percent of the intended target.
That experiment is significant precisely because it crossed the boundary from computational proposal to physical evidence.
As generative systems improve, however, the number of proposals could grow much faster than the number of candidates that can receive comparable experimental attention.
A model can create another structure almost instantly. A laboratory cannot create another characterization campaign almost instantly.
Synthesis requires equipment, expertise, raw materials and time. Characterization requires instruments. Manufacturing introduces defects and process variability. Environmental testing requires additional facilities. Component qualification may take months or years.
AI can compress one part of the scientific pipeline without automatically compressing the rest.
A Stable Crystal Is Not Yet a Technology
One source of confusion is that the word “discovery” can describe very different stages of evidence.
A machine-learning model may predict that a crystal structure is energetically stable. That is scientifically useful. But stability under a computational method does not automatically establish that the material can be synthesized economically, manufactured reproducibly or maintained under operating conditions.
Different layers of uncertainty enter at different stages.
First comes model uncertainty. A machine-learning system is most reliable in regions sufficiently represented by its training and validation data. Materials discovery is difficult precisely because genuinely interesting candidates may lie outside that familiar domain.
Then comes reference-physics uncertainty. Machine-learning models are often trained against calculations based on methods such as density functional theory. Those calculations are extraordinarily valuable, but they remain approximations whose accuracy can vary with chemistry, structure and the property being predicted.
Next comes physical and manufacturing uncertainty. A perfect computational crystal is not necessarily the material produced by an industrial process. Defects, grain boundaries, impurities, temperature histories and manufacturing tolerances can change behavior.
Finally comes application uncertainty.
A battery material must survive repeated electrochemical cycling. A turbine material must tolerate extreme heat and mechanical loading. A semiconductor must perform reliably at manufacturing scale. A space material may face radiation, vacuum and severe temperature cycling. A defense material may need to survive shock, vibration, corrosion, aging, extreme thermal conditions or other highly demanding environments.
No single confidence score from a discovery model can represent that entire chain.
The Missing Infrastructure Is Evidence
This suggests that the next important innovation in AI-enabled materials science may be less glamorous than another generative model.
The field needs a common way to record what has actually been demonstrated.
One approach would be a materials AI evidence passport: a standardized evidence record that follows an AI-generated candidate from computational discovery through experimental validation and, where appropriate, manufacturing and qualification.
The passport would not decide whether a material is “good” or “bad.” It would make the status of the evidence legible.
At the computational stage, it could record the model and version used, the relevant training-data domain, the reference calculation, the conditions under which the model was validated and an empirically tested estimate of uncertainty.
If a candidate lies substantially outside the model’s validated domain, that fact should travel with the candidate rather than disappearing behind a high prediction score.
Recent work on AI-assisted alloy discovery points in the same direction: useful discovery systems increasingly need to distinguish confidence from uncertainty and identify regions where the available evidence is insufficient, rather than merely rank candidates.
The next layer would document independent computational verification where appropriate.
After that would come the physical record: whether synthesis has been achieved, whether composition and structure have been confirmed, whether the predicted properties have been measured and whether results have been reproduced independently.
Later stages could record whether a manufacturing route has been demonstrated and whether the material has survived testing under conditions representative of its intended use.
A scientist, investor, manufacturer, government laboratory or program manager should be able to look at a candidate and immediately distinguish between three very different statements:
The model predicts this should work.
We have made it and measured the relevant property.
We can manufacture it reproducibly and it works in the environment for which it is intended.
All three statements are valuable. They are not equivalent.
Why This Would Accelerate Science Rather Than Slow It
Standardized evidence can sound bureaucratic, particularly in a field whose attraction lies partly in accelerating discovery. But the purpose would be the opposite.
As computational candidate generation becomes cheaper, experimental capacity becomes relatively more scarce.
The scientific problem becomes one of allocation.
Which candidates deserve expensive synthesis? Which deserve synchrotron time? Which should proceed to manufacturing experiments? Which require additional calculation first? Which apparently spectacular result is simply too far outside a model’s validated domain to justify immediate investment?
An evidence passport would allow laboratories to direct scarce physical resources toward candidates with the strongest combination of potential value and credible supporting evidence.
It could also make results more portable.
Different universities, national laboratories and companies do not need to use identical AI models or surrender proprietary datasets. But they could use a common grammar for describing what a model has established and what physical tests remain incomplete.
That distinction is important. Scientific standardization does not require methodological uniformity.
Researchers can disagree about models while still agreeing that provenance, uncertainty, synthesis and physical validation should be visible.
The same principle already operates throughout mature engineering disciplines. A component rarely becomes trustworthy because its designer announces a confidence score. Trust accumulates through documented testing, traceability, calibration, independent measurement and experience under increasingly representative conditions.
AI-generated materials should not be exempt from that logic simply because the front end of discovery has become computational.
The Strategic Implications Are Larger Than Defense
My original interest in this problem came from defense applications, where the consequences of weak validation can be unusually severe.
AI can help search for energetic compounds, thermal-protection materials, armor, radiation-tolerant components and materials designed for extreme environments. But a computational prediction cannot substitute for the destructive and environmental testing required before such materials enter operational systems.
Defense is therefore a useful stress test for the broader problem. It is not the only sector facing it.
The energy transition will depend on new battery chemistries, catalysts, photovoltaic materials and materials for electricity transmission and storage.
Semiconductor progress increasingly depends on materials with carefully controlled electrical and thermal characteristics.
Fusion systems require materials capable of surviving environments that are extraordinarily difficult to reproduce.
Space exploration requires lightweight structures, radiation tolerance and long-duration reliability.
Medical technologies introduce their own requirements for safety, biocompatibility and reproducibility.
Across all of these fields, AI can accelerate the search. Physics still decides whether the result works.
The Next Bottleneck
The history of technological development repeatedly shows that discovery and deployment operate at different speeds. Artificial intelligence may make that mismatch far more visible.
The computational side of materials science is entering an era in which proposing a new candidate can become extremely cheap. That does not make experimental science obsolete. It makes experimental science more valuable.
When millions of possible candidates compete for limited laboratory attention, the ability to determine which computational claims deserve physical verification becomes a strategic scientific capability in its own right.
The most successful AI-for-science systems will therefore not simply generate the largest number of new materials. They will create better loops between prediction and experiment.
Models will propose. Experiments will test. Failures will return information to the models. Manufacturing will reveal forms of uncertainty invisible in idealized calculations. Real operating environments will expose limits that neither simulation nor laboratory characterization could fully anticipate.
That closed loop – not generation alone – is where the real acceleration of materials science will occur.
AI may be able to propose tomorrow’s battery cathode, semiconductor, radiation shield, catalyst or armor material in hours.
The important question is whether science can tell, nearly as quickly, what has actually been proved.

Credit: NATO
By Burak Oktenli
Key Takeaways:
Vendor diversity does not equal dependency diversity: a portfolio can appear multinational while still containing single points of failure in models, clouds, authentication, updates, or jurisdictions; commercial digital services can become unavailable overnight due to export controls, outages, or policy changes even when hardware remains intact.
Critical AI-enabled systems should carry an explicit dependency budget and be stress-tested for substitution time rather than supplier count, so allies can practice controlled dependence—using the best available technology while ensuring rapid, operable fallbacks—before the October implementation plan is finalized.
Ankara’s new industry strategy is designed to make defense production more scalable, interoperable, and resilient. The same discipline should apply to the AI models, clouds, identity services, and software dependencies that can disappear without a factory shutting down.
A missile shortage is easy to see. A missing cloud service may not become visible until the capability depending on it stops working.
That difference matters as NATO turns the commitments made at its July summit in Ankara into an implementation plan for a stronger transatlantic defense industrial base. The Alliance’s new Strategy for Industry-NATO Cooperation is unusually concrete: it calls for modularity and open architectures, stronger interoperability, more resilient supply chains, and tabletop exercises that stress-test whether production can surge and endure under crisis conditions.
Ankara also produced two practical mechanisms. The NATO Front Door for Industry is intended to simplify how companies find procurement, innovation, testing, and engagement opportunities. The NATO Engine is meant to connect industrial demand with available manufacturing capacity across the Alliance. Both respond to the same strategic reality: deterrence depends not only on possessing capability, but on being able to scale, sustain, and replace it.
There is one layer of the industrial base that deserves the same treatment before the October implementation plan is completed: the digital infrastructure underneath AI-enabled military capability.
Modern military AI increasingly arrives as a stack rather than a box. A system may depend on one company for the model, another for cloud hosting, a third for identity and access management, proprietary interfaces for integration, external services for updates and evaluation, and data pipelines governed in yet another jurisdiction. The nationality of the prime contractor tells only part of the story.
A procurement portfolio can therefore look diversified while retaining a single digital point of failure.
Vendor Diversity Is Not Dependency Diversity
The problem became visible in June when a U.S. export-control directive required Anthropic to restrict access to its Fable 5 and Mythos 5 models for foreign nationals. Because the order took effect immediately and the company said it had no reliable way to verify nationality in real time, Anthropic suspended the models for all users. The controls were lifted on June 30, and access began returning the next day.
The point is not that NATO should avoid American AI services, nor that this particular episode predicts a future alliance crisis. The lesson is narrower and more useful: a commercially available digital capability can change availability because of a government order, export restriction, licensing decision, security incident, provider outage, contract dispute, or technical change even when every physical component remains intact.
Replacing the service may require much more than buying another subscription. A second model may expose different interfaces. Its outputs may need fresh validation. Security controls may have to be rebuilt. Data may need to move between jurisdictions. Operators may require retraining. Existing software may have been optimized around one provider’s architecture.
A replacement that exists commercially may therefore be unavailable operationally for weeks or months.
This is the digital equivalent of discovering that several weapon systems depend on the same scarce component. NATO already treats concentration risk in physical supply chains as a resilience problem. Digital concentration deserves the same precision.
Five contractors do not create resilience if all five ultimately rely on the same cloud, the same model provider, the same authentication layer, the same update service, or the same jurisdiction for a mission-critical function. A multinational supply chain can still contain a single switch.
Give Critical AI Capabilities a Dependency Budget
NATO’s implementation plan offers an opportunity to make that exposure measurable. Every critical AI-enabled capability should carry a dependency budget: an explicit account of how much operational capability rests on any single provider, technical service, interface, or jurisdiction, and how quickly that dependence can be substituted.
This does not require one universal percentage. A logistics-planning tool can tolerate a different dependency profile from cyber defense, intelligence analysis, air defense, or command-and-control support. What matters is that concentration becomes visible before a system is embedded deeply enough to make replacement prohibitively difficult.
For each critical capability, planners should map the model provider, compute environment, cloud operator, identity service, update authority, proprietary interfaces, data dependencies, evaluation services, cryptographic credentials, jurisdictional constraints, and fallback options.
Then ask operational questions. How much capability remains if the primary provider disappears for 24 hours? What remains after 30 days? Can an allied operator move to another model without rebuilding the surrounding software? How long would revalidation take? Can the system continue in a degraded local mode? Who controls the credentials, keys, updates, and interfaces required to make the transition?
Those questions turn digital sovereignty from a political slogan into an engineering property.
They also create a more useful measure than national origin alone. An American service may be entirely appropriate for a European military mission if substitution paths are credible and the conditions governing access are understood. A nominally European system may create greater vulnerability if its compute, software dependencies, or update chain ultimately converge on one external provider.
Stress-Test Substitution Time, Not Supplier Count
The Ankara strategy already calls for tabletop exercises that stress-test defense production under heightened demand and crisis conditions. Digital dependency should be added to those exercises.
One scenario could remove a major cloud or model provider from an allied workflow without warning. Another could impose a jurisdictional restriction on a critical software component. A third could assume that a commercial provider remains online but stops issuing trusted security updates.
The metric should be substitution time, not whether an alternative vendor exists on paper.
A fallback model that requires three months of integration and validation offers little resilience during the first week of a crisis. The same is true of an alternative cloud environment that cannot accept existing data, identities, credentials, or workloads without extensive reengineering.
This is where NATO’s emphasis on modularity, open architectures, digital standards, testing, verification, and lifecycle interoperability becomes strategically important. Open interfaces reduce switching costs. Common evaluation procedures make alternative models easier to qualify. Portable data and identity architectures reduce migration time. Contract terms can require providers to document critical dependencies and preserve workable exit paths.
Controlled Dependence, Not Digital Autarky
None of this requires NATO to abandon American technology or ask every ally to reproduce the frontier-AI ecosystem nationally. That would consume enormous resources and could fragment the Alliance technologically.
The more practical objective is controlled dependence.
Allies can continue using the best available models, clouds, and software while designing systems that remain operable when one layer changes. The United States benefits as well: allied confidence in American technology is stronger when reliance comes with tested continuity arrangements rather than an assumption of permanent availability.
The strategic issue is larger than procurement preference. Software services and AI infrastructure will increasingly determine whether physical military assets can be coordinated, maintained, upgraded, and used effectively. A defense industrial strategy that measures only factories, inventories, and production lines will miss part of the capability chain.
Ankara gave NATO a serious framework for strengthening the industrial base behind deterrence. The implementation plan due in October should recognize that part of that industrial base has no factory floor.
A missile shortage is visible in the warehouse. A digital dependency becomes visible when the mission stops. NATO should find it first.
Science Has Discovery Thresholds, It Also Needs Stop Rules – OpEd

Image: ChatGPT
Key Takeaways:
- Science is skilled at defining what would count as a discovery but often fails to specify in advance what evidence would close, downgrade, or pause an extraordinary claim; historical cases such as OPERA and BICEP2 show that the disappearance of an anomaly is itself a scientific success.
- AI makes anomaly hunting nearly limitless by generating unlimited candidates from vast datasets, which shifts the scarce resource from detection to adjudication and risks turning research programs into narratives that continually relocate rather than face decisive tests.
- Critical AI-enabled or high-profile searches should therefore carry explicit exit conditions—discriminating observables, conventional alternatives, calibration failures, independent replication standards, and null-result sensitivity thresholds—so that closure, downgrade, or pause become recognized scientific outputs rather than after-the-fact improvisations.
AI can make anomaly hunting nearly limitless. Researchers should define in advance what would close, downgrade, or pause an extraordinary claim not only what would count as a discovery.
Science is very good at celebrating the moment a result becomes interesting. It is less practiced at deciding when continued pursuit is no longer justified.
When the OPERA experiment reported a neutrino timing result that appeared to challenge the speed of light, the scientific response was not to protect the anomaly. Researchers attacked the timing chain, checked the instrumentation, and sought independent measurements. Later measurements were consistent with neutrinos traveling at light speed. The disappearance of the anomaly was not a failure of science. It was the science.
The BICEP2 episode made the same point in a different way. An apparent B-mode polarization signal was widely discussed as possible evidence of primordial gravitational waves. A joint BICEP2/Keck and Planck analysis later found strong evidence for dust and no statistically significant evidence for tensor modes in the analyzed data. The valuable result was not simply that a spectacular interpretation weakened. Science had narrowed what the observation could responsibly mean.
These cases expose a missing half of falsifiability. Scientists spend enormous effort defining what would count as evidence for a claim. High-cost and high-profile searches should also define what would count as enough evidence to close, downgrade, or pause one.
Falsifiability Needs an Exit Condition
Before an extraordinary physical claim consumes years of attention, researchers should be able to answer uncomfortable questions in advance. What observation would materially weaken the hypothesis? What calibration failure would invalidate the signal? What conventional explanation would be sufficient to end the extraordinary interpretation? At what sensitivity would a null result close the parameter range being tested? How many genuinely independent failures to replicate would lower the priority of the claim?
“More data” is not a falsification criterion.
There are good reasons to resist rigid stopping rules. Premature termination can bury real discoveries. Instruments improve. Background models change. A null result at one sensitivity may become a detection at another. Some theories remain scientifically valuable even when the decisive experiment is not yet technically possible.
But the opposite failure is real as well: a research program can become structurally incapable of losing. An anomaly appears and a conventional explanation removes most of it, so attention moves to a residual. The residual disappears and the search moves to another dataset. A replication fails and the failure is attributed to different conditions. A predicted signature is absent and the parameter range moves. None of those moves is automatically illegitimate. Taken together without an exit condition, however, they can turn an empirical program into a narrative that changes faster than it can be decisively tested.
A scientific program that cannot say what would make it stop is in danger of protecting a claim rather than testing it.
Stopping also does not have to mean abandoning a field. It can mean changing the status of a claim. “Discovery” becomes “candidate.” “Candidate” becomes “calibrated anomaly.” “New physics” becomes “model preference under stated assumptions.” One parameter region can close while another remains open. A team can conclude that an instrument lacked sufficient sensitivity, that a proposed signature was not discriminating, or that a conventional mechanism explains the observation without meaningful residual structure.
Those are scientific outputs. Closure is not the opposite of discovery; it is one way evidence becomes useful.
AI Makes Anomalies Cheap
Particle physics already recognizes part of this problem through stringent significance conventions and corrections for the look-elsewhere effect. Searching many channels makes an apparently striking local fluctuation less surprising, which is why the scope of the search has to be part of the evidence rather than an afterthought.
Artificial intelligence makes the broader stopping problem more urgent because it changes the economics of anomaly hunting. A system can scan enormous collections of spectra, images, light curves, detector events, candidate materials, or simulated physical states and rank the strangest examples. That is useful. It also means the supply of interesting outliers can become effectively unlimited.
When finding candidates becomes cheap, adjudicating them becomes the scarce resource.
An AI system can always produce another unusual point, another model fit, another candidate cluster, or another corner of parameter space worth inspecting. If every failed lead merely authorizes the next search without changing the status of the underlying claim, automation can make an already weak scientific habit scale much faster.
Synthetic data sharpen the problem. Imagine a classifier trained to distinguish simulated wormholes from simulated black holes. It performs spectacularly on a synthetic test set and then flags a real astronomical observation as wormhole-like. That may justify follow-up. It does not establish that a wormhole was detected. There are no confirmed wormhole examples on which to validate the label. The classifier may have learned differences between simulation pipelines, omitted astrophysical effects, or artifacts of the generators. Its success proves that it can separate the synthetic worlds it was given. Nature has not promised to resemble either one.
This is why anomaly detection and claim authority should remain separate. A search system can help decide where scientists look next. It should not decide what the observation is called—or whether a search has earned unlimited continuation.
Make Stopping a Scientific Output
The practical discipline is straightforward. Before the result becomes institutionally or emotionally expensive to lose, a project should specify the discriminating observable, the serious conventional alternatives, the calibration failures that would invalidate the signal, the search scope that must be corrected for, what replication would count as independent, the sensitivity at which a null result closes the tested region, and what evidence would trigger a downgrade or pause.
For expensive or extraordinary-claim programs, funders and review panels could ask for those continuation and termination conditions alongside the discovery criteria. The purpose would not be to impose a bureaucratic kill switch on scientific curiosity. It would be to make it harder to invent a new survival condition only after the old one fails.
The same principle should shape publication. Null results should be treated as positive scientific products when they close a meaningful parameter range. A calibration failure can close an anomaly. A conventional explanation can close an extraordinary interpretation. A failed replication can identify which dependency mattered. A search that reaches its prespecified sensitivity without detecting the predicted effect has produced information even if it does not produce a headline.
This matters increasingly in AI-assisted science because search capacity is growing faster than the scientific community’s capacity to investigate every candidate. The bottleneck is shifting from finding unusual things to deciding which unusual things deserve continued belief, money, instrument time, and attention.
Discovery is one way science advances. Elimination is another.
A mature research program should know not only what would make it celebrate, but what would make it stop.

About Burak Oktenli
Burak Oktenli holds an MBA and a Master of Professional Studies in Applied Intelligence from Georgetown University. His research addresses the governance of authority in autonomous and AI-enabled systems, and his writing has appeared at the Modern War Institute at West Point, RUSI, RealClearDefense, RealClearMarkets, and Geopolitical Monitor. He is the author of Authority Architectures for Autonomous Systems, a ten-volume series on how authority in autonomous systems is delegated, monitored and recovered, at authority-architecture.me.
View all posts by Burak Oktenli

