Wednesday, August 19, 2026

When Machines Search For Discoveries, Science Has To Count The Search – OpEd


A section of the LHC (Large Hadron Collider). Photo by alpinethread, Wikipedia Commons.

August 19, 2026

By Burak Oktenli

Key Takeaways:

The 750 GeV LHC photon excess illustrated the classic “look-elsewhere effect”: a locally impressive anomaly can lose significance once the full breadth of a search is accounted for.

Adaptive AI pipelines amplify this problem because they generate and prune hypotheses during the search, making the true multiplicity hard to see if only the winning result is reported.

Robust AI-assisted discovery requires logging the full search trace, calibrating against realistic null replays, freezing rules before final confirmation, and disclosing search history so machine-scale exploration remains scientifically accountable.

AI systems can scan millions of hypotheses before surfacing one anomaly. The statistical evidence must account for the search that produced it.

In late 2015, physicists at CERN found something irresistible in the new data from the Large Hadron Collider: an excess of photon pairs near 750 GeV, where no known particle was expected. ATLAS reported a local significance approaching four sigma for one interpretation of the bump. CMS saw a smaller compatible excess. For a few months, the possibility of a new particle became one of the hottest questions in high-energy physics.

Then the 2016 data arrived. The bump did not. CERN concluded that the much larger data set had failed to reproduce the excess and that it appeared to have been a statistical fluctuation.


The episode was a vivid lesson in a problem particle physicists know well: the look-elsewhere effect. A peak can look impressive at one location while becoming much less surprising once you account for all the places a search could have found a peak. The local question is, “How unusual is this result here?” The global question is, “How often would a search this broad find something at least this unusual somewhere?” Those are different probability questions. In the 750 GeV case the gap between them was public at the time. Against its local figure approaching four sigma, ATLAS reported a global significance of roughly 2.3 sigma once the range of masses the search covered was taken into account, and the CMS global number was lower still.

That distinction matters more now because the searching is no longer at human scale.

Machine-learning systems increasingly sweep scientific archives for rare events, faint patterns, useful materials, unusual spectra, and unexpected correlations. An adaptive pipeline can change its next move after seeing its last result. It can alter a time window, choose a new frequency band, retrain a model, change preprocessing, generate a new hypothesis, and discard unpromising branches before a human ever sees the final candidate.


Adaptive search does not repeal classical multiple-testing theory. It makes the object that theory must account for much harder to see.

Bonferroni corrections, false-discovery-rate methods, trial factors, selective inference, and related tools remain valuable. But a correction needs a defensible account of the family or search process. In an adaptive AI pipeline, that family may be generated while the search is running. The effective number of opportunities for a favorable anomaly may be buried in the trajectory of the search rather than written in a methods section.

A pipeline that reports only its winning anomaly and a local p-value can therefore make a machine-scale search look like a single prespecified test. The failure lies in the reporting rather than in the searching, when a machine-scale search is presented afterward as though it had looked only once.

Science has confronted a related problem before. Genome-wide association studies test enormous numbers of genetic variants for association with disease and other traits. The field eventually converged on an unusually stringent convention, roughly P < 5 x 10^-8 for common-variant GWAS, precisely because a conventional 0.05 threshold would be absurd when a study searches across the genome. The family is not perfectly independent, but the multiplicity is visible enough to be priced.

AI-assisted discovery complicates that bargain. A machine may not start with one fixed list of a million hypotheses. It can create new branches from the data itself. Multiplicity, in that setting, should be measured rather than casually assumed.


Three practices would go a long way.

First, log the search. A scientific AI pipeline should preserve the realized search trace: which data slices it inspected, which preprocessing choices it tried, which models and hyperparameters it tested, which branches it abandoned, and which decisions caused new branches to be created. Search history is part of the evidence. The same four-sigma-looking anomaly means something different after ten opportunities than after ten million.

Second, when independent confirmation is unavailable, replay the full search on a defensible null. Run the same adaptive policy on signal-free, permuted, time-shifted, simulated, or otherwise scientifically justified null data and record the most extreme result produced by each complete search. That can empirically price the breadth of an adaptive procedure without pretending that every branch is independent or forcing a fictional count of “effective tests.” The null must reproduce the features of the real search that determine what can become reportable.

That last condition matters. Null replay is not magic. A beautifully computed global p-value is only as trustworthy as the null model and calibration behind it. If the noise distribution, dependence structure, preprocessing, or instrument state has drifted, a stale threshold can be confidently wrong.

Third, freeze the rule before the final evidence is inspected. If the discovery threshold, search policy, or calibration changes after the result disappoints, the change itself becomes another look. The statistical bar has to travel with the process that generated the claim.

There is also an important escape hatch. If a machine searches freely on development data, freezes one hypothesis, and that hypothesis is tested exactly once on genuinely untouched evidence, the exploratory search need not contaminate the final test. The problem returns when the same evidence helps select and confirm the winner, when the holdout is queried repeatedly, or when no independent confirmation exists.

That is why the solution is not to make scientific AI timid. Search aggressively. Let machines explore parameter spaces humans could never enumerate. Let them generate hypotheses and chase strange corners of data. But separate the freedom to explore from the authority to claim discovery.

Journals already ask increasingly sophisticated questions about reproducibility. Is the data available? Is the code available? Can the analysis be rerun? AI-assisted science suggests a natural next line in that checklist: search availability.

A discovery paper should disclose not only the winning analysis but the search capable of producing it. Show the search family or trace. State whether the final evidence was untouched. Report the multiplicity target. Identify the null calibration. Preserve failed branches when they were eligible to become the headline result.


Disclosure of that kind is what makes machine-scale exploration scientifically legible, and it is worth the paperwork it costs.

Five sigma was never a law of nature. It was a social technology: a convention built to protect extraordinary claims against the search practices of its era. The search practices are changing. The machines now look farther, faster, and more adaptively than any collaboration of humans could.

That is good news for discovery. But the evidence conventions must keep up. The machine looked everywhere. The paper should say so.



About Burak Oktenli

Burak Oktenli holds an MBA and a Master of Professional Studies in Applied Intelligence from Georgetown University. His research addresses the governance of authority in autonomous and AI-enabled systems, and his writing has appeared at the Modern War Institute at West Point, RUSI, RealClearDefense, RealClearMarkets, and Geopolitical Monitor. He is the author of Authority Architectures for Autonomous Systems, a ten-volume series on how authority in autonomous systems is delegated, monitored and recovered, at authority-architecture.me.
View all posts by Burak Oktenli →

No comments: