Tuesday, September 22, 2026

The Search For Alien Technology Has A Measurement Problem – Analysis


Image: ChatGPT


September 22, 2026
By Burak Oktenli

Key Takeaways:

Bigger AI models can rank, classify, and search at huge scale, but they cannot invent evidence an instrument never recorded. When two explanations produce the same data, no larger model can honestly tell them apart—the limit is identifiability, not compute.

Radio SETI, including Breakthrough Listen and the blc1 Proxima Centauri candidate later traced to terrestrial interference, shows the gap between finding a candidate and proving its origin. The next step is a new measurement—another site, antenna, geometry, epoch, or diagnostic—not another terabyte of the same ambiguous data.

Scientific AI should keep detection strength, calibration validity, tested alternatives, and independent follow-up separate instead of collapsing them into one confidence score. Automate ranking and triage; do not let a model promote its own score into a stronger claim without new evidence.


Bigger models can rank, classify, and search at extraordinary scale. They still cannot manufacture evidence that an instrument never captured.

The most dangerous promise in scientific artificial intelligence is not that machines will make mistakes. It is that they will make certainty out of missing information. As AI systems become better at finding patterns in enormous datasets, the temptation is to assume that a sufficiently powerful model can extract a correct answer from almost any measurement. That is false. Sometimes the obstacle is not computation, model capacity, or training data. Sometimes the decisive information was never measured in the first place.

Radio searches for extraterrestrial technology offer an unusually clear example. Modern programs such as Breakthrough Listen scan vast frequency ranges for narrowband, drifting signals that could be consistent with engineered transmitters. The field has also embraced machine learning. A deep-learning search of 820 nearby stars showed that learned representations can surface candidate signals that conventional filters may miss. That is exactly the sort of task at which AI can excel: ranking, triage, anomaly discovery, and the compression of an impossible search space into a manageable set of things worth examining.


But candidate discovery is not origin discovery. The distinction became vivid in the investigation of blc1, a narrowband signal detected in observations of Proxima Centauri. Its morphology was interesting enough to demand serious follow-up. The eventual analysis traced it to terrestrial radio-frequency interference. The lesson was not that unusual signals are unimportant. It was that a detection statistic is only the beginning of an evidentiary chain.

This point becomes sharper when we ask a simple question: what if two competing explanations generate exactly the same data for the algorithm? Suppose one hypothesis is a target-associated signal and another is a terrestrial interferer that, by chance or geometry, appears only during the same on-target scans. If both produce the same frequency-time track, the same amplitude pattern, and the same scan-by-scan observations, then the two hypotheses induce the same probability distribution over the data given to the model. In that case, no classifier can separate them honestly. Not a larger transformer. Not a deeper neural network. Not a quantum computer. The distinction is absent from the measurement.

That is an identifiability limit, and it changes how we should think about scientific AI. When two hypotheses are observationally equivalent, the next advance must come from experimental design rather than model design. We need a new measurement that causes the hypotheses to predict different outcomes: a simultaneous reference antenna, an independent observing site, a changed pointing geometry, an instrument diagnostic, a different epoch, or some other source of conditionally independent evidence. Another terabyte of the same ambiguous data does not solve the problem.


The same principle appears in more ordinary ways. Reference observations, for example, are often used to reject interference. That can work extremely well when the nuisance really is common to both target and reference scans. But the protection has a cost. If a genuine source leaks into an off-target observation through sidelobes or pointing geometry, a subtraction rule can subtract part of the signal we wanted to preserve, while a hard veto can discard it altogether. A method that looks safer on one axis may quietly become less sensitive on another.

Calibration creates a second trap. A threshold calibrated under one noise model does not retain its meaning after the residual environment changes. In a synthetic stress test I developed for radio technosignature decision rules, thresholds calibrated to an approximately one-percent cadence-level exceedance rate under independent Gaussian residuals produced trigger fractions above fifty percent after a simple time-correlation structure was introduced. Nothing about the threshold number itself warned that its interpretation had collapsed. The problem was not that the detector suddenly became stupid. The statistical conditions supporting the threshold no longer held.

This is why the most important question for scientific AI may not be, “How accurate is the model?” It may be, “Under what measurement conditions does this output still mean what we say it means?” That is a different discipline. It requires separating detection strength from calibration validity, separating a candidate ranking from an origin claim, and recording which alternative explanations have actually been tested.

Engineering already has language for this. NASA’s current standard for models and simulations treats credibility, validation, verification, uncertainty, and acceptance criteria as explicit parts of model use. Scientific AI needs the same instinct. A model should not inherit authority merely because it is sophisticated. Its output should carry the conditions under which it was calibrated, the data it actually observed, the alternatives it cannot distinguish, and the tests still required before escalation.


That discipline matters far beyond SETI. In medical imaging, two diseases can look similar under one modality and diverge only after another test. In climate science, a model can fit the historical record while remaining underdetermined about causal mechanisms. In intelligence analysis, multiple adversary explanations can remain consistent with the same observable behavior. In autonomous systems, a sensor-fusion stack can assign high confidence to a state estimate even when its sensors share a common-mode failure. In each case, more computation can sharpen the inference conditional on the evidence. It cannot supply the missing discriminator.

This should change how we evaluate AI progress in science. Benchmark culture often rewards the highest score on a fixed dataset. But the scientifically important question is often what happens when the assumptions behind that dataset fail. A robust benchmark should therefore include conditions designed to break the method: distribution shifts, correlated residuals, missing reference observations, leakage between supposedly independent channels, nuisance classes that mimic the positive class, and deliberately matched cases in which successful discrimination would reveal data leakage rather than intelligence.

The technosignature community already has many of the ingredients. Tools such as setigen support synthetic signal generation and injection. Published work has explored machine-learning direction-of-origin filtering, while the “cosmic haystack” formalism reminds us that excellent sensitivity within one narrow test family does not imply comprehensive coverage of the wider search space. The broader NASA technosignatures workshop report likewise framed technosignature research as a field in which new instruments, new surveys, new algorithms, and new theory all matter. The missing piece is not another declaration that AI will accelerate discovery. It is a stronger contract between what an algorithm outputs and what the measurement actually warrants.

Such a contract would be simple in principle. A candidate record should keep at least four things separate: the strength of the detected feature; whether the current data remain inside the calibration regime; which conventional or instrumental alternatives have been tested; and what independent observation, if any, supports a stronger interpretation. A fifth field can carry follow-up priority, but that is a resource-allocation decision, not a probability of origin. Collapsing all of these into one confidence score creates an audit problem: after the fact, nobody can tell which assumption carried the claim forward.

The same asymmetry should govern automation. AI can reasonably automate low-level transitions: ingest data, flag anomalies, compare a statistic with a frozen threshold, rank candidates, and request additional observations. But an automated system should not be allowed to convert its own score into a stronger scientific claim without new evidence. Promotion should require a named test or an independent measurement. Demotion, by contrast, should be easy. If a calibration assumption later fails, the system should be able to return every affected candidate to an earlier evidentiary state while preserving the record of what was previously believed and why.

This is not an argument against AI in science. It is an argument for using AI where it is strongest. Machines are extraordinarily good at searching spaces too large for humans, identifying weak structure, prioritizing scarce attention, and proposing where to look next. Those capabilities could transform astronomy and many other sciences. But scientific authority should attach to evidence, not to model scale.


The phrase “AI for discovery” therefore needs one amendment. AI can accelerate the path to discovery. It can expose patterns that deserve investigation. It can help design the next observation. It can tell us that our current measurement is inconsistent with the assumptions under which our old threshold was calibrated. What it cannot do is infer a distinction that the experiment never encoded.

That boundary is not a limitation to be embarrassed about. It is a design instruction. When the model cannot know, the answer is not always a bigger model. Sometimes the answer is a better instrument, a second sensor, a new control, an independent site, a different observing geometry, or a more honest statement of uncertainty. The future of scientific AI will depend as much on improving what we measure as on improving what we compute.

If the evidence is not in the measurement, intelligence cannot conjure it into existence.


About Burak Oktenli
Burak Oktenli holds an MBA and a Master of Professional Studies in Applied Intelligence from Georgetown University. His research addresses the governance of authority in autonomous and AI-enabled systems, and his writing has appeared at the Modern War Institute at West Point, RUSI, RealClearDefense, RealClearMarkets, and Geopolitical Monitor. He is the author of Authority Architectures for Autonomous Systems, a ten-volume series on how authority in autonomous systems is delegated, monitored and recovered, at authority-architecture.me.
View all posts by Burak Oktenli →


No comments: