Tuesday, September 15, 2026

The Trigger Is Now Curious: Physics Should Calibrate Its Taste – Analysis


Image: ChatGPT

September 15, 2026

By Burak Oktenli

Key Takeaways:

After 27 June the LHC entered Long Shutdown 3 toward the High-Luminosity machine (~June 2030). In August CMS reported two Level-1 unsupervised anomaly algorithms (AXOL1TL, CICADA) that in 2024 kept more than four billion events; about 740 million of those would not have been saved by any other Level-1 path.

A trigger already decides what physics can later exist. A false positive can still be studied; a miss is gone. Calibration, the author says, must cover what the model fails to keep, not only what it flags.

Three records are p
roposed: which model and threshold saved an event; tests of what the selector can miss; and why a shadow algorithm was allowed to change the archive. An anomaly score is not a discovery. Other fields that use AI to triage data face the same problem.

CMS is already using unsupervised machine learning to decide which collisions survive real-time data selection. Once a learned system helps shape the scientific archive, its blind spots and version history become part of the calibration problem.

On 27 June, CERN operators dumped the Large Hadron Collider’s final beams before the machine entered the four-year transformation that will turn it into the High-Luminosity LHC. The collider is quiet, but the decisions being made during Long Shutdown 3 will help determine what its experiments can preserve, reconstruct and eventually claim when beams return.

One development makes that question unusually concrete. In August, the CMS Collaboration reported the first sustained use of unsupervised anomaly detection in real-time LHC data selection. During 2024, two Level-1 algorithms, AXOL1TL and CICADA, selected more than four billion collision events for storage and analysis. Nearly 740 million of those events were not selected by any other Level-1 trigger algorithm and otherwise would not have entered the retained record.


That result moves the discussion beyond whether machine learning can run fast enough inside a trigger. It can. The harder scientific question is what calibration should mean when a learned selector helps decide which observations remain available for future physics.

The model does not decide what an event means. It acts earlier. It helps decide whether an event looks unusual enough to deserve scarce bandwidth before most collision data disappear. Once artificial intelligence sits that early in the measurement chain, calibration has to include not only what the model finds, but also what its selection process can fail to preserve.


A trigger is part of the measurement

The LHC produces collisions at a bunch-crossing rate of 40 MHz. CMS cannot store the full detector readout for every crossing. In the Phase-2 design, the upgraded Level-1 trigger is designed to accept up to about 750,000 events per second, while the High-Level Trigger reduces the stream to roughly 7,500 events per second for permanent storage.

Triggering has therefore always encoded scientific judgment. Physicists specify signatures worth retaining, including energetic muons and photons, jets, missing transverse momentum and more elaborate combinations. Every trigger menu answers an unavoidable practical question: what deserves bandwidth?

Learned anomaly detection changes how part of that judgment is produced. AXOL1TL is an unsupervised, signal-agnostic autoencoder operating in the CMS Level-1 Global Trigger. Trained on Zero Bias data, it assigns anomaly scores to incoming trigger objects in real time. Its production model entered operation in 2024 and was updated during 2024 and again for 2025 data taking. CICADA approaches the same general problem through calorimeter information. Together, the recent CMS results show that machine-learned selection is no longer a laboratory demonstration at the edge of collider physics. It is operating inside the machinery that allocates scientific attention.


The asymmetry of the event you never saved

The most important assurance problem follows from an asymmetry between false positives and false negatives.

A false positive is comparatively forgiving. The trigger retains an event that later proves ordinary. Physicists can reconstruct it, inspect it, compare it with backgrounds and decide that the anomaly was uninteresting.

A false negative is different. If a physically important collision receives an ordinary score and is rejected before full readout, the complete event cannot later be reconstructed as though the detector had stored it. What was missed has to be studied indirectly through diagnostic channels that were preserved by design.


This is the scientific difference between a model that recommends what to examine and a model that helps determine what can still be examined. The second system shapes the future evidence base before anyone knows which event will matter.

CMS already has important pieces of the answer. Zero Bias samples support training and evaluation. Firmware behavior can be compared with software emulation. Test-crate operation gives new algorithms a shadow environment before they influence ordinary readout. For Phase 2, Level-1 Data Scouting is being developed to capture trigger information at the full 40 MHz collision rate, creating a high-statistics stream for monitoring, diagnostics and selected analyses. The opportunity during Long Shutdown 3 is to treat these practices as a coherent calibration discipline for learned selection.
Three records a learned trigger should leave behind

First, preserve selection provenance. For an event retained through a learned anomaly path, the collaboration should be able to recover which model and firmware version acted, which threshold applied, the relevant detector and calibration state, and the score that placed the event on that path. This does not require attaching a large payload to every event. Durable references to versioned artifacts are enough if they remain recoverable. A trigger decision should be reproducible as part of the measurement history.

Second, preserve evidence about the misses. An anomaly detector can become highly sensitive to the wrong kind of strange. Detector noise, calibration transitions, unusual multiplicities or other artifacts may attract high scores while a physically important signature remains comparatively ordinary in the model’s representation. Validation should therefore test detector pathologies as well as experiment-specific signal injections, and blind benchmarks can help prevent teams from tuning the selector to the very anomalies used to judge it.

That concern is already visible in CMS research. Recent work on debiasing ultrafast anomaly detection quantified how model-selection procedures can bias an anomaly detector toward simulated benchmark signals and explored an alternative criterion based on agreement across learned representations. The lesson reaches beyond any one method: validation does not merely measure what the trigger recognizes. The validation design can influence what the trigger learns to treat as interesting.

Third, preserve the evidence behind changes in role. Whenever a learned selector moves from shadow operation into a path that affects retained data, or when its threshold, model or firmware changes materially, the evidence supporting that change should remain visible. Stability across runs, agreement between firmware and software emulation, sensitivity to pileup and detector faults, benchmark coverage and reproducibility across updates define the conditions under which the selector is allowed to shape the scientific record.

An anomaly score does not make a discovery

The conceptual boundary is simple but important. An anomaly score can earn an event a place in the retained stream. It cannot by itself give the event evidentiary weight as new physics.

A surprising event still has to survive detector reconstruction, calibration, background modeling, statistical controls, independent analyses and the judgment of the collaboration. The learned trigger preserves possibilities; the rest of the scientific process determines whether any of them deserve belief.

I argued recently in Eurasia Review’s Science section that scientific AI needs an explicit way to preserve uncertainty and surface observations that do not fit existing categories. Learned triggering is a particularly demanding version of that problem because the system acts before the full scientific record exists.

The assurance goal should therefore be inspectability. What did the selector find unusual? Which model and threshold were active? What detector state surrounded the decision? Which classes of events did validation show the model could miss? A useful learned trigger should leave enough evidence for later scientists to interrogate the selection that shaped their dataset.

Long Shutdown 3 is the design window

The timing is favorable. CERN’s updated accelerator schedule places the start of High-Luminosity LHC operations around June 2030. CMS is using the intervening shutdown to rebuild major parts of its trigger and data-acquisition architecture for higher rates, richer Level-1 information and pileup approaching 200 proton-proton interactions per bunch crossing.

Phase 2 will bring tracking information into the Level-1 decision and expand the diagnostic possibilities available through scouting. Calibration practices designed into that architecture now can become ordinary instrumentation rather than an after-the-fact response to a model problem.

The lesson extends beyond CERN. Astronomy, genomics, materials science and other data-intensive fields are moving machine learning upstream from analysis toward triage, prioritization and selection. The more an algorithm determines which observations survive for later scrutiny, the more its provenance, blind spots and change history become properties of the measurement system itself.

Particle physics is unusually well placed to establish that norm. Calibration, control samples, blind analyses, versioned software and adversarial scrutiny are already part of its experimental culture. Learned selection can be absorbed into the same discipline as another component capable of shaping what an experiment will later be able to know.

Calibrate the curiosity

Particle physics has always advanced by choosing what to ignore. The trigger makes that choice at extraordinary speed because the experiment has no alternative.

Machine learning makes part of that choice adaptive. Once a learned selector helps determine what survives long enough to become evidence, its validation envelope, version history and diagnostic channels belong inside the experiment’s calibration story.

The trigger has become curious. Physics should be able to show how that curiosity was calibrated.



About Burak Oktenli
Burak Oktenli holds an MBA and a Master of Professional Studies in Applied Intelligence from Georgetown University. His research addresses the governance of authority in autonomous and AI-enabled systems, and his writing has appeared at the Modern War Institute at West Point, RUSI, RealClearDefense, RealClearMarkets, and Geopolitical Monitor. He is the author of Authority Architectures for Autonomous Systems, a ten-volume series on how authority in autonomous systems is delegated, monitored and recovered, at authority-architecture.me.
View all posts by Burak Oktenli →

No comments: