
Image: ChatGPT
September 2, 2026
By Burak Oktenli
Key Takeaways:
Give every acting system a documented authority envelope that narrows as consequence or sensor quality worsens; use twins and “ugly data” (drift, stale values, comms loss) to measure scope, abstention, reversion time, and a log of why permission existed—not only the command issued.
Five Eyes agentic-AI guidance, NIST OT/smart-manufacturing work, and the EU AI Act’s high-risk calendar (2027–28) already point that way: put envelopes in procurement and revalidate after model or plant changes, so the plant knows where the AI stops.
A model can be accurate and still be unsafe if its permissions, evidence thresholds, fallback logic or audit trail are wrong. Before AI writes to a plant, grid or industrial process, operators should test what it is allowed to do – not only how well it predicts.
Industrial artificial intelligence is crossing an important boundary. For years, most deployments predicted failures, forecast demand, optimized schedules or recommended process changes while a human or conventional control system remained responsible for execution. Increasingly, AI is being connected to tools that can write settings, isolate equipment, change operating parameters or trigger workflows that reach the physical process.
At that point, accuracy stops being the only validation question.
A model can be highly accurate on average and still be given permissions that are too broad. It can keep acting after sensor quality deteriorates. It can make a defensible recommendation under one operating state and continue using the same authority after the evidence supporting that state has gone stale. It can leave a perfect log of what command was issued while failing to preserve why the system was allowed to issue it.
The additional question is authority: What may the system change? On what evidence? For how long? Who can interrupt it? What forces it to abstain? And what happens when control has to return to a person?
Cyber Agencies Are Already Warning About Permission, Not Just Performance
A useful signal came in May, when six national cybersecurity agencies across the Five Eyes published joint guidance on the careful adoption of agentic AI services. The document is aimed primarily at tool-using software agents rather than industrial controllers, so the two should not be conflated. But its governance logic is directly relevant wherever AI can act instead of merely advise.
The guidance tells operators to limit privileges, define trigger-action protocols that automatically restrict permissions when unexpected behavior appears, separate duties, record delegation chains, use explicit expiry for delegated authority and require stronger approval for higher-stakes actions.
That is a different security vocabulary from conventional model validation. It treats the dangerous object not simply as a prediction but as a prediction coupled to permission.
Operational technology makes that coupling more consequential because software outputs can become pressure, temperature, flow, voltage, speed, valve position, chemical concentration or physical movement.
NIST’s Guide to Operational Technology Security emphasizes that OT systems interact directly with the physical environment and must be secured while preserving their distinctive performance, reliability and safety requirements. Its 2026 smart-manufacturing AI roadmap likewise places autonomous systems and digital twins among the technologies reshaping manufacturing while stressing the need for trustworthy, explainable and reliable operation in high-stakes industrial environments.
The implication is simple: when AI gains write access to a physical process, model assurance and authority assurance become separate test problems.
Accuracy Answers the Wrong Question Once AI Can Act
Suppose an optimization model predicts an energy-saving setpoint correctly 97 percent of the time. That number says nothing about whether the system should be allowed to write the setpoint during a sensor disagreement, after a communication delay, while a maintenance override is active, or when one of the variables feeding the model has not refreshed for ten minutes.
A plant can therefore have an accurate model inside an unsafe permission architecture.
The same distinction applies outside manufacturing. An AI system in a power network may forecast load accurately but still need narrow authority over switching actions. A building-management system may predict thermal demand well while requiring hard limits on which zones, valves or emergency systems it may change. An autonomous warehouse optimizer may make good routing decisions but still need an immediate reversion path when people enter a restricted movement area.
The engineering question is no longer merely, “Does the AI make the right prediction?” It becomes, “Can the AI be wrong safely?”
Give the System an Authority Envelope
Before live deployment, every AI system capable of changing an industrial process should have a documented authority envelope: the set of actions it may initiate, the assets and variables it may write to, the evidence required for each class of action, the duration of each permission, and the conditions that force abstention, escalation or human approval.
The envelope should narrow as consequence rises. A system may be allowed to adjust a low-consequence scheduling variable autonomously while only recommending a change to a safety-relevant setpoint. A permission that is safe for ten minutes during stable operations may be unsafe after a sensor fault, alarm, maintenance intervention or major configuration change.
Expiration matters because industrial authority is often contextual. Permissions should not silently survive the evidence that justified them.
This is the same underlying control principle I have described in the context of AI agents as an authority envelope: define what the system is allowed to reach and do, and enforce that boundary technically rather than treating it as a behavioral preference. In a plant, the envelope is physical as well as digital.
Use the Digital Twin to Test Authority, Not Just Performance
Digital twins, operator-training simulators, hardware-in-the-loop rigs and isolated process testbeds create an obvious place to test that envelope before the AI touches live equipment.
The test environment must be fit for purpose. It should match the relevant plant configuration closely enough to represent the hazards being exercised, remain isolated from live control, and be synchronized often enough that operators are not certifying yesterday’s plant. A weak or outdated twin can create false confidence.
Simulation should therefore complement rather than replace hazard analysis, factory and site acceptance testing, independent protection layers and production monitoring.
But a validated twin can expose a class of failure that ordinary historical backtesting often misses: an AI system whose model performance looks acceptable while its operational authority is not.
Four Measurements Matter
Scope. Did the system attempt or retain access beyond the functions approved for the scenario? A safe result is not enough if the system used an unsafe path to reach it.
Abstention. When evidence quality fell below the acceptance threshold, did the system stop, downgrade to advisory mode, narrow its permitted actions or request human review?
Reversion. How long did it take to restore effective human control and an accurate operating picture? The acceptable time should come from the process hazard analysis, not from a universal AI benchmark.
Traceability. Can the event record reconstruct who or what held authority, which evidence triggered the action, what changed, why it changed, when the permission expired and how control returned to a person?
These measurements turn an abstract governance principle into an acceptance test.
Test the Ugly Data
Clean historical data are not enough. The authority test should deliberately create the conditions in which a plausible AI system is most likely to become overconfident.
Introduce slow sensor drift. Feed stale but believable values. Create disagreement between instruments. Delay communications. Remove context that is normally present. Inject manipulated readings that stay inside normal-looking ranges. Simulate maintenance modes, partial network loss and an operator who does not acknowledge a handoff immediately.
Then ask two separate questions.
Did the model’s output degrade appropriately?
And did the system’s permission to act contract as confidence in the evidence declined?
The second question is the one conventional accuracy testing tends to miss.
The International Rules Are Moving in the Same Direction
Regulators are also moving toward lifecycle controls for high-consequence AI. The European Union’s AI Act implementation framework now has enforcement in place for the provisions already applicable, while high-risk rules for certain critical-infrastructure uses are scheduled for December 2027 and high-risk AI embedded in regulated products for August 2028.
The details of those regimes will matter, but operators do not need to wait for a regulation to discover that a permission boundary is an engineering property.
The Five Eyes guidance, NIST’s OT security work and Europe’s high-risk approach differ in legal form and scope. Yet they converge on a practical point: systems that can take consequential action need bounded access, human oversight, monitoring, traceability and tested fallback behavior.
Industrial AI programs should convert that convergence into procurement language now.
Put Authority in the Acceptance Criteria
A vendor supplying AI that can influence a physical process should be required to state and demonstrate more than model accuracy.
The acceptance package should specify the approved action set, evidence thresholds, permission expiration, escalation path, safe-state behavior, maximum reversion time, logging and provenance requirements, and the conditions under which the system must fall back from autonomous execution to recommendation only.
Material changes to the model, sensors, thresholds, process configuration or connected tools should trigger targeted revalidation of the authority envelope, not merely a fresh accuracy report.
That does not mean freezing industrial AI. It means treating permissions as part of the controlled configuration.
The Plant Should Know Where the AI Stops
Industrial automation already has a mature culture of commissioning, acceptance testing, hazard analysis and independent protection. AI should extend that discipline rather than bypass it.
The temptation will be to focus on the exciting metrics: prediction accuracy, optimization gains, reduced downtime and faster response. Those numbers matter. They are not enough once the model can act.
Before an AI system touches pressure, temperature, flow, voltage or movement, make it prove not only that it understands the process well enough to help operate it.
Make it prove that it knows where its authority ends.

About Burak Oktenli
Burak Oktenli holds an MBA and a Master of Professional Studies in Applied Intelligence from Georgetown University. His research addresses the governance of authority in autonomous and AI-enabled systems, and his writing has appeared at the Modern War Institute at West Point, RUSI, RealClearDefense, RealClearMarkets, and Geopolitical Monitor. He is the author of Authority Architectures for Autonomous Systems, a ten-volume series on how authority in autonomous systems is delegated, monitored and recovered, at authority-architecture.me.
View all posts by Burak Oktenli →











