3.1 Operational orchestration
3.1.1 Why this sounds like science fiction
This chapter introduces systems that observe a live operation, make bounded decisions, coordinate other agents, and act before a human has opened a conference call. That description sounds like science fiction, so it is worth starting somewhere real: electric utilities have been running systems with this operating pattern on their distribution grids for years, with one difference that will matter.
The function is called FLISR — fault location, isolation, and service restoration. Sensors and automated switches on distribution feeders detect a fault, the system isolates the faulted section, and customers on healthy sections are restored from alternate feeds, typically within thirty seconds to two minutes.1 The system acts on its own, but only inside an envelope that engineers designed in advance: it may operate specific switches on specific feeders, and it may do nothing else. It restores most customers automatically, and engineers review its response afterward. That sequence — detect, contain, review — is the operating pattern this chapter is about, and it already exists as a production discipline, without anyone calling it orchestration.
The difference is judgment. Utility engineers work out the switching plans in advance — for the faults they anticipate, which sections to isolate and which alternate feeds to restore from — and the system executes those plans deterministically. An agentic system is asked to do something harder: exercise judgment inside the envelope, for cases the designer did not enumerate, using retrieved evidence instead of a precomputed plan. That difference is why the vocabulary in this chapter is new even though the operating pattern is not. A system that executes pre-engineered responses needs a switching plan. A system that makes decisions inside a live operation needs a contract, an envelope, and a record of what it decided and why.
Many readers already use agentic systems at home, running OpenClaw, Hermes Agent, or something similar — watering the lawn, checking a camera, sorting messages, turning lights on and off, starting the laundry, opening the door for a delivery.2 That is real automation, often useful and sometimes surprisingly capable, and people are comfortable with it for two reasons that go together: the consequences stay local, and they are small. Nobody is using a home agent to automate a hospital. They are automating sprinklers, garage doors, laundry reminders, and the recurring buy order on one brokerage account. No customer commitment is involved, no regulated process, no contractual service level, and no formal audit trail.
There is a reason the stakes feel manageable, and it inverts at production scale. If someone hacks your sprinkler controller one night, that is not the end of the world. It might be the end of your lawn, but it is not the end of the world. When the same class of tool starts rerouting medical shipments or throttling power to a substation, the failure stops being an anecdote — and a whole new vocabulary of accountability arrives with it: who authorized the action, what the system was permitted to touch, who reviews the result, who can stop it. That vocabulary is what this chapter is building, because the moment the work has real impact, it is no longer optional.
The next step for this technology is to graduate from the household into things that are more meaningful, and parts of that graduation have already happened. Electric grids already detect faults and restore customers on their own, through the FLISR systems described a moment ago. Financial markets have run on automated trading for decades — program trading, computers set up to execute large baskets of stocks when conditions are met, goes back to the 1980s.3 Take those one-off approaches and generalize them and you have the next decade: more and more systems, in more industries, automated in similar ways, with inference engines where the pre-programmed logic used to be.
An agentic system supporting medical shipments, payment processing, warehouse fulfillment, customer identity, or a public-service application is making decisions inside an operating business. It may affect a customer who never learns an agent was involved. It may need to respond in seconds, preserve an audit trail, honor a regulatory commitment, protect sensitive data, and stop when the situation falls outside its tested conditions.
Meeting that requires more than a more forceful system prompt. The underlying problem is entirely practical: give software limited authority over a live system without pretending that limited authority is no authority.
Development and production can share the same models, tools, skills, and runtimes while carrying entirely different responsibility. So the question stops being whether the agent can produce a good answer and becomes what authority the system holds while the business is running, which decisions it can affect without a person, how quickly it must respond, and how the organization will explain the result afterward.
Call the first category agentic development, where delegated systems help create, test, review, and prepare software, and the second operational orchestration, where delegated systems observe and act within a live business environment. Agentic development is usually supervised at the artifact boundary. Operational orchestration has to be supervised at the decision and consequence boundary.
An operationally orchestrated system needs a few terms that development-focused discussions often leave implicit:
- Operational agency is the bounded authority to affect a live system, business process, or customer outcome.
- A permission envelope defines exactly what an agent may read, change, approve, communicate, spend, or trigger. AWS’s security team calls the same idea the security box — controls outside the agent, deterministic in enforcement, covering every interaction between the agent and the outside world.4
- An escalation boundary defines the conditions under which the system must stop acting and ask a human or a higher-authority system to decide.
- A decision record preserves the inputs, context, instructions, model or agent involved, available permissions, output, approval, action, and result. Design has a precedent: architecture decision records, which Michael Nygard named in 2011, document a decision and its rationale at build time; the decision record here is the run-time twin, one per consequential action.5
- An operational envelope defines the conditions in which an automated response is expected to be safe: the signals, thresholds, actions, rollback, and response time have all been specified in advance. Autonomous driving standardized this shape years ago: SAE J3016 calls the conditions a driving system is designed to function in its operational design domain, and calls the stable stopped condition it brings itself to when it must stop its minimal risk condition.6
The point of all that vocabulary is not to make a production system autonomous in the romantic sense but to make its authority bounded, inspectable, reversible, and accountable. A self-operating system is not a system without people. It is a system in which people have designed what can happen before an incident, what happens during one, and what has to return to human judgment at the edge of the known case.
Operational agency also has a lifecycle, because authority is not something an agent receives once and keeps forever. The production record should say who granted it, for what purpose, against which version of the contract, with what expiration, and under which conditions it must be reconfirmed or revoked. Identity systems already run this cycle for humans: Microsoft Entra and its peers schedule recurring access reviews — access recertification, in the vendor’s words — and the same cadence applies to an agent’s authority.7 For the routing system, that record might look like this:
| Field | Example |
|---|---|
| Authority owner | Order-service operations manager |
| Granted to | routing-agent-v3 using service account order-routing-proposer |
| Allowed action | Isolate affected order cohorts or stores and propose routing them to manual review |
| Not allowed | Change global routing, merge code, read payment data, or contact customers |
| Validity | During Sev-2 incidents while the contract and health checks are current |
| Reconfirmation | After a contract change, model change, schema change, or incident review |
| Revocation | Any failed evaluation, unexplained action, expired credential, or edge case outside the envelope |
| Human owner at the edge | Named order-service on-call engineer |
Parts of that record already have a protocol-level precedent. OAuth 2.0 Token Exchange, published by the IETF as RFC 8693, lets a middle-tier service trade a user’s token for a new one scoped to the next hop, and the standard defines a claim that carries, in its words, “a statement that one party is authorized to become the actor and act on behalf of another party.” The claim nests, so a service several hops downstream can see the whole chain of actors behind a request. What the protocol deliberately does not decide is when the chain should stop, who reviews it, or when it expires — those lifecycle questions belong to the arrangement, not the token.8
3.1.2 A note about why this may become regulated
Some readers will wonder why a production system needs this much bookkeeping. The answer is that the authority may not stay an internal engineering preference for long. A government or regulator may eventually require an organization to reaffirm on a schedule that an agent still has permission to make a particular class of decision, and to suspend or revoke that authority when the system misbehaves. A production system that can reroute medical shipments, control a power-grid response, approve financial transactions, or affect access to a public service may have to be licensed, extensively documented, and inspected by an independent investigator, with an audit trail showing what the system knew, which permissions were active, who confirmed them, what it decided, and whether anyone altered the record afterward.
That is not a universal requirement today, which is precisely why it is worth designing the records now. If the authority matters, the organization has to be able to show how it was granted, reaffirmed, narrowed, suspended, and restored — and that is not paperwork wrapped around the architecture but part of the architecture. When the model changes, the data schema changes, the agent gains a new tool, the business changes its tolerance for customer impact, or the system meets a case that was not in its evaluation set, the authority may need narrowing or reconfirmation. The contrast is easiest to see in production operations, so consider the same failure handled two ways.
3.1.3 The current system: the outage and the incident call
The current model at its worst is documented, not hypothetical. On August 14, 2003, a software bug stalled the alarm processor in the control room of FirstEnergy, the utility monitoring northern Ohio. The grid was already in trouble that afternoon — a generating unit had tripped, and transmission lines were sagging into trees — but the operators did not know it. The system that was supposed to alert them had failed silently, and nothing told them it had failed. For over an hour they ran a complex power system with no alarms, reasonably assuming that no news meant normal conditions, while a failure that should have stayed local cascaded across the regional grid and cut power to roughly fifty million people in eight states and Ontario — the largest blackout in North American history. The task force that investigated the cascade documented the sequence in detail, and its account of that hour is the clearest statement of the current model’s limit: the equipment failures were manageable, and what was missing was awareness that they were happening.9
The blackout is the extreme case. The same shape — a room full of people reconstructing what the system is doing while the failure spreads — shows up at ordinary scale wherever the response depends on a conference call, and the home-goods company from the earlier chapters had a retail-sized version of it. The order service fails. The failure begins in the Toronto fulfillment route, orders from the Hamilton and Mississauga stores start timing out, and nothing in the system immediately distinguishes whether the problem is the order API, the payment dependency, or the route between the warehouses and the carrier.
The on-call engineer opens a P1 incident, and the pages go out to the payment team, the order team, infrastructure, release engineering, and the warehouse operations lead. Hundreds of people eventually join the call, frustrated, because nobody can quickly answer the basic questions. Which stores are affected? Which orders are affected? Did the morning deployment cause this? Can the Toronto route be isolated without stopping Hamilton and Mississauga orders? Who is allowed to change the routing rule, and who knows the rollback procedure?
The outage lasts nineteen minutes, and it lasts that long because a human response has mechanical delays built into it: someone has to notice, someone has to page, people have to join a call, and the facts have to be assembled by hand from dashboards and memory while customers wait. None of that is a technology failure. It is what response looks like when the system cannot operate itself — the time goes to people instead of the work — and the era this chapter describes is the one in which that stops being the default.
3.1.4 The orchestrated system: detect, contain, review
The orchestrated version of the same failure does not start with a page. The operating software is already watching the order API’s error rate, queue latency, payment-dependency health, fulfillment-route status, and recent deployment records. A diagnostic agent recognizes the combination of signals that has preceded this failure before and identifies the Toronto route, the Hamilton and Mississauga stores, and the orders created after the routing change. Notice what it does not do: it does not call that group “an affected branch.” It names the stores, the route, and the order cohort the response will touch.
A containment agent can then route that cohort to manual review — and nothing outside the envelope in the lifecycle table above. A separate evaluator checks the evidence, a deterministic health check confirms that the proposed isolation matches the known failure conditions, and the system records the inputs, the agents, the permissions, the proposed action, the approval, and the rollback path.
There is still a follow-up meeting, where the operational orchestrators review the anomaly and ask whether the permission boundaries were appropriate, whether the response worked, whether the system learned anything valid, and whether the contract should be narrowed or reconfirmed.
The autopilot precedent. Flying an airliner through turbulence is the classic case where automation beats reaction, and aviation wrote the conclusion down years ago. Airbus’s guidance to crews in severe turbulence is to keep the autopilot ON: the autopilot is designed to cope with turbulence, holds the aircraft close to its intended flight path, and reacts faster than any pilot without the risk of overcorrection.10 FLISR is the same instinct on a power grid, and program trading is the same instinct in a market. The practice of using technology to react before things become disasters already exists. What is new is that the same kind of envelope is about to be attached to inference engines — so that almost everything can run on autopilot, under contracts like the ones this chapter describes.
Delegation is the most consequential design activity an orchestrator owns. Every line of the contract — what the system can read, what it may change, what it must return, when it must stop, who reviews it, and how the action is reversed — is a design decision that used to live inside the code and now lives inside the arrangement. The chapter’s job is to make those decisions inspectable.
Eaton, Feeder Automation Manager (FLISR), https://www.eaton.com/us/en-us/catalog/software/feeder-automation-software.html — automated fault location, isolation, and service restoration on distribution feeders, with restorations in approximately 30 seconds to 2 minutes; and Schweitzer Engineering Laboratories, “Fault Location, Isolation, and Service Restoration (FLISR),” https://selinc.com/solutions/p/flisr. Multiple vendors ship FLISR as a standard distribution-automation function; cited as evidence that bounded automated containment is an existing production discipline.↩︎
OpenClaw, “OpenClaw,” https://docs.openclaw.ai/; Nous Research, “Hermes Agent Documentation,” https://hermes-agent.nousresearch.com/docs; HKUDS, “nanobot,” https://github.com/HKUDS/nanobot. These are examples of agent runtimes with tools, memory, skills, and automation capabilities. The chapter uses domestic automation as a contrast in consequence and accountability, not as a claim that every home deployment is low-risk.↩︎
Federal Reserve Board, “A Brief History of the 1987 Stock Market Crash with a Discussion of the Federal Reserve Response,” Finance and Economics Discussion Series 2007-13, https://www.federalreserve.gov/pubs/feds/2007/200713/. The paper describes program trading as “computers set up to quickly trade particular amounts of a large number of stocks, such as those in a particular stock index, when certain conditions were met,” with portfolio insurance as the strategy most tied to the crash; the Brady Commission (1988) reached the same conclusion about the destabilizing interplay of index arbitrage and portfolio insurance.↩︎
Hart Rossman et al., “Four security principles for agentic AI systems,” AWS Security Blog, https://aws.amazon.com/blogs/security/four-security-principles-for-agentic-ai-systems: “We describe this as the security box. It’s external to the agent, deterministic in its enforcement, and comprehensive in its coverage. Every interaction between the agent and the outside world passes through it.” See also the Agentic AI Security Scoping Matrix, https://aws.amazon.com/ai/security/agentic-ai-scoping-matrix.↩︎
Michael Nygard, “Documenting Architecture Decisions,” November 2011, which named the architecture decision record; per Martin Fowler’s bliki, “Michael Nygard coined the term ‘Architecture Decision Record’ with an ADR-formatted article in 2011” (https://martinfowler.com/bliki/ArchitectureDecisionRecord.html). ADRs capture design decisions and their rationale; the run-time decision record this chapter describes is the per-action twin.↩︎
SAE International J3016, “Taxonomy and Definitions for Terms Related to Driving Automation Systems,” defines the operational design domain as the “operating conditions under which a given driving automation system or feature thereof is specifically designed to function, including, but not limited to, environmental, geographical, and time-of-day restrictions,” and the minimal risk condition as “a stable, stopped condition to which a user or an ADS may bring a vehicle after performing the DDT fallback in order to reduce the risk of a crash when a given trip cannot or should not be continued.” Quoted per NHTSA, “Automated Driving Systems: A Vision for Safety,” https://www.nhtsa.gov/sites/nhtsa.gov/files/documents/13069a-ads2.0_090617_v9a_tag.pdf.↩︎
Microsoft Entra ID Governance documentation: access packages “can require regular access reviews,” and access rights “can also be regularly reviewed using recurring Microsoft Entra access reviews for access recertification,” https://learn.microsoft.com/en-us/entra/id-governance/identity-governance-overview. Identity governance already schedules the reconfirmation cycle this chapter applies to delegated authority.↩︎
J. Bradley, M. Jones, B. Campbell, and H. Tschofenig, “OAuth 2.0 Token Exchange,” RFC 8693, Internet Engineering Task Force, https://datatracker.ietf.org/doc/html/rfc8693. The
may_actclaim is defined as making “a statement that one party is authorized to become the actor and act on behalf of another party”; nestedactclaims record the delegation chain. The RFC standardizes the token mechanics of delegated authority, not its governance.↩︎U.S.-Canada Power System Outage Task Force, “August 14th 2003 Blackout: Causes and Recommendations,” final report, https://www.ferc.gov/sites/default/files/2020-05/ch5.pdf. Chapter 5 documents the silent failure of the alarm processor in FirstEnergy’s GE/Harris XA/21 energy management system and the operators’ loss of situational awareness for over an hour. Approximately 50 million customers lost power across eight U.S. states and Ontario.↩︎
Airbus, “Managing Severe Turbulence,” Safety First, https://safetyfirst.airbus.com/managing-severe-turbulence: “Keep autopilot ON. Autopilot is designed to cope with turbulence and will keep the aircraft close to the intended flight path without the risk of overcorrection.”↩︎