5.4 What Maya is doing
Maya is the running example this book follows; the earlier chapters showed her recognition and her conflict. What they did not show is the part of the work that never makes the demo: maintaining the system itself.
Maya is not a developer with better tools. She writes skills as SKILL.md files, freezes mishandled cases into an evaluation set as a LangSmith dataset or an Inspect-style task, adds a critique agent that is not permitted to see raw PII, wires MCP tools with revocable scopes, and curates memory in MEMORY.md and Letta blocks about what has worked and what has not. She may well write code. She also spends real time comparing outputs, rewriting an evaluation rule, changing a handoff, and narrowing a permission in an external policy file. As the system matures, she encourages skills that have stabilized to migrate into tools — a script, a function — so the same procedure runs without interpretation variance and can be tested under fixed conditions.
Here is what that means for the org chart. To a manager, or to anyone outside this transition, the work of an orchestrator looks almost indistinguishable from the work of a programmer right now. If you walk past a room of developers, it may not be obvious that one or two of them have made a complete shift to a new way of working. But they are not building websites and they are not building components. They are building the systems that themselves construct those outputs. The gap between building components and building the systems that construct them is the whole shift, and it is wide.
5.4.1 The same idea in operations: responding to an incident
Everything so far in this chapter has been about development — building and shipping a feature. Orchestration also changes operations: what happens when production breaks. Same website, 2:14 in the morning, and the checkout error rate spikes.
The old path pages an on-call engineer, who opens six dashboards and pages a second person. The second engineer reads the alerts, takes another five minutes to figure out which team needs to be paged first, then pages the specialist on duty. Fifteen or twenty minutes later, somebody is finally on the phone who actually understands the part of the system that failed. That is a typical diagnosis timeline for a service the company does not have a watching system for — and most companies do not have a watching system yet. The delay is not slowness on anybody’s part; it is the cost of relying on people to reconstruct what a system that was paying attention the whole time would have known.
Companies have improved reliability, but incident response is still far more manual than the tooling suggests. Uptime Institute’s 2026 survey found fewer impactful outages, yet one in ten outages was still serious or severe and the cost of failures continued to rise.10 A 2024 PagerDuty-commissioned survey of 500 large-company IT leaders found that more than 70 percent had not fully automated remediation, responder mobilization, cross-team collaboration, or internal communication; the average reported incident took 175 minutes to resolve.11 The opportunity is not to remove judgment from incident response. It is to stop making people rediscover the same context and coordination path while the clock is running.
The orchestrated path looks different. The system is already watching production. It sees the error spike within a couple of minutes, and it is configured to automatically start checking whether it can identify a cause: one agent reads the checkout service’s recent traces, one checks the payment provider’s status page, one compares the code in the last two deploys with the errors that just started, and one drafts a customer status message the system will not post until a human approves it. Within five minutes it has usually found the root cause and drafted a suggested patch. It is still configured to stop and wait for a human before taking any action — a person decides whether to roll back — but the diagnosis arrives with the evidence attached, instead of six dashboards and two levels of paging the wrong team to figure out who even knows what is going on.
None of that depends on technology that does not exist yet. The monitoring that detects the spike, the agents that investigate, the approval gate that stops them before acting — every piece can be built with tools available today. What usually does not exist is the contract stating how a human owner can authorize a rollback action through the system, with a recorded decision and a clear escalation path when something goes wrong.
Uptime Institute, 2026 Global Data Center Survey, https://uptimeinstitute.com/about-ui/press-releases/16th-annual-2026-global-data-center-survey-deployment-of-high-density-racks-rising-fast-operators-face-continued-recruiting-and-retention-pressures. Fewer impactful outages than prior years, but one in ten outages serious or severe, with the cost of failures continuing to rise.↩︎
PagerDuty, 2024 incident-cost study, https://www.pagerduty.com/newsroom/study-cost-of-incidents/. Commissioned survey of 500 large-company IT leaders: more than 70 percent had not fully automated remediation, responder mobilization, cross-team collaboration, or internal communication; average reported incident resolution time 175 minutes. Vendor-sponsored research; the cost figures are reported, not independently audited, and are not evidence that automation alone caused better outcomes.↩︎