3.5 Design review and the profession
3.5.1 The orchestration design review
Before a consequential topology goes live, review the design itself rather than the outputs. Ask what the topology assumes about the world, which contract can change that assumption, where authority enters, and how a person will know when a result has landed outside the expected envelope. The review should end with a deliberately boring statement — what may this topology do without asking, what must it ask about, and what may it never do? — because boring boundaries are the sign that an architecture has become operational rather than aspirational. The model may generate the plan; the orchestrator decides whether the plan deserves authority.
Three questions the review has to answer before approval:
- What is the smallest change to inputs, models, or data that would silently produce a wrong answer with high confidence? If the team cannot name one, the evaluation set is not testing the dangerous failure mode.
- Which contract, if compromised or simply mistaken, would cause the most consequential downstream action? That contract gets the tightest authority and the strongest independent check.
- What does the topology look like when its orchestrator is unreachable for a week? If the answer is “it waits,” the design is sound. If the answer is “it accumulates,” the design has built a hidden dependency on a person.
Consider a medical-shipment system. An agent may be allowed to reroute a routine shipment when a warehouse loses power, but that permission should not automatically cover temperature-sensitive medication, controlled substances, pediatric vaccines, or a shipment whose delivery address conflicts with the prescription record. Those cases cross an escalation boundary. The system must preserve the shipment record, identify the affected patients or facilities without exposing unnecessary personal data, and ask a qualified human to decide. Speed matters, but speed does not erase the obligation to know what is being moved, for whom, under which conditions, and with what evidence.
3.5.2 The vocabulary will become ordinary
The terms sit close together and almost nobody uses them consistently yet. People talk about agents helping developers, about automation, workflows, copilots, and AI operations, and none of that adds up to a settled everyday vocabulary for a set of delegated systems operating a production process on behalf of a company — even as the stack described earlier makes that vocabulary unavoidable.
The vocabulary will become necessary the moment these systems start affecting customers and critical services, because someone will ask: what was the decision envelope for that system? What was it allowed to decide without a human? Which stores, shipments, accounts, or patients could its decision affect? How did it relate to the other systems around it? Was the authority properly delegated, and who confirmed it was still appropriate after the model, the data, the contract, or the operating conditions changed? Who was watching the audit logs for anomalies? What happened when the system met a case it had never seen before?
Those are not philosophical questions. They are the questions an incident investigator, an operations manager, a customer, a regulator, or a court asks after an automated decision has caused harm or disrupted a business.
Current AI governance frameworks already point in this direction. The European Union’s AI Act requires human oversight for high-risk AI systems and says that people assigned to that oversight must have the necessary competence, training, and authority. It also includes requirements related to logging, record-keeping, monitoring, and traceability.53 NIST’s AI Risk Management Framework similarly treats governance, accountability, human oversight, monitoring, and documentation as organizational responsibilities rather than features a model can provide by itself.54
3.5.3 Where the profession may harden
Neither framework creates a universal professional license called “certified orchestrator,” and this book is not forecasting one. What the frameworks do establish is precedent: regulators already impose competence, training, and oversight requirements on the humans who operate high-risk automated systems. Once a person is responsible for configuring the agents that operate a power grid, a medical network, or a financial market, those same questions — trained for it, authorized for it, reviewable in the role — arrive whether or not anyone issues a license. The honest form of the claim is institutional, not professional: formal qualification, periodic reconfirmation, and independent review are arriving as organizational requirements first, and licensing, if it comes, will follow the liability rather than precede it.
The cost of getting orchestration wrong is already visible, and it affects employment and compensation. Klarna’s sequence — an assistant announced at 700 agents’ equivalent, then a chief executive conceding fifteen months later that cost had been “a too predominant evaluation factor” and rehiring so a customer could always reach a human — is an orchestration failure rather than an AI failure: the design had not decided when volume stops being enough. It is the kind of failure that matters more as these systems take on work closer to money, health, and safety.55
None of which means a human has to personally approve every routine action, since that would defeat the purpose of building systems designed to respond faster than a meeting can be assembled. The requirement is that humans remain responsible for the decision boundaries — what the system may do, what it must explain, when it must stop, who can override it, and how the organization proves those boundaries were maintained.
3.5.4 Operational responsibilities expand the role beyond development
All of which is why the orchestrator’s work is broader than software development. A developer writes the service that gets deployed; an orchestrator defines the operating arrangement around that service — the agents, tools, memory, permissions, handoffs, evaluations, logs, escalation paths, and human responsibilities that determine what happens after deployment.
A well-designed system may operate itself for long stretches. It may detect a failing dependency, route around it, and return to normal before any customer notices. It may keep a medical shipment moving while refusing to guess about a temperature excursion, or flag an unusual payment pattern and hold the transaction for review rather than approving or rejecting it blindly. When the situation crosses the decision envelope, though, the system has to know how to stop, and the organization has to know who is accountable for what happens next. [^ch07_oasp]: NVIDIA, “NVIDIA Launches Open Agent Safety Platform to Secure Agents From Testing to Deployment,” September 28, 2026, https://nvidianews.nvidia.com/news/open-agent-safety-platform — OpenShell verified-policy runtime plus Sentry enforcement on BlueField-4 DPUs; 100+ partners including Anthropic, Scale AI, Salesforce, SAP, Citi, and JPMorganChase. The assessment — credential-half progress, governance layer still absent — is this book’s; NVIDIA describes the platform as full-stack governance, a claim this chapter reads narrowly because delegation records and cross-organization authority transfer are outside what the launch documents. Verified September 29, 2026.
HQ 6 — Assembled. The human and AI each wrote portions of this chapter. I assembled, reviewed, and take responsibility for the whole; the voice and arguments are mine, and I know which parts are which.
Regulation (EU) 2024/1689, Articles 12, 14, and 26, EUR-Lex, https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng. The regulation includes requirements concerning record-keeping and logs, human oversight, and deployer responsibilities; Article 14 requires high-risk systems to support effective human oversight, while Article 26 requires deployers to assign oversight to people with the necessary competence, training, and authority. This chapter does not claim that the regulation creates an “orchestrator” license.↩︎
National Institute of Standards and Technology, “AI Risk Management Framework,” https://www.nist.gov/itl/ai-risk-management-framework, and “AI RMF Playbook,” https://airc.nist.gov/docs/AI_RMF_Playbook.pdf. NIST presents the framework as voluntary guidance and describes governance, accountability, human oversight, monitoring, documentation, and role definition as organizational practices. It is not a licensing regime.↩︎
Klarna’s February 2024 release called its assistant’s output equivalent to 700 full-time agents; that was company-reported task capacity, not headcount eliminated. In May 2025 CEO Sebastian Siemiatkowski told Bloomberg that cost had been overemphasized and announced a pilot ensuring human access. Cited for escalation and authority design, not as proof the deployment failed.↩︎
- 3.1 Operational orchestration
- 3.2 Contracts and grants
- 3.3 Topology
- 3.4 The governance layer that does not exist yet
- 3.5 Design review and the profession