6.5 Safeguards and credentials
6.5.1 The fallback plan
An orchestrator has to plan for the loss of AI, and it should be said plainly that this is not a way of giving residents a hard time on the road to a certification. It is a fundamental part of orchestration. The tools are genuinely amazing — that is this book’s premise — but there will be times when you need them the most and they are unavailable: a provider can go down, credentials can expire, a model can become unsafe or unreliable, a network connection can fail, a company can lose access to a vendor, and a regulator can prohibit a use. The system may need to keep operating while the AI layer is offline.
The goal, to be clear, is not that the organization operates at 100 percent productivity without AI. The goal is that it operates.
There is a limit to this, and it should be stated honestly. If AI becomes unavailable for a year, you are probably not going to be able to continue operating the same way. What you are preparing for is the single-day or multi-day outage that arrives during the most critical period of the year — the holiday rush, the tax deadline, the cold snap. It might seem silly to prepare for that, until the day it happens. And the requirement is part of something larger than any one system: as we move into this new world, organizations have to keep thinking about contingency and keep thinking about risk, because automating the routine work did not remove the risk. A fundamental understanding of the technical system has to be preserved, in people, so that the organization can demonstrate independence from any particular AI system.
So every resident should learn to write and test an AI-unavailable plan identifying:
- which workflows stop safely;
- which workflows fall back to deterministic software;
- which decisions move to a human queue;
- which records remain available without the model;
- who has authority to activate the fallback;
- how customers and staff are informed;
- how work is reconciled when AI returns; and
- how the organization practices the transition.
In practice a claims system might stop automated approvals, preserve intake, route every uncertain file to trained claims specialists, and keep payment processing running through the existing deterministic rules. A support system might stop automated replies and put cases in a human queue. A deployment system might stop autonomous merges while leaving engineers able to build, test, review, and release changes by hand.
NIST’s AI Risk Management Framework playbook explicitly names contingency options, recovery, incident response, and safe removal from operation as part of managing AI risk, and an AI-unavailable plan is the operational form of that requirement.9
The people who will disagree. There are popular developers who say the opposite, and Steve Yegge is one of the loudest. Yegge — the author of “Revenge of the Junior Developer” and the creator of Gas Town — has argued that the era of reading code is over: “Code is a liquid. You spray it through hoses. You don’t freaking look at it.”10 He may turn out to be right about the oceans of generated code no human will ever inspect. This book is not proposing that an orchestrator in training has to understand every line of what the system generates. The claim is narrower: there is a necessity to understanding how the thing operates — the architecture, the approach, the failure modes — such that on the morning a massive outage takes down a provider, or a hosted model comes under attack and goes dark, the person responsible can operate the system without it. As time goes forward that will be less likely to be needed. It will still matter, because that is what being responsible means.
6.5.2 The incentive problem
Organizations also have to be honest about an incentive problem sitting underneath all of this. If the system writes the routine code, the resident never sees the routine code. If the system investigates the incident, the resident never learns how an investigation works. And if the senior engineer uses an agent to finish every small task alone, that senior engineer has quietly removed the teaching moments that produced the next senior engineer.
The answer is not to make people do pointless work by hand. It is to make learning a required part of the system. A resident should sometimes write the code manually, sometimes inspect generated code, sometimes reproduce a failure without an agent, and sometimes operate the fallback — not out of nostalgia for typing, but for understanding. Measure whether residents can explain a change, find a boundary, diagnose a failure, and operate without the model, rather than measuring only how many lines the agent produced or how fast a ticket closed.
DORA’s 2025 research describes AI as an amplifier of an organization’s existing strengths and weaknesses rather than a substitute for the underlying system, and that is the right training principle.11 A strong engineering practice makes an agent more useful, while a weak one just makes the agent produce work faster than the organization can verify.
6.5.3 The system-specific qualification exam
The final qualification is closer to a bar exam than a course certificate. The candidate is not tested on reciting the features of an agent framework; they are tested on whether they can responsibly operate one particular system, using its real goals, code, data, permissions, evaluation set, telemetry, incidents, and operating procedures, across ordinary cases, edge cases, and questions designed to expose a shallow understanding.
A candidate must be able to:
- explain the system’s goal and the customer or business consequence;
- trace a decision through code, data, context, tools, and telemetry;
- describe the system in wide mode, across the whole workflow;
- inspect one component in deep mode and explain its local failure;
- answer a question from audit about what happened and what evidence exists;
- answer a question from legal about data use, authority, and responsibility;
- explain to security which credentials and permissions apply;
- explain to operations how to detect, pause, roll back, and recover the system;
- explain which improvements the system made and what evidence supports them;
- distinguish a bug fix from a new capability or changed permission;
- find a plausible but incorrect result in the evaluation set;
- identify an assumption the system failed to check;
- diagnose and modify a relevant piece of generated code;
- stop the system and activate the AI-unavailable procedure; and
- hand the system to another qualified person without relying on private knowledge.
Build trick cases into it. An API returns 200 while the business action never completed. A model produces valid JSON with the wrong customer identity. A new instruction improves the average result and makes one small but important customer group worse. A retry creates two refunds. A permission that appears read-only exposes data the system should never see. A model provider goes down during the busiest operating period of the year. In each case the candidate has to recognize why it is dangerous and choose the next action.
The senior orchestrator should sit the review alongside the relevant domain specialist, security, operations, audit, and legal or compliance owners wherever their authority applies. A candidate who passes a coding exercise and cannot answer an audit question is not qualified for a critical system, and neither is one who can explain the policy but cannot diagnose the implementation.
All of it is specific to a system and a scope; a bar exam is a generic license, and this is the layer that sits on top of one. Passing the exam for a customer-support workflow does not authorize anyone to operate a hospital logistics system, and qualification should be repeated after a material model change, a new tool, a new data source, expanded authority, a serious incident, or a long absence from the domain.
National Institute of Standards and Technology, AI RMF Playbook, “Manage,” including guidance on contingency options, incident response, recovery, and safe removal from operation. https://airc.nist.gov/airmf-resources/playbook/manage↩︎
Steve Yegge, in conversation with Tim O’Reilly, “Steve Yegge Wants You to Stop Looking at Your Code,” O’Reilly Radar, 2026, https://oreillyradar.substack.com/p/steve-yegge-wants-you-to-stop-looking; and “Revenge of the Junior Developer,” Sourcegraph, March 2025, https://sourcegraph.com/blog/revenge-of-the-junior-developer. The same conversation supplies the Formula One framing: “If you’re looking at your code, then you’re in a Formula One race and you’ve parked your car and opened the hood.”↩︎
DORA, “State of AI-assisted Software Development 2025.” https://dora.dev/dora-report-2025. DORA describes AI as an amplifier of existing organizational strengths and weaknesses and emphasizes underlying organizational capabilities.↩︎