5.3 Before and after: two ways to change a website

Book 1 · Your Next Job TitleChapter 5 · section 3 of 7

The difference between a developer and an orchestrator is entirely practical, and it becomes visible the moment you hand both of them the same problem.

So take two people responsible for the same website, and a product manager who asks each of them to add a saved-search feature. Same requirement, same codebase, same deadline. The difference is not in what they have been asked to do; it is in how they approach it — what they think about first, what they delegate, what they measure, and what they understand the job to be.

5.3.1 Developer A: direct implementation

Developer A uses an AI coding agent to draft the code, reviews it, writes tests, and opens a pull request. With today’s tools that takes two or three days, which is not slow at all — it is genuinely impressive compared with where we were five years ago. The times are illustrative.

Look at what Developer A is actually doing across those two days, though. Developer A holds the architecture in their head: how the search index is built, which services read it, what the migration has to avoid breaking. Developer A prompts the agent, reads what comes back, catches the hallucinated API call, asks for a correction, and reads it again. Developer A thinks through the security question — can one customer see another customer’s saved searches — and writes the test that would catch it if the answer changed. Then Developer A walks the change through design review, answers security’s questions, fixes what QA finds, babysits the deploy, and watches the dashboards afterward. The agent writes code far faster than Developer A ever could. Everything else — the plan, the checks, the handoffs, the judgment — is Developer A, in sequence, by hand.

5.3.2 Orchestrator B: delegated implementation

Orchestrator B is responsible for the same website and has a system. Orchestrator B did most of the work earlier. Designing it meant deciding which agents exist, what each is responsible for, what context each receives, what tools and permissions each holds, how they hand off, what happens when they disagree, and what requires a human decision. Then came preparing it: writing the skills, wiring the tools, giving it memory of past cases, and testing that it behaves inside its permissions. That preparation is the orchestration, and it takes real time — Orchestrator B may have spent as long building the system as this one feature would have taken the ordinary way. What Orchestrator B buys with that investment is not one fast feature. It is the capability to move fast on every change from then on, because the checking and testing that Developer A does by hand are already built into the system.

When the request arrives, it does not necessarily go through Orchestrator B at all. One of the biggest differences in the system Orchestrator B has built is that Orchestrator B has told product management to go ahead and ask the system to implement a feature. Orchestrator B is still monitoring what the system is being asked to do, and in some cases the system is configured to come back and ask Orchestrator B a question when it needs more guidance.

The system runs multiple paths at once. One agent inspects the repositories, API contracts, and recent support cases while a planning agent compares implementation approaches and picks one. A coding agent implements the change, a test agent writes regression cases from the requirements and from prior failures, a security agent checks data access and permissions, and an adversarial test agent writes tests specifically designed to break the implementation and find out whether the strategy was actually right. A deployment agent prepares a feature-flagged release. All of that happens in parallel rather than in sequence, and none of it needs Orchestrator B.

What makes them a system rather than a collection is that they are connected on defined terms. The planning agent hands its approach to the coding agent, the coding agent hands a diff to the security and test agents, the test agent returns failures to the implementation agent, and the deployment system cannot promote anything until the checks pass.

Out the other end comes the evidence: the plan, the diff, the test results, the adversarial results, the security review, the permission changes, the monitoring plan, and the rollback path. Along the way the system flags a proposal that would expose another customer’s saved searches, because Orchestrator B configured a security agent to catch exactly that. Then Orchestrator B comes back, reviews the evidence package, and approves or rejects. Orchestrator B’s involvement in the change is the preparation beforehand and the decision at the end.

This is the level of capability that can be assembled from the pieces that exist today. A bounded change that costs Developer A two or three days of sequential work can be completed in thirty or forty minutes by an orchestrated system running agents in parallel, testing its own work, and delivering the evidence along with the change. The difference is that the orchestrated system can act autonomously, it can parallelize, and in some cases you can have systems with dynamic topology that are created as needed to scale and meet demand. Counted honestly — preparation plus the run — the first change may cost about what it would have cost the ordinary way. The return shows up on the second change, and the fiftieth, when all of that preparation is already in place.

The two models look like this side by side:

Direct implementation Delegated implementation
Primary implementer Developer A Multiple connected agents and tools
Human’s main work Understand the architecture and construct the change Define the objective, configure the agents, set authority, evaluate evidence, and decide when to stop or release
Typical output Code, tests, pull request, release A proposed change plus style review, content review, tests, security review, evidence of who decided what, monitoring, and rollback evidence
Human technical role Decide how to implement the requirement Decide whether the interrelated systems understood and safely implemented the requirement
What the person operates The website and its development process The website plus the systems that develop, review, release, and monitor it
Typical path for a bounded change Implementation followed by several human handoffs Parallel agents followed by connected checks and one accountable decision

5.3.3 The pieces of this pattern already run in production

Orchestrator B is an illustration, not one named engineer. The pieces are real, and they run in production; the complete loop is still assembled by hand. The frameworks named earlier in this chapter — Gas Town, Microsoft Agent Framework, Google ADK, Strands, MetaGPT, CrewAI, LangGraph — are how teams are wiring it up today: a planner that breaks work into tasks, workers that run in parallel, monitors that replace stuck ones, merge discipline between parallel changes, evidence handed back to a person who decides.

5.3.4 What Orchestrator B’s stack looks like in concrete systems

The table below is not a vendor endorsement. It maps each job in B’s run to technologies whose documentation you can open today.

Job in B’s run Concrete systems you might use today What they give you What they still leave to you
Plan → fan-out → join LangGraph supervisor/swarm graphs; CrewAI hierarchical crews; OpenAI Agents SDK handoffs Explicit topology, checkpoints, human-in-the-loop gates Who owns each node’s authority after a handoff
Durable wait / crash recovery Temporal workflows (including agent plugins that memoize tool calls) Append-only event history; timers; “resume after human approval” Semantic meaning of why a step was authorized
Connecting tools and data MCP servers (the Model Context Protocol — a standard way for an agent to discover and call external tools and data sources) for the repository, Jira, the payments sandbox, the analytics warehouse One standard way for agents to discover and call tools and data sources Deciding while agents run what to allow, deny, or escalate — across the whole agent, not just one tool call
Agent-to-agent discovery A2A Agent Cards (/.well-known/agent-card.json) advertising skills and interfaces How independent agents find each other across runtimes Whether the receiving agent inherits the sender’s permission envelope
Procedures Agent Skills SKILL.md directories (and product variants in Claude Code, Cursor, Hermes, OpenClaw, Goose recipes) Versionable procedural memory with progressive disclosure Binding authority; who may edit the skill; audit of why it was selected
Persistent facts Hermes/OpenClaw MEMORY.md / USER.md layouts; Letta Memory Blocks Something that survives the session A portable interchange standard; enforced write policy; expiry as a first-class field
Policy outside the model External policy files in Rego or Cedar; NeMo Guardrails / RAIL for output rails Fail-closed verdicts at intervention points; structured output checks Fleet-wide ownership registry; learning policy; automatic linkage to outcomes
Evidence of the run OpenTelemetry GenAI spans; OpenInference; LangSmith / Langfuse / Phoenix Vendor-neutral-ish shape of model/tool/agent events; eval on traces A decision record: proposal → policy verdict → human → outcome
Regression before release LangSmith datasets (inputs / outputs / metadata); Inspect AI tasks; product agent-eval suites Reproducible scoring of “did this path still work?” Continuous outcome-linked compliance for live decisions and follow-through

Every row in that table is a real technology, and every entry in the last column is work the orchestrator still has to do. Later chapters go deeper on instructions (Chapter 6), topology and protocols (Chapter 7), and nondeterminism and telemetry (Chapter 8). Designing the system means choosing and wiring pieces like these, not inventing an orchestration operating system from scratch.

Here is a fragment Orchestrator B might check into a codebase. It is written in the format the Agent Skills specification already defines:8

---
name: pr-authz-review
description: Review a change for cross-tenant data exposure before merge.
compatibility: Requires read access to the PR diff and authorization tests; must not deploy.
---

# PR authorization review

1. Load the PR diff and the authorization test suite paths listed in `references/authz.md`.
2. Flag any query path that can return another tenant's records.
3. Require a failing test that reproduces the exposure before recommending merge.
4. Separate *observed* leakage from *hypothesized* risk.
5. Stop and escalate if the change touches payment or PII export tools.

The skill is useful the day it lands, and it is incomplete. Nothing in SKILL.md records who owns the skill, who may let an agent rewrite it overnight, or which policy has to pass before the coding agent can reach a deploy tool. The skill format covers the procedure. Ownership, authority, and audit have to come from somewhere else.

This is an area where no single widely adopted, interoperable, end-to-end permission standard covers agent skills yet — though the frontier is active. Researchers are already publishing candidates: SkillGuard, a 2026 paper, proposes dual-plane runtime governance of what a skill may load and what side effects it may cause; AgentBound specifically proposes Android-style permission manifests and enforcement for MCP servers and tools; NIST’s National Cybersecurity Center of Excellence has a project and a concept paper on software and AI agent identity and authorization, asking how least privilege, delegated authority, and prompt-injection defenses should work for agents. These are proposals and concept efforts, not adopted standards.9 That is something you will find frequently in this edition of the book: ideas and concepts that are obviously needed but have not yet been standardized.

We are not far from a machine-readable permission envelope that travels with every agent and skill. It will declare identity, allowed tools, data scope, network reach, delegation rights, memory access, side effects, spending limits, and the actions that require confirmation. The runtime — not the prose description — will enforce it. This book argues for one, and the chain in Section 4.6.7 is what I expect it to look like.

The first partial versions have begun to ship. NVIDIA’s September 2026 agent-safety launch compiles the operator’s file, network, tool, process, and credential limits into a runtime-verified policy — an envelope in exactly this sense for the resources an agent may touch, though it carries no authority record, no learning policy, and no cross-organization story.[^ch04_oasp_env] Microsoft’s ACS covers verdicts; the runtime-verified envelope covers resources. The full object — identity, tools, side effects, spend, confirmation thresholds, and the grantor behind all of it, portable across vendors — is still the open artifact this chapter is arguing for.


  1. Agent Skills, “Specification,” https://agentskills.io/specification; repository https://github.com/agentskills/agentskills. Defines SKILL.md (YAML frontmatter + Markdown body), optional scripts/references/assets, and progressive disclosure. Cited as the open packaging format for procedural skills — not as a permission or audit standard.↩︎

  2. SkillGuard (arXiv:2606.03024, https://arxiv.org/abs/2606.03024) — a 2026 research proposal for dual-plane skill permissions; AgentBound (arXiv:2510.21236, https://arxiv.org/abs/2510.21236) — proposes Android-inspired permission manifests and enforcement for MCP servers and tools specifically; NIST NCCoE Software and AI Agent Identity and Authorization project, https://www.nccoe.nist.gov/projects/software-and-ai-agent-identity-and-authorization — a project and concept effort, not a completed standard. All cited as candidates toward an end-to-end permission envelope, not adopted standards.↩︎