3.3 Topology

Book 2 · The Delegation ContractChapter 3 · section 3 of 5

3.3.1 Topology is the architecture

A delegation contract is a single row, and a topology is the relationship among rows: the order in which contracts run, the data they pass, the boundaries they cross, and the human gates sitting between them. The home-goods company’s incident runbook is a topology of four contracts.

The industry is starting to draw these arrangements under a name of its own. Anthropic’s engineering guidance calls them workflows and names five — prompt chaining, routing, parallelization, orchestrator-workers, evaluator-optimizer — describing, for instance, how “a central LLM dynamically breaks down tasks, delegates them to worker LLMs, and synthesizes their results.”38 Microsoft’s Azure architecture guidance catalogs multi-agent patterns — sequential, concurrent, group chat, handoff — with the same advice this chapter has been making: use the lowest level of complexity that meets the requirement.39 What this book adds by calling the arrangement a topology is the part the pattern catalogs leave out: the contract at each node, the human gates between them, and the records the whole thing leaves behind.

Topologies come in two broad shapes. In a static topology, the components and their relationships are fixed before the work starts, so the same return workflow always runs intake → eligibility → action → message with the same handoffs and permissions for every request — which makes it easier to test, because the orchestrator knows in advance which agents should run and which records should exist. In a dynamic topology, the system creates or selects components during the work, so a coordinator receiving an unusual order investigation might start three research agents, ask a fourth to compare their findings, and shut those workers down when the report is complete, while a routine request uses only one. The arrangement changes according to the task.

Dynamic does not mean uncontrolled, though. Every dynamically created agent still needs a record of who created it, what task it received, what context and permissions it had, what it returned, whether it created another agent, and when its authority ended, and the parent and child work should share a trace or correlation ID so an investigator can reconstruct the run.

detected anomaly signal
        |
        v
[routing agent] -- proposes isolation --> [approval gate]
        |                              |
        |                     deterministic health check
        |                              |
        v                              v
   operating record <-------- [routing agent]
                                      |
                              changes one routing rule
        |
        v
   on-call human page

The contracts differ at every node. The routing agent has read access to schema and incidents; the approval gate has read access to the proposal and to a health-check service; write access extends only to the routing rule for the affected store cohort; and the on-call page fires only after a structured isolation proposal has cleared the health check. The on-call human can reverse any action with a single command from a known runbook.

Relax any one of those boundaries and the topology produces a different failure mode. Allow the routing agent to change a global routing rule and the company can divert unaffected orders during a false alarm and lose revenue. Allow it to page a human without a structured explanation and the on-call engineer gets an alert without knowing which stores or customers are at risk. Make the health check a model instead of a deterministic test and the approval gate inherits the model’s nondeterminism, so the human is approving noise rather than evidence.

3.3.2 The smallest useful topology

A system that is running perfectly — stable, never changing, never having outages — does not need any of this, and an orchestrator who reinvents a working system for the sake of making it agentic has not improved it.

So start with the smallest arrangement that exposes the risk. If one agent and one deterministic test answer the question, adding six agents is not sophistication; it is a new system to operate. The home-goods runbook has four contracts because four was the smallest number that separated read from approve from route from page. Fewer would have collapsed a boundary, and more would have added coordination debt without exposing a new risk.

Coordination debt is the accumulated cost of keeping delegated components aligned — the multi-agent version of a metaphor software has used since Ward Cunningham coined technical debt in 1992: expedient choices that speed things up now and cost more later, compounding until someone pays them down.40 It rises when roles overlap, outputs are untyped, state travels through prose, and nobody knows which result is authoritative; it falls when interfaces are explicit, evidence is structured, and the workflow has a clear stopping point. Before adding a component, an orchestrator should ask one blunt question — what failure becomes less likely because this exists? — and if the only honest answer is that the demo looks more agentic, the component is decoration with an API key.

Four signs that a topology is carrying too much of that debt:

  • The team can describe the topology’s job in one sentence. If they cannot, the topology has grown past its name.
  • Each contract has a typed output. If contracts pass prose between them and rely on a later model to interpret, debt is rising.
  • The state can be reconstructed from a single log. If not, multiple agents are silently maintaining parallel truths.
  • There is a named human who can stop the topology at any moment. If not, the topology has removed the institution from its own decision.

Three out of four is a yellow light, and two out of four is a red one.

Drawn as the artifact it is, each handoff between contracts carries the same five fields — the handoff object, which is what makes the topology auditable line by line:

 contract A ──► [ handoff object ] ──► contract B
                  principal:  who acts next (agent id, human name)
                  scope:      what that principal may do with it
                  budget:     what the step may spend (tokens, money, time)
                  evidence:   what must accompany the result
                  owner:      who is called when it stops

A handoff missing any field is the topology confessing where its next failure will surface: no principal, and nobody acts; no scope, and anything acts; no budget, and the cost chapter’s surprise invoice writes itself; no evidence, and the result cannot be checked; no owner, and nothing stops it.

3.3.3 Operational orchestration is starting to have a stack

The vocabulary in this chapter still sounds ahead of the tooling, but the tooling is beginning to catch up — not as a finished profession, and not yet as a settled standard for consequential operations, but as a recognizable category. Vendors, open protocols, and durable-execution platforms are starting to talk about what this chapter calls operational orchestration: agents and workflows that run inside a live business process, coordinate with other agents, survive failure, and sit behind human gates.

Three strands are visible now.

Managed runtimes for multi-agent business workflows. Cloud platforms are no longer only offering a chat box with tools. Microsoft’s Agent Framework and Foundry Agent Service describe agents and graph-based workflows with checkpointing, human-in-the-loop steps, and deployment into a managed runtime — and they name production-shaped uses such as financial transaction processing and supply-chain automation, not only coding assistants.41 That is operational topology language entering first-party documentation.

The same machinery is arriving in the open-source stack. LangGraph’s interrupt primitive — one line of code that pauses the graph mid-run, surfaces the proposed action, and resumes only when a person responds — turns the approval gate from a diagram element into a shipped feature.42

Protocols that make the edges of a topology explicit. Two open protocols are starting to separate concerns that used to be mashed into one integration mess. Anthropic’s Model Context Protocol (MCP) — now under the Linux Foundation’s Agentic AI Foundation alongside Goose and AGENTS.md — standardizes how an agent reaches tools and data — the vertical edges of the arrangement.43 Google’s Agent2Agent protocol (A2A), also under the Linux Foundation, standardizes how independent agents discover one another and collaborate across frameworks and servers — the horizontal edges. Microsoft’s own Agent Framework materials describe using both: MCP for tools, A2A for agent-to-agent collaboration. That open stewardship matters for the same reason Chapter 5 borrowed O’Reilly’s phrase: topology edges that only one vendor will let you speak are not really edges you own. A topology in this chapter’s sense is no longer only a whiteboard diagram; it is beginning to have wire formats for “what this node can call” and “which other agents exist.”

Durable execution for long-running operational work. An operational agent that dies mid-run is not an academic inconvenience. It is a half-applied routing change, a payment held without a record, or a shipment left without an owner. Durable-execution platforms such as Temporal are positioning themselves as the substrate that keeps agent workflows alive across crashes, retries, and human waits — including first-party integrations that wrap agent SDKs so a tool call or model step that already completed is not paid for twice after a restart.44 That is the operating arrangement starting to borrow reliability machinery from the rest of distributed systems, which is exactly where operational orchestration belongs.

None of this yet equals the full discipline this chapter is naming. A workflow builder that can draw four agents does not, by itself, produce a permission envelope, an authority owner, or a decision record a regulator can read. A protocol that lets agents talk does not decide what they are allowed to decide. Durable execution keeps a run alive; it does not decide whether the run should have been allowed.


  1. Anthropic, “Building Effective Agents,” December 2024, https://www.anthropic.com/engineering/building-effective-agents. Names five workflows — prompt chaining, routing, parallelization, orchestrator-workers, evaluator-optimizer — including the orchestrator-workers arrangement in which “a central LLM dynamically breaks down tasks, delegates them to worker LLMs, and synthesizes its results,” and advises adding complexity only when a simpler shape cannot meet the requirement.↩︎

  2. Microsoft, “AI agent orchestration patterns,” Azure Architecture Center, https://learn.microsoft.com/en-us/azure/architecture/ai-ml/guide/ai-agent-design-patterns. First-party guidance on multi-agent orchestration patterns (including sequential, concurrent, group chat, and handoff) and on choosing the lowest complexity that meets the requirement. Cited for the industry beginning to document operational topology choices, not as proof that any one pattern is dominant.↩︎

  3. Ward Cunningham coined the technical-debt metaphor in 1992: “a little debt speeds development so long as it is paid back promptly with refactoring. The danger occurs when the debt is not repaid.” See Cunningham, “The WyCash Portfolio Management System,” Addendum to the OOPSLA 1992 Proceedings, and http://wiki.c2.com/?TechnicalDebt.↩︎

  4. Microsoft, “Introducing Microsoft Agent Framework,” Azure Blog, https://azure.microsoft.com/en-us/blog/introducing-microsoft-agent-framework/; Microsoft Learn, “Microsoft Agent Framework overview,” https://learn.microsoft.com/en-us/agent-framework/overview/; and Foundry Agent Service / workflow concepts, https://learn.microsoft.com/en-us/azure/foundry/agents/concepts/workflow. Cited for managed multi-agent workflows, checkpointing, human-in-the-loop support, MCP/A2A integration, and named business-process uses (e.g. transaction processing, supply-chain automation). Product surfaces are early and changing; the load-bearing claim is that vendors are shipping operational orchestration stacks, not only developer copilots.↩︎

  5. LangChain, “Making it easier to build human-in-the-loop agents with interrupt,” https://www.langchain.com/blog/making-it-easier-to-build-human-in-the-loop-agents-with-interrupt, and LangGraph interrupts documentation, https://docs.langchain.com/oss/python/langgraph/interrupts.↩︎

  6. Anthropic, Model Context Protocol, https://modelcontextprotocol.io/ and https://github.com/modelcontextprotocol. Cited for the emerging standard that connects agents/applications to tools and data sources (the vertical edges of a topology). Complementary to agent-to-agent protocols; not a complete operational governance model.↩︎

  7. Temporal, “AI Applications & Agents With Temporal,” https://temporal.io/solutions/ai, and “Production-ready agents with the OpenAI Agents SDK + Temporal,” https://temporal.io/blog/announcing-openai-agents-sdk-integration. Cited for durable execution as a substrate for long-running agent workflows: crash recovery, retries, preserved progress across model/tool steps, and human-in-the-loop waits. Not cited as the only durable runtime, and not as a substitute for authority design.↩︎