4.3 Topologies and decision records in the wild
The previous chapter treated topology as the relationship among contracts — who runs, in what order, what they pass, and where a human sits — and decision records as the evidence that makes a consequential action explainable afterward. Those ideas are older than the current tooling, and the tooling is finally starting to catch up. Both are early. Names collide, frameworks churn, and a “trace” is still not the same thing as an operating record an adjuster or a regulator can defend.
4.3.1 How systems relate to one another
The industry has begun naming a small set of relationship patterns. They are not mutually exclusive and none is automatically correct, because each one relocates where variation appears and who is responsible for joining the results.
A single agent with tools remains the honest default. Anthropic’s guidance on building effective agents is blunt about this: start simple, add multi-step and multi-agent structure only when a simpler system fails evaluation.20 For many operational workflows, one bounded agent plus deterministic checks is already a topology.
Orchestrator–worker / supervisor. A central agent decomposes work, calls specialists, and synthesizes results. The nondeterminism problem concentrates at the center: the supervisor’s routing decision and its join of worker outputs.
Swarm / handoff. Control moves from agent to agent. Variation is harder to localize, because there may be no single join point — whoever is active is the system until the next handoff.
Pipeline and critique loops. Sequential chains (retrieve → draft → check → act) and evaluator–optimizer loops (one agent produces, another critiques, the first revises) show up in the same pattern catalogs. Pipelines make intermediate variation inspectable; critique loops can improve quality or amplify correlated agreement if both agents share the same flawed source.
The durable question for the orchestrator is not which library won last quarter. It is whether the topology makes authority, handoff, and stop conditions legible — and whether a second person can see, from the record, which agent was allowed to do what when the answer varied. Run the Castellano test against each pattern and the differences stop being architectural taste: in the single-agent arrangement, her three answers come from three runs and the record shows one join point; in the swarm, the variation is spread across handoffs and nobody owns the join at all; in the critique loop, a second agent is structurally positioned to catch the partial-denial drift — if the envelope told it what drift looks like. The topology is not how the system is drawn. It is where the variation lands and who is standing there when it does.
4.3.2 How decisions are kept
A topology without a record is a system that can be trusted only while somebody is watching it. The industry’s answer so far has mostly been observability — capture the run so engineers can debug it — which is necessary and not sufficient for consequential orchestration, because the operating record this chapter asks for also needs authority, human decisions, and outcome, three fields that most LLM traces still treat as optional metadata.
Two layers are forming around that need.
A shared telemetry substrate. OpenTelemetry’s generative-AI semantic conventions are the closest thing to a vendor-neutral vocabulary for what happened: model calls, agent invocations, and tool executions as spans with agreed attributes.21 Instrument once, send traces to more than one backend. The conventions are still evolving (many GenAI span definitions remain in development status), but they are the right place to bet if an organization does not want every agent framework to invent a private log format — another small instance of the architecture of participation: the record should outlive the dashboard you bought this year.
Trace and evaluation platforms. On top of that substrate — or beside it — sit tools that turn runs into something a team can inspect and score: Langfuse, Arize Phoenix, LangSmith, AgentOps, Braintrust, and OpenAI’s own trace grading for end-to-end agent workflows among them.222324 Prefer at least one path you can self-host or export from. When the operating record is trapped in a single SaaS UI, “who can stop the system” quietly becomes “who still has a login.”
Had the claims workflow run on one of these platforms, the Castellano load test would have left a different artifact: not a meeting where three results surprised a team, but three traces with envelopes attached, each showing which rules were active when the 0.83 confidence denial was proposed. The gap an orchestrator still has to close by design is the gap between engineering telemetry and an operating record. A beautiful waterfall of tool calls that omits which permission envelope was active, which human owned the boundary, and what customer outcome followed is useful for latency debugging and incomplete for defending a claim decision. The fields that matter for nondeterminism are the ones Section 8.5 names: input class, versions, instructions, authority, evidence, human decisions, exceptions, and outcome. Choose platforms that can carry those fields, or wrap them until they can.
Anthropic, “Building effective agents,” https://www.anthropic.com/engineering/building-effective-agents, and Claude Cookbook, “Orchestrator-workers,” https://platform.claude.com/cookbook/patterns-agents-orchestrator-workers. Cited for the pattern catalog (including orchestrator-workers, parallelization, and evaluator-optimizer style loops) and for the guidance to prefer simpler systems until evaluation shows otherwise. Not cited as a claim that any one pattern dominates production.↩︎
OpenTelemetry, generative AI semantic conventions (spans for model, agent, and tool operations), https://opentelemetry.io/docs/specs/semconv/gen-ai/ and https://github.com/open-telemetry/semantic-conventions/blob/main/docs/gen-ai/gen-ai-spans.md. Cited for the emerging vendor-neutral vocabulary (
create_agent,invoke_agent,execute_tool, and related attributes). Many GenAI conventions remain in development status; the chapter cites the direction of standardization, not finished stability.↩︎Langfuse, “AI Agent Observability with Langfuse,” https://langfuse.com/blog/2024-07-ai-agent-observability-with-langfuse, and observability documentation, https://langfuse.com/docs/observability/overview. Cited for open-source tracing, agent-graph views, evaluation on captured runs, and OpenTelemetry / OTLP ingestion across agent frameworks. Product surface evolves; cited for the category of self-hostable agent observability.↩︎
Arize Phoenix documentation and OpenInference instrumentation model (open-source agent/LLM tracing with evaluation support), https://arize.com/docs/phoenix and https://github.com/Arize-ai/openinference. Cited as an OpenTelemetry-oriented tracing and evaluation stack for multi-step agents. Not cited as uniquely sufficient for regulatory operating records.↩︎
Representative tools in the same category as of the research date: LangSmith (LangChain/LangGraph-native tracing and evals), https://docs.smith.langchain.com/; AgentOps (agent session/event replay), https://docs.agentops.ai/; Braintrust (trace-linked experiments and regressions), https://www.braintrust.dev/docs. Cited collectively to show a forming category rather than to rank vendors. Capabilities and hosting models change quickly; none replaces a designed operating record with authority and outcome fields.↩︎
- 4.1 Variation and fences
- 4.2 The failure envelope
- 4.3 Topologies and decision records in the wild
- 4.4 Training the envelope
- 4.5 The operating record