7.5 Evidence packs: the operating record as infrastructure
There is one more layer of infrastructure an orchestrated organization cannot buy finished: the assembled evidence pack that lets a meeting, a reviewer, an auditor, or a regulator reconstruct what a delegated system did, why it was allowed to do it, and what the consequence was. Without it, every review meeting reverts to guessing, every incident becomes archaeology, and every audit becomes a negotiation.
Chapter 8 named the gap: the industry built engineering telemetry, and the operating record is still assembled by hand.
When the prepared evidence is wrong
The failure that proves the need arrives as a customer complaint, not as a slide. A customer writes that a lamp arrived broken and asks where the refund is. Maya’s return workflow from Chapter 10 uses four agents — intake, eligibility, action, and message — and the message agent is not allowed to say a refund is complete until the payment system confirms it. Something in that chain told a customer “Your refund has been completed” while the payment system still showed the request as pending.
The first piece of infrastructure is the joined record itself. For one order, the operating record should assemble what each agent extracted, approved, submitted, and sent, and which authority let the message go out in the first place — intake classification, eligibility decision, action submission with payment status, the message sent to the customer, and the review check that exposes the mismatch between a successful API request and a completed business outcome.
Maya needs two views over that record, which operations people will recognize under other names, system-level and component-level. This book calls them wide mode and deep mode. They are how she decides whether she is investigating a workflow or one agent inside it.
Wide mode follows one outcome. She starts with the joined record for the affected order, then asks whether the problem is one order or a cohort — how many customers got the same message, which stores were affected, which agent version produced it, and whether the behavior started after an instruction, skill, memory, model, or schema change.
Deep mode follows one component. Once the wide view points at a single component, she selects the message agent and asks a narrow question: when payment status was pending, did this agent tell customers the refund was complete? The answer arrives as counts — how many messages, how many incorrect completion claims, broken down by the payment status the agent saw at decision time.
Now she has a specific investigation with a handful of candidate causes: whether the agent received the pending status at all, whether a schema renamed a field, whether a skill example equated submitted with complete, whether memory taught the wrong assumption, or whether one node was still running an old configuration. Valid JSON can be entirely wrong, because parseability is not truth.
She may hand the grouping work to a telemetry-analysis agent and ask it to return links to evidence, and that agent may not change refund policy or deploy a new message-agent version unless she granted it that authority — the same approval discipline that governs the meeting.
The join keys that make both modes possible
Neither mode works without shared identifiers across agents, payment services, and support records. Every consequential event needs at least these fields:
trace_id— joins every event for one customer request;decision_id/parent_decision_id— identifies one agent decision and its cause;agent_id— the component that acted;entity_refs— orders, customers, stores, shipments affected;input_refs/output— evidence used and what was produced;authority— what the agent was allowed to do;versions— model, skill, memory, policy, schema active at the time; andoutcome— pending / confirmed / rejected, the real result.
A dashboard that shows token counts and cannot identify affected orders may help with cost, and it will not help Maya answer a customer complaint or give a review meeting evidence worth arguing about.
How far today’s standards get you — and where they stop
Those fields are not a private invention, since pieces of them already have homes in public specs. The mistake is assuming the homes are finished houses.
| Field you need | Closest public artifact today | What it still will not do for you |
|---|---|---|
trace_id + agent/tool spans |
OpenTelemetry GenAI semantic conventions; OpenInference span kinds (AGENT, TOOL, LLM…); product tracers (Langfuse, Phoenix, LangSmith)13 | Guarantee entity_refs or payment outcome are on the span — you must attach them |
| Durable step history | Temporal workflow event history14 | Explain why a step was authorized; encode refund completed-vs-pending semantics unless your workflow code does |
| Handoff from intake → message | A2A Task/Message/Artifact exchange; OpenAI Agents SDK HandoffSpanData15 | Transfer the permission envelope with the task; record who owns the decision after the handoff |
authority / allow-deny-escalate |
Microsoft Agent Control Specification manifests and verdicts (allow, deny, warn, escalate, transform); host-specific guardrails16 | A mature, cross-vendor fleet registry of owners, budgets, and revocation; automatic join from verdict → customer outcome |
| Tool edges | MCP tool descriptors + OAuth for HTTP transports17 | Runtime policy over the whole agent loop; MCP’s own docs: the protocol cannot enforce implementor security principles |
| Skill / memory versions | Agent Skills SKILL.md; Hermes/OpenClaw memory files; Letta blocks18 |
A standard “this version was active on decision D” record that every tracer emits the same way |
| Offline regression | LangSmith datasets; Inspect AI tasks; product agent-eval suites19 | Prove that Tuesday’s live cohort matched policy — evals are not operating audit |
What the industry has not standardized is the object a skeptical meeting actually needs:
proposal → policy verdict → human decision → customer or system outcome
joined on identifiers a support lead and a payments engineer both recognize.
A second room: the incident
The same discipline shows up when the room is an incident channel instead of a review meeting. At 02:14 checkout errors spike. The old war room starts with “is it us or the provider?” and burns twenty minutes collecting screenshots. The prepared path: a Temporal-owned triage workflow with MCP read tools for status and deploy diffs, GenAI spans already carrying service.name and deploy SHA, an ACS policy that denies the “post to status page” tool without a human verdict, and a single evidence pack that already answers whether the canary cohort and the provider outage window overlap. People still decide whether to roll back. They do not spend the decision window reconstructing timelines from memory.
Either way the operating sequence is the same, whether the pack is preparing a meeting or cleaning up after a bad message:
- Start with the customer or business outcome.
- Join the records.
- Inspect the whole workflow (wide).
- Select one component if necessary (deep).
- Check the active state — instructions, skills, memory, tools, permissions, versions.
- Change one thing.
- Check the full workflow again.
Miss the trace_id and none of this works cleanly, which is not a tooling preference but an orchestration failure: the system can act without leaving any path for people to review what it did. Everything in this pack also lives inside one organization’s systems. The moment a delegated decision crosses into a system somebody else runs, the authority and outcome fields have to cross with it, and nothing yet carries them — which is where the next chapter begins. [^ch16_oasp]: NVIDIA, “Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring,” NVIDIA Technical Blog, September 28, 2026, https://developer.nvidia.com/blog/nvidia-open-agent-safety-platform-a-reference-for-continuous-in-silicon-agent-monitoring/ — BlueField-4 “on the node’s only path to the model” in the Vera Rubin POD reference design; continuous out-of-band observability and real-time policy enforcement at line speed; the DOCA gateway “continuously verifying each agent’s identity and delegated authority”; quarantine in milliseconds. Announced September 28, 2026, one day before this paragraph was written; cited for the topology — monitoring and enforcement in infrastructure the agent’s host does not control — not as evidence of maturity. Verified September 29, 2026.
HQ 6 — Assembled. The human and AI each wrote portions of this chapter. I assembled, reviewed, and take responsibility for the whole; the voice and arguments are mine, and I know which parts are which.
OpenTelemetry generative AI semantic conventions, https://github.com/open-telemetry/semantic-conventions-genai (development status); OpenInference, https://arize-ai.github.io/openinference/spec/semantic_conventions.html. Cited for span shape — not as a complete operating-record schema for meetings.↩︎
Temporal, AI / durable agent solutions, https://temporal.io/solutions/ai. Cited for workflow event history and human-in-the-loop waits — not for semantic refund or authority fields.↩︎
Agent2Agent (A2A) Protocol, https://a2a-protocol.org/v1.0.0/specification/; OpenAI Agents SDK tracing/handoffs, https://openai.github.io/openai-agents-python/tracing/. Cited for handoff mechanics — not for authority transfer.↩︎
Microsoft Agent Control Specification, https://github.com/microsoft/agent-governance-toolkit/blob/main/policy-engine/spec/SPECIFICATION.md. Cited for portable policy manifests and verdicts; early relative to fleet-wide ownership registries.↩︎
Model Context Protocol, https://modelcontextprotocol.io/specification/2025-11-25/ (incl. authorization). Cited for tool connectivity; protocol-level enforcement limits apply.↩︎
Agent Skills, https://agentskills.io/specification; OpenClaw memory, https://docs.openclaw.ai/concepts/memory; Letta Memory Blocks, https://docs.letta.com/guides/core-concepts/memory/memory-blocks/. Cited for versionable procedure and persistence layouts.↩︎
LangSmith datasets, https://docs.langchain.com/langsmith/; Inspect AI, https://inspect.aisi.org.uk/. Cited for regression/capability evals — not live cohort compliance audit.↩︎
- 7.1 Does the agentic era still need source control?
- 7.2 Task systems: from bug trackers to agent work boards
- 7.3 How fleets communicate
- 7.4 Where the standards are being written
- 7.5 Evidence packs: the operating record as infrastructure