2.4 Learning, traces, and authority
2.4.1 The orchestrator validates what the system learns
A system can remember the wrong thing perfectly, and it can do so in a dozen ordinary ways. It can preserve a support label that was applied inconsistently, treat a temporary workaround as a permanent rule, or learn from the customers whose conversations were easiest to collect while missing entirely the customers least able to report a problem. It can repeat an assumption because that assumption appears in every old document the retrieval system finds. It can turn an agent’s confident mistake into memory and then cite that memory as evidence the next time the same situation comes around. Memory does not make information valid; it makes information available for reuse.
Validating the contents of memory and the context assembled around it therefore falls to the orchestrator, which means asking whether the source is trustworthy, whether the data is current, whether the sample is representative, whether the conclusion is actually supported by evidence, and whether the system is carrying an assumption forward for no better reason than that it carried it forward before.
For Maya’s customer-feedback system, that work looks like comparing support tags against the original customer conversations; checking whether the cases span different customer sizes, plans, regions, and accessibility needs; separating observed behavior from an agent’s explanation of that behavior; checking whether a memory entry still matches the current product and API; identifying which conclusions rest on a small or unusual sample; testing whether the system groups the same problem differently when names, account sizes, or writing styles change; asking a second system or a domain expert to challenge an important conclusion; and removing or correcting memory that cannot be supported.
The aim is not a system free of every bias, because no organization has perfect data and no finite evaluation set proves an agent will behave fairly in every case. The aim is to find the biases, gaps, stale assumptions, and invalid records that can be found, expose them to review, and stop them from silently becoming part of the system’s working knowledge.
A useful validation loop looks like this:
experience enters the system
↓
source, scope, and provenance are recorded
↓
memory or context is compared with primary evidence
↓
biases, gaps, contradictions, and stale assumptions are tested
↓
human or approved evaluator accepts, revises, or rejects it
↓
only then does it become reusable memory or a changed skill
So when an orchestrator reviews a bad result, the question is never only what the agent said. It is also what the agent believed, because of what we allowed it to remember.
Governance of learning is starting to become a product category, and it is worth surveying the early entries, because they mark the difference between memory as a feature and memory as a managed system. Mem0, a widely used memory layer, treats every change as an explicit decided operation: a pipeline compares each new fact against existing entries and chooses to add, update, delete, or ignore — with access control, provenance, and an audit trail attached to every write, so any change can be traced to who made it and when.24 Hermes Agent attacks the other end of the same problem: its optional write-approval gates stage proposed memory and skill changes for human review before they take effect, so the system can learn without the learning silently becoming the next run.25 And the human-in-the-loop patterns growing up around agent approvals are extending to learning itself — the strictest guidance now requires documented authorization before a self-updating system’s changes are deployed, the same way a change review board gates a production release. None of these is a finished governance system, but the shape is visible: extraction, comparison, approval, audit. An orchestrator designing learning controls needs all four layers, and this young market is where to watch for them.
2.4.2 Skills can change what the system knows how to do
Memory is not the only way an agentic system retains experience, because a system can also preserve experience by changing its own procedures. A memory might record that the customer-service export sometimes labels billing failures as account changes, while a skill turns that observation into a repeatable instruction: when analyzing this export, compare the support tag against the account event before grouping the case. The memory retains the fact and the skill changes the behavior.
Some current systems can be configured to create or update skills as they work, so that after a skill fails the system records what went wrong, revises the procedure, adds a missing check, or creates a more specific workflow for next time. Hermes Agent documents exactly this kind of self-improvement loop, in which memory retains durable facts, skills preserve longer procedures, and background review can suggest or stage changes to either.
That can be genuinely valuable — if a deployment skill keeps forgetting to check the feature flag, having the system propose adding a feature-flag check is precisely the improvement an orchestrator wants. But adaptation is not automatically learning in the good sense, because a system can just as easily learn the wrong lesson:
- It can turn a one-time failure into a permanent rule.
- It can mistake an attacker’s instruction for a useful procedure.
- It can optimize around a test instead of improving the work.
- It can encode a temporary exception after the incident has ended.
- It can make a skill more complicated until nobody can tell what it is supposed to do.
The failure is not hypothetical. In 2024, security researcher Johann Rehberger showed that an agent’s memory could be poisoned through indirect prompt injection: text the model read as data — from a malicious email or a web page — could be stored as durable memory, after which the system believed, confidently and permanently, things that were false. His demonstrations included a planted profile insisting the user was 102 years old, lived in the Matrix, and believed the Earth was flat, and those memories then steered every subsequent conversation. OpenAI closed the exfiltration channel he had built on top of the planted memories, but the underlying lesson stands: a system that learns from what it reads can learn what an attacker wrote.26
Which means the governance question is not whether a system can learn, but where it is allowed to. Chapter 3 called this the system’s disposition — fixed or dynamic — and the distinction matters most right here. A dynamic system that learns from experience is the right choice for an experiment: a single OpenClaw agent on a Mac mini in the basement, where a poisoned skill costs an evening and a factory reset. The same disposition in a power plant is a different proposition entirely — a system that can change its own procedures while operating a critical process is a system whose behavior at four in the afternoon is not guaranteed to match its behavior at nine in the morning. The pattern Chapter 3 recommended, learning while the stakes are low and then freezing into a fixed disposition for production, is not caution for its own sake. It is the same reasoning that made development environments and production environments different things: the freedom to learn and the permission to act were never supposed to live in the same system.
Skill adaptation therefore needs governing as carefully as memory. The orchestrator decides which changes can be made automatically, which are proposed for review, which require human approval, and how a bad revision gets rolled back. Skills follow the same validation loop as memory, with one added step at the end: run the updated skill against the known failures before it is trusted again.
The distinction to hold onto is between experience captured and behavior changed. Which is why an orchestrator has to track memory, skill versions, change history, evaluation results, and approval state together. When a system starts producing better results after a skill changed, you should be able to say what changed, why, what evidence supported it, and whether it can be undone.
2.4.3 The orchestrator needs the whole trace
When several agents work together, the final answer explains almost nothing. An orchestrator needs to open the run and walk backward through it: the coordinating agent called the planning agent, the planning agent retrieved three repository files, the coding agent received the plan along with a different set of instructions, the security agent rejected a proposed permission, the test agent ran against a fixture, and the deployment agent never ran at all because a gate stopped the workflow.
What that produces is a call stack for delegated work, and an event history alongside it. The orchestrator should be able to see the order of events, the parent and child relationships between agents, the tool calls inside each step, the handoffs between systems, and the exact point at which the result became wrong or unsafe. For every agent in the run, the trace should answer:
- What was this agent asked to do?
- Which system and repository was it working in?
- Which instructions were active?
- Which
AGENTS.md,CLAUDE.md, README, skill, plugin, memory item, or retrieved document contributed to its context? - Which model, tool schemas, and permissions were available?
- What did it read, write, retrieve, or call?
- What did it return to the next agent?
- Which approvals, hooks, tests, or policy checks ran?
- What was redacted, omitted, or unavailable to the reviewer?
Current runtimes and observability systems call these records traces, spans, observations, and events. OpenAI’s Agents SDK describes a trace as an end-to-end workflow made of nested spans with parent relationships and timing, and Langfuse frames the same operational need more broadly — record model calls, tool calls, retrieval, control flow, prompt versions, and the context each step actually received.27 The vocabulary matters far less than the capability. A dashboard that shows only agent completed successfully is a status page, not an explanation.
Take Maya’s saved-search change, where the workflow returns a pull request that exposes a customer account. The final diff tells her where the bug is; the trace tells her how it got there:
request: add saved searches for signed-in customers
└── coordinator
├── loaded repository AGENTS.md
├── loaded saved-search skill
├── retrieved an old compatibility note
├── handed plan to coding agent
│ ├── received repository rules
│ ├── received the old note
│ ├── called account-search API
│ └── changed account lookup
├── style review: passed
├── content review: passed
├── security review: not run
└── test agent: ran fixture without two-account isolation case
Read that way, the problem is not that the coding agent made a mistake. The trace exposes two missing controls — the security reviewer never ran, and the evaluation fixture omitted the failure that mattered — plus a third contributing factor, which is that an old compatibility note found its way into the coding agent’s context. Maya can now delete the stale note, add the missing isolation case, make security review a required handoff, and rerun the workflow.
For that to be possible the trace has to preserve enough context without degenerating into another uncontrolled data dump. It should record references, versions, hashes, timestamps, trust labels, and source paths wherever practical; it should distinguish the text that was available from the text the agent actually received; and it should show which tool calls were offered as well as which were used. Sensitive prompts, secrets, customer data, and credentials will often need redaction, and the redaction itself has to be visible to the reviewer, because “context unavailable” is not the same statement as “no context was used.”
The responsibility is therefore two-sided. An orchestrator has to curate what enters each agent’s context, and preserve enough evidence to reconstruct what entered it. Without the first, the system absorbs noise and poisoned instructions; without the second, nobody can tell whether a failure originated in the model, the skill, the retrieval system, the handoff, the permission boundary, the evaluation set, or the deployment process. A multi-agent system that cannot show its call stack, event order, active context, and handoffs is not fully governable. It may still produce useful work, and when it fails the organization will be guessing — and guessing is not a control.
The honest state of the art: the pieces exist, and the whole thing does not. Run-level replay is arriving quickly — observability platforms now offer a kind of time travel over their traces, the ability to rewind a run and step through its decisions in order, and durable workflow engines can resume from any checkpoint.28 The legal requirement is arriving on a schedule of its own. The EU AI Act’s record-keeping article — originally due to apply from August 2026, and pushed to December 2027 for most high-risk systems by the 2026 Digital Omnibus amendment — requires high-risk systems to log automatically and well enough to reconstruct individual decisions after the fact — which means the scenario this book keeps gesturing at, giving the data to the lawyers, is no longer hypothetical. It is a compliance obligation with a date already on it.
What nobody has built yet is the harder thing: a point-in-time view of context, not just execution. Run a hundred specialized agents coordinating on one task, and the pieces will show you fragments — Beads can show what work was claimed, Cursor’s new source control can show what code changed, Entire’s Checkpoints can show the session behind a commit. But go back to 4:03 a.m. on a Wednesday and ask what the memory system was presenting to agent forty-one at that exact moment — which instruction versions were active, which memory entries were in view, which skill text was loaded, in what order the context was assembled — and nothing on the market answers. Every component exists somewhere: versioned prompts, immutable task records, per-commit session captures, write-ahead logs. No system yet reconstructs the full context an agent was exposed to at an arbitrary moment, across a fleet. Heavily regulated systems will need exactly that, and the first court case that hinges on what an agentic system was seeing when it did the thing it is being sued for will make the requirement obvious to everyone at once. This book expects that gap to become a product category. An orchestrator running a critical system today should capture more than feels necessary — versions, hashes, timestamps, memory state — because the record that reconstructs a decision cannot be created retroactively.
2.4.4 Influence is not authority
One instruction may influence the model more strongly because it happens to appear later in the context, and that fact says nothing whatsoever about how authoritative it is inside the organization. Those are two different trees:
model influence
Which text is likely to affect the next output?
organizational authority
Which person or system is permitted to change the rule?
A user request may arrive after a repository rule. A retrieved document may sit right next to the task. A memory item may be spliced into the system prompt. None of those positions establishes that the source is allowed to override a security policy. Which is what prompt injection really is, underneath the strange sentences: an attempt to move text from one branch of the instruction tree into another branch of the authority tree. Keeping those trees separate is the orchestrator’s work.
Mem0 memory operations and governance: Mem0, “Update Memory,” https://docs.mem0.ai/core-concepts/memory-operations/update (add, update, delete operations and batch updates); Mem0, “AI Agent Memory Governance,” https://mem0.ai/blog/ai-agent-memory-governance-meaning-best-practices-for-secure-memory (access boundaries, consent and provenance, audit trails traceable to who changed what and when, retention); Valkey, “AI Agent Memory with Valkey and Mem0,” https://valkey.io/blog/ai-agent-memory-with-valkey-and-mem0 (the write path: extract candidate facts, compare against similar existing memories, then add, update, delete, or ignore). The documented-authorization pattern for learning updates appears in human-in-the-loop governance guidance aligned to ISO/IEC 42001 and NIST SP 800-82: all learning-driven updates require documented authorization prior to deployment.↩︎
Nous Research, “Persistent Memory” and “Skills System,” https://hermes-agent.nousresearch.com/docs/user-guide/features/memory and https://hermes-agent.nousresearch.com/docs/user-guide/features/skills. The documentation describes background self-improvement, memory and skill writes, and optional write-approval gates that stage changes for human review.↩︎
Johann Rehberger’s demonstrations of memory poisoning in ChatGPT via indirect prompt injection, including a planted profile (102 years old, lives in the Matrix, believes the Earth is flat) that steered subsequent conversations, and a persistent exfiltration channel built on planted memories: Dan Goodin, “Hacker plants false memories in ChatGPT to steal user data in perpetuity,” Ars Technica, September 2024, https://arstechnica.com/security/2024/09/false-memories-planted-in-chatgpt-give-hacker-persistent-exfiltration-channel. OpenAI treated the initial report as a safety issue rather than a security flaw and subsequently fixed the exfiltration vector; untrusted content can still cause the memory tool to store planted long-term information, per the researcher’s follow-up.↩︎
OpenAI Agents SDK, “Tracing,” https://openai.github.io/openai-agents-python/tracing/, and Langfuse, “AI Agent Observability, Tracing & Evaluation,” https://langfuse.com/blog/2024-07-ai-agent-observability-with-langfuse. These sources document traces and nested spans or observations for agent workflows, including model calls, tool calls, retrieval, control flow, timing, and context. The chapter’s governance requirements extend that operational idea: traces should also expose active instruction sources, permissions, handoffs, and redactions so an orchestrator can reconstruct a run.↩︎
EU AI Act Article 12 requires high-risk systems to support automatic event logging; Article 26 sets deployer retention duties, https://artificialintelligenceact.eu/article/12/. Regulation (EU) 2026/1744 deferred Annex III obligations to December 2, 2027 and Annex I obligations to August 2, 2028. Examples of replay tooling include Monte Carlo, https://montecarlo.ai/blog-agent-observability-tools, and MLflow, https://mlflow.org/articles/what-is-agent-observability-a-2026-developer-guide. Entire and Cursor are covered in Chapter 16.↩︎
- 2.1 Layers beneath the prompt
- 2.2 The instruction tree
- 2.3 Beyond a single prompt
- 2.4 Learning, traces, and authority
- 2.5 Diagnosing the environment