5.6 Concrete artifacts — what exists, and why it is not enough

Book 1 · Your Next Job TitleChapter 5 · section 6 of 7

Chapter 5 surveys products; this section is about formats, meaning the files and schemas the industry is beginning to share. Use them, and do not mistake any of them for the finished governance layer.

Before going through them one at a time, it helps to know the three formats that show up over and over in orchestration work. MCP, the Model Context Protocol from the table above, is the connectivity layer. The July 2026 specification (2026-07-28) made the protocol core stateless and added multi-round-trip requests, header-based routing, cacheable discovery results, a formal extension model, and authorization hardening — evidence, along with the very large SDK download volumes the project reports, that MCP has moved quickly from an experimental integration format toward shared agent infrastructure.12 It does not decide the agent’s goal, decompose the work, manage a fleet, provide durable memory, or enforce the organization’s complete policy by itself. Agent Skills (SKILL.md with YAML frontmatter and a Markdown body) is an open standard for procedural memory: a small, versionable description of how to do one thing, loaded only when needed. AGENTS.md is a freeform convention — a Markdown file in a repository that tells an agent how to work in that repo. MCP, Agent Skills, and AGENTS.md are not competing with one another; an orchestrated system uses all three at once.

5.6.1 Skills: SKILL.md and friends

The open Agent Skills specification, with product implementations in Claude Code, Cursor, Hermes, OpenClaw, Pi, and YAML recipes in Goose, is the closest thing we have to a portable runbook for agents.13 AGENTS.md — now under the Linux Foundation’s Agentic AI Foundation alongside MCP and Goose — is a useful convention for describing how to work in a given repository, though it remains freeform Markdown with no required schema.14 Anthropic has also started shipping something called a plugin, which is related to a skill but not the same thing: a plugin bundles slash commands, subagents, hooks, MCP server configurations, and skills into one installable unit, distributed through a marketplace.15

A short warning for readers. Skills are still an emerging standard. There is very little standardization in how skills are interpreted by agents and tools from different vendors. A lot of these new primitives are released by one model provider and then catch on quickly: for a while skills were called “Claude skills” because they were first made popular by Anthropic’s Claude, and the term has only recently broadened. The descriptions in this chapter are accurate to the best of the author’s knowledge at the time of writing; treat them as a snapshot, not a stable taxonomy.

What they get right is procedure you can version, a home for “how we do this kind of work” that is not trapped inside one person’s chat history, and progressive disclosure — the name for how a skill loads its own content. The way progressive disclosure is defined by the open Agent Skills specification is that the agent sees only the skill’s name and description at startup; the full instructions load only when the skill is activated; auxiliary files load only when they are needed. The agent never carries instructions it is not using, which keeps context small and behavior focused. What they are not is a permission envelope. An allowed-tools hint in frontmatter is an experimental pre-approval signal whose enforcement depends on the runtime, not a fleet policy engine, and nothing in the skill format records who authorized widening the agent’s tool set, or links a skill revision back to the outcomes that justified it.

Chapter 3’s distinction between deterministic and nondeterministic parts matters when you build with these tools. A skill and a tool are not the same kind of thing. Tools are what an agent uses to impact the outside world: read a file, call an API, run a query, post a message. A skill is a natural-language description of a procedure — usually a Markdown file with frontmatter — interpreted by an agent when the procedure is needed. The interesting thing is what happens to a skill over time. Run the same skill over and over, and the orchestrator will often convert it into a tool — a Python script, or a function in another language — that reduces interpretation variance and can be tested under specified inputs and conditions. Skills are a good starting point because they are cheap to write and easy to revise. Tools are what you graduate to when reliability matters more than flexibility.

5.6.2 Memory: Markdown files and Memory Blocks

Hermes and OpenClaw treat memory as files: USER.md, MEMORY.md, dated notes, sometimes a dreaming or consolidation step. Letta, from the MemGPT line, treats memory as structured blocks with labels, size limits, and read-only flags that can be shared across agents.16

What they get right is persistence a human can open in a text editor or an API, and a clean separation between the facts we keep and the procedures we run. What they are not is a portable standard. Moving memory from OpenClaw’s layout into Letta blocks is a migration project rather than a checkbox, and none of these formats enforce what may be written or who may write it. OpenClaw’s own documentation is honest that its memory guidance is not policy enforcement. Today, memory is often attached to an individual agentic system: a conversation thread, a framework checkpoint, an agent-specific store, or state managed by one provider. These memory systems are at the stage the web was when a browser first offered to remember your password. Memory is powerful — it is what lets an agent get better at your particular work over time — but the industry is still developing ways to curate and inspect what ends up in it: what was written, when, by which part of the system, and whether it is still true. The ability to audit memory — to take snapshots of memory, to reconstruct what the system believed at the moment it made a decision — is important for us to develop, and there is no common, mature standard for it yet. For anything touching eligibility, identity, or medical and financial context, memory you cannot inspect or correct is a liability.

Memory is beginning to separate from the agent itself. Mem0 sits between an application and its model and preserves memory across sessions, tools, runs, and multiple agents. LangGraph distinguishes thread-scoped checkpoints from long-term stores that can be read across threads. Amazon offers AgentCore Memory as a managed service usable independently of a particular model or agent framework. These systems point toward memory as shared infrastructure rather than a private feature of one agent instance.17

Vector databases became a distinct infrastructure category when retrieval stopped being an application detail and became a shared service. Agent memory may follow the same path. At one end, systems such as Mem0 can support durable memory shared across large fleets; at the other, self-hosted and local stores can keep memory close to a user, device, or workload. The result may be a new database category with its own schemas, retention rules, permissions, retrieval strategies, consolidation processes, and lifecycle. That last sentence is a forecast, not a report — memory as an independent infrastructure layer, and especially memory at the edge, is still emerging, and none of these systems has reached the maturity or scale of an established database category. But the direction is visible.

Dreaming, or memory consolidation. A memory file that just accumulates every event the agent ever saw is not useful — it is a firehose. Many systems now run a periodic step that reviews the memory, summarizes what is still relevant, and lets older entries fall off. Some systems call this step dreaming, which is a metaphor — not an established standard — borrowed from the way human sleep consolidates and reorganizes memories overnight.18 Treat it as a name for offline consolidation. Hermes, Nanobot, and Letta all ship some version of this; Letta frames its variant as sleep-time compute and has published a paper on the technique.19 In practice the dreaming step is where you decide what an agent is allowed to remember, for how long, and under whose authority. The same governance questions you have about a decision record apply here: who scheduled the dreaming run, what was rewritten, what was discarded, and how would you reconstruct the agent’s state at a given moment in time.

The context window is finite. Memory also runs against the context window of the model the agent is talking to: a maximum amount of text in a single interaction. It limits active model input — not what can be stored outside the model — but once the limit is exceeded the agent simply cannot see anything that does not fit in the window. The job of memory management, then, is not just storage — it is curating what is preserved and what is discarded when the window is full. Most modern memory systems are designed for exactly this: a retrieval step selects the relevant subset of memory and hands it to the model inside the context window, and the rest stays out of sight. As systems run longer and accumulate more memory, the curation step becomes the difference between an agent that remembers what matters and one that drowns in its own history. This is one of the reasons dreaming is becoming standard practice: it is how systems control what survives the long run.

5.6.3 Evals: datasets and Inspect tasks

LangSmith-style datasets, OpenAI’s agent-eval surfaces, and Inspect AI from the UK AI Security Institute all give you reproducible ways to ask whether an agent still behaves on the cases you care about.20

What they get right is treating regression as engineering, with versioned datasets, scorers that can fail a build, and repeated runs that record versions, thresholds, and observed variance rather than promising strictly reproducible scores from a stochastic model. What they are not is an operating record for what actually happened to the people the system touched. Passing an offline suite does not establish that last Tuesday’s run was carried out inside policy; it only shows the model still knows how to handle the cases in the suite.

5.6.4 Audit and traces: OTel GenAI, OpenInference, Temporal history

OpenTelemetry’s generative-AI semantic conventions and Arize’s OpenInference conventions define a vendor-neutral shape for trace data, and Temporal’s workflow event history supports deterministic replay and resumption of workflow logic from recorded events and activity results — completed activities are not rerun to replay the code, and arbitrary external side effects still have to be accounted for outside the history.21 Product tracers — Langfuse, LangSmith, Phoenix, OpenAI’s trace UI — sit on top of those.

What they get right is that you can stop inventing a private log format per framework, and you can finally see tool calls and handoffs. What they are not is a complete audit standard for what the agent decided and why. The GenAI conventions remain in development status and attribute churn is real. A trace can show every step the system took. It does not record which policy verdict allowed the step, which human owned the decision, or what happened in the world afterward — whether the action the system took actually had the result it intended. That is the difference between debugging and governance, and it is the gap this book keeps returning to. Chapter 8 returns to that gap and Chapter 10 turns it into meeting evidence.

5.6.5 Permission: MCP and ACS

The Model Context Protocol specification is explicit that MCP is connectivity, not enforcement: it standardizes how tools connect, not what they are allowed to do. Its HTTP authorization profile is real OAuth work, but the spec cannot enforce security principles at the protocol level, and a good many implementors will not. And publishing an MCP server does not make it available to every agent: the agent needs a compatible host or client, explicit configuration, supported capabilities, and the authentication and authorization the deployment requires.

ACS is the Agent Control Specification, a Microsoft-led public-preview/beta specification and policy-decision interface from the Agent Governance Toolkit for portable permission envelopes around agents. Where MCP says “here is how to call a tool,” ACS says “here is how to decide whether this agent is allowed to call this tool, in this situation, under this policy, with this verdict.” An ACS manifest describes an agent’s allowed scope; intervention points let you check the agent’s intent before risky calls; policies written in Rego or Cedar return a verdict such as allow, deny, warn, escalate, or transform. Its verdict matters only when the host or runtime invokes it and enforces the result; the specification itself does not reach into your deployment. NeMo Guardrails and Guardrails AI’s RAIL sit closer to conversational and output-structure rails; the IETF drafts on agent authorization and agent identity are early enough that you cannot hand one to a regulator and call it done.22

One more thing before the format tour goes further, because agent security is often reduced to key storage and it is bigger than that. Each agent needs a distinct identity, narrowly scoped authority, explicit boundaries around the tools, data, network locations, and side effects it may reach, and a way to revoke that authority. Tokens must be audience-bound rather than blindly passed to downstream services. Tool descriptions, tool results, retrieved documents, and web pages must be treated as potentially hostile, because indirect prompt injection can turn legitimate capabilities against the system. Irreversible actions need stronger confirmation, while code and browser execution need isolation, input and output validation, rate limits, and a durable record connecting the requesting identity to the action and result.23

What these standards get right is that tools become typed edges, policy can live outside the model, and output can be structurally constrained. What they are not is the full loop this book keeps naming: ownership, delegation lifecycle, learning policy, fleet-wide envelopes, outcome linkage.

5.6.6 Topology: A2A Agent Cards

A2A, originally from Google and now open, defines Agent Cards and a task/message/artifact exchange so that agents can discover one another and collaborate across servers.24

What it gets right is the horizontal edge: who else exists, what skills they advertise, and how to hand work across frameworks. What it is not is authority. An Agent Card advertises claimed capabilities and security requirements; discovery alone does not verify them. Discovering that a downstream action agent exists tells you nothing about whether your intake agent may invoke it, under whose budget, or with which escalation when something goes wrong downstream.

5.6.7 The long-term bar

This book is aimed at systems that operate under human authority for years. For that, interoperability is not enough. You need an inspectable chain:

owner → permission envelope → escalation contract
      → learning policy (what may self-edit)
      → audit record (proposal → verdict → actor)
      → outcome (what happened to the customer or system)

Today’s standards cover the middle of that chain — tools, skills, spans, emerging policy manifests — but do not yet form one integrated orchestration standard, and the two ends, where ownership registries, learning policies, and outcome records belong, are the thinnest. So use the available pieces now, prefer the open ones you can inspect and swap, and write contracts yourself to cover the gaps. Choosing open-weight models, open harnesses, and open audit formats where they fit the job keeps an exit ramp open, and Chapter 5 returns to that preference under the name O’Reilly gave it: the architecture of participation.


  1. Model Context Protocol specification, 2026-07-28, https://modelcontextprotocol.io/specification/2026-07-28 (stateless core, multi-round-trip requests, header-based routing, cacheable list results, authorization hardening, extension framework); release notes, “The 2026-07-28 Specification,” https://blog.modelcontextprotocol.io/posts/2026-07-28/; authorization profile, https://modelcontextprotocol.io/specification/2025-11-25/basic/authorization. The download figure is the project’s own, from the release notes: “close to half-a-billion downloads a month” across Tier 1 SDKs, with the TypeScript and Python SDKs each past one billion total downloads — a project-reported number, not independently audited. Cited for tool/resource connectivity and OAuth patterns; the specification’s own security posture is that MCP cannot enforce implementor security principles at the protocol level.↩︎

  2. AGENTS.md, https://agents.md/. Convention for repository-level agent instructions; freeform Markdown without a required schema. Stewardship associated with the Linux Foundation Agentic AI Foundation (AAIF) founding projects.↩︎

  3. Linux Foundation, “Linux Foundation Announces the Formation of the Agentic AI Foundation,” https://www.linuxfoundation.org/press/linux-foundation-announces-the-formation-of-the-agentic-ai-foundation; OpenAI, “Agentic AI Foundation,” https://openai.com/index/agentic-ai-foundation/. Cited for stewardship of MCP, goose, and AGENTS.md — not as a claim that AAIF defines authority, memory portability, or audit of decisions.↩︎

  4. Anthropic, “Use plugins in Claude,” https://support.claude.com/en/articles/13837440-use-plugins-in-claude (plugin definition and marketplace behavior); “Claude Code Plugins” (official), https://github.com/anthropics/claude-plugins-official (manifest structure, .claude-plugin/plugin.json). Plugins bundle slash commands, subagents, hooks, MCP server configurations, and skills into one installable unit distributed through a marketplace. Skills inside a plugin work across Claude surfaces; hooks and subagents run only where the runtime supports them.↩︎

  5. OpenClaw, “Memory,” https://docs.openclaw.ai/concepts/memory (workspace Markdown: USER.md, MEMORY.md, dated notes; docs note memory does not enforce policy); Hermes Agent memory/skills documentation under https://hermes-agent.nousresearch.com/docs/; Letta, “Memory Blocks,” https://docs.letta.com/guides/core-concepts/memory/memory-blocks/. Cited for concrete persistence layouts and their governance limits.↩︎

  6. Mem0, “Introduction,” https://docs.mem0.ai/introduction, and “How It Works,” https://docs.mem0.ai/core-concepts/how-it-works (memory layer between application and model, persisting across sessions, tools, runs, and multiple agents; managed and self-hosted deployments, https://docs.mem0.ai/platform/platform-vs-oss); LangGraph memory overview, https://docs.langchain.com/oss/python/concepts/memory (thread-scoped checkpoints versus long-term stores read across threads); AWS, Amazon Bedrock AgentCore prescriptive guidance, https://docs.aws.amazon.com/prescriptive-guidance/latest/agentic-ai-frameworks/amazon-bedrock-agent-core.html (AgentCore Memory as a managed service usable independently of a particular model or agent framework). Cited as evidence that memory is separating from individual agent runtimes — not as a claim that any has reached the maturity of an established database category.↩︎

  7. UC Irvine News, “Dreaming is linked to improved memory consolidation and emotion regulation,” https://news.uci.edu/2024/05/13/dreaming-is-linked-to-improved-memory-consolidation-and-emotion-regulation; Nature Scientific Reports, “Evidence of an active role of dreaming in emotional memory processing,” https://www.nature.com/articles/s41598-024-58170-z. Cited for the metaphor that motivates agent “dreaming” steps — human dreaming appears to consolidate and reorganize memories during sleep.↩︎

  8. Letta, “Sleep-Time Compute,” https://www.letta.com/blog/sleep-time-compute (research paper and Letta 0.7.0 release shipping sleep-time agents); Fast Company, “Why sleep-time compute is the next big leap in AI,” https://www.fastcompany.com/91368307/why-sleep-time-compute-is-the-next-big- leap-in-ai. Cited as one of the names for periodic memory consolidation in modern agent runtimes.↩︎

  9. LangSmith dataset JSON types and evaluate APIs, https://docs.langchain.com/langsmith/; Inspect AI, UK AI Security Institute, https://inspect.aisi.org.uk/; OpenAI, agent evals / trace grading guides under https://developers.openai.com/. Cited for regression and capability evaluation — not as complete production decision audit.↩︎

  10. OpenTelemetry generative AI semantic conventions, https://github.com/open-telemetry/semantic-conventions-genai (Development status; opt-in stability flags apply); OpenInference semantic conventions, https://arize-ai.github.io/openinference/spec/semantic_conventions.html; Temporal on AI / durable agents, https://temporal.io/solutions/ai. Cited for trace and durable-execution substrates — not as a finished audit standard.↩︎

  11. Microsoft Agent Governance Toolkit — Agent Control Specification, https://github.com/microsoft/agent-governance-toolkit/blob/main/policy-engine/spec/SPECIFICATION.md (versioned manifests, intervention points, allow/deny/warn/escalate/transform). Cited as the closest open portable policy-envelope artifact at time of writing; still early/beta relative to long-term fleet governance needs.↩︎

  12. MCP security best practices, https://modelcontextprotocol.io/docs/2025-11-25/tutorials/security/security_best_practices (per-client consent, confused-deputy attacks, token audience validation, prohibition of token passthrough); MCP 2026-07-28 specification, https://modelcontextprotocol.io/specification/2026-07-28 (implementers must address explicit consent, data control, and arbitrary data-access and code-execution paths); NIST NCCoE, “Accelerating the Adoption of Software and AI Agent Identity and Authorization,” concept paper, February 2026, https://www.nccoe.nist.gov/sites/default/files/2026-02/accelerating-the-adoption-of-software-and-ai-agent-identity-and-authorization-concept-paper.pdf; OWASP Top 10 for Agentic Applications, 2026, https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/ (tool misuse, identity and privilege abuse, memory poisoning, insecure inter-agent communication, cascading failures — risks well beyond secrets management).↩︎

  13. Agent2Agent (A2A) Protocol specification, https://a2a-protocol.org/v1.0.0/specification/ (Agent Cards, Task/Message/Artifact). Cited for discovery and agent-to-agent task exchange — not for authority ownership after handoff.↩︎