2.3 Beyond a single prompt
2.3.1 Skills are procedural memory
A skill is a reusable procedure. The open Agent Skills specification defines one as a directory containing a SKILL.md file with metadata and instructions, optionally accompanied by scripts, references, and assets, and it recommends progressive disclosure — expose the name and description first, load the instructions when the skill activates, and load supporting material only when needed.17 Claude Code describes skills as reusable workflows that load on demand, OpenClaw as Markdown instruction files that teach an agent how and when to use tools, and Codex as reusable workflows with supporting resources and invocation policies.18
Set against the neighboring concepts, the distinctions are clean:
| Thing | What it does |
|---|---|
| Prompt | States the current request |
| Skill | Encodes a repeatable procedure |
| Reference | Supplies supporting knowledge or evidence |
| Template | Constrains the shape of the output |
| Script | Performs deterministic work |
| Tool | Exposes an action or data source |
| Policy | Defines what is allowed or forbidden |
| Hook | Intercepts a lifecycle event |
Maya’s “read customer feedback” skill might say:
- Load the original support conversations, not only the weekly summary.
- Group cases by the customer’s underlying problem, not by the tag assigned by support.
- Preserve representative quotes and links to the source records.
- Separate observed evidence from proposed explanation.
- Identify contradictions and return them rather than smoothing them away.
- Produce a structured report with affected service, confidence, unresolved questions, and recommended next investigation.
That is more durable than a prompt because it can be evaluated, revised, versioned, and invoked by several systems. It is also more dangerous than a prompt, for exactly the same reason: it gets reused when the person who wrote it is not in the room. A skill is procedural memory, and procedural memory needs an owner, a review cycle, and eventually a retirement date.
One design choice in that specification deserves its own explanation, because it was born from economics. Progressive loading exists because loading the full instructions of every available skill into every request burns tokens — a system with thirty skills would pay for thirty skill files on every call, whether it used them or not. The consequence is that the description carries real weight. It is the only thing an inference engine sees when deciding whether to use the skill, so it has to be descriptive enough to give the model the context for that decision — without being so broad that the skill triggers when it should not. There is a craft to a well-made skill, and most of the craft lives in the parts the model reads before it commits.
2.3.2 Plugins package capability with instructions
A plugin is not simply a bigger skill. A plugin can package skills together with tools, channels, model providers, hooks, MCP servers, credentials, user-interface components, or arbitrary runtime code, and OpenClaw’s documentation draws the distinction directly: tools are callable actions, skills teach agents how to work, and plugins add runtime capabilities.1920
The difference shows up concretely in Maya’s system. A skill might teach the agent how to inspect a pull request, while a plugin adds the pull-request tool, authenticates it against GitHub, registers a post-tool hook, and exposes a skill explaining the review workflow. The skill is a procedure; the plugin is an installable capability package.
Installing one therefore changes considerably more than the agent’s vocabulary. It may change what the agent can see, what it can call, which events can trigger code, and which credentials are in reach. OpenClaw recommends inspecting the live runtime after installation rather than trusting that a package manifest describes what is actually active, and OpenAI likewise describes plugins as bundles that may contain skills, connectors, MCP servers, hooks, and presentation assets.21
Which means the questions are not optional. What does this package add? What code runs? Which credentials does it request? Which tools become visible? Which hooks can block or modify behavior? Who approved the source? How do we remove it? Installing a plugin without understanding what it does is a common mistake, and a consequential one — and the format itself is part of the problem. A plugin rolls a lot of functionality into one unit: tools, skills, hooks, credentials, provider configurations, sometimes arbitrary code. A bundle that size is difficult to review as a unit, and installing one is closer to adding a dependency that runs with production credentials than to adding a browser extension. That is the level of scrutiny it deserves.
2.3.3 Hooks are not suggestions
A hook runs because an event happened, not because the model remembered to ask for one, and that single property is what separates it from every instruction discussed so far. Claude Code documents hooks at session start, prompt submission, before and after tool calls, permission requests, subagent start and stop, task completion, compaction, and other lifecycle events, and a hook can inspect the event, add context, return a decision, or block an action outright.22
Compare the two mechanisms on the same requirement. A skill can say run the security scan before opening the pull request; a pre-commit hook or CI gate can refuse the pull request when the scan has not passed. That is the difference between guidance and enforcement, and Maya can put it to work with a hook that checks every generated pull request for a change to authentication or authorization, a new external dependency, access to customer data, a database migration, a disabled test, or a change to the deployment configuration. The hook does not need the model to understand why any of those matter. It stops the workflow and asks for a human.
None of which makes hooks automatically safe. A hook can block legitimate work, leak sensitive event data, or quietly become an unreviewed execution path of its own. The point is that hooks operate at a different layer from instructions and should be reviewed as both code and policy.
One more point, because it is easy to assume that instructions have to live where the work happens. They do not. A hook does not need to be distributed to every repository, every machine, and every workstation — for some checks, distributing it makes no sense at all. A security check or a compliance verification that every change must pass is often better built into a continuous integration and continuous deployment server, where it runs once, centrally, against every change, and where a single owner can update it without touching a hundred repositories. That is where hooks become genuinely powerful: global checks, scripts triggered in response to a change, verification of compliance or operational sanity — all enforced from one place. The general principle is that the various kinds of instructions you pass to a model can be assembled from several different sources. Do not assume they need to be bundled with the source code, or with the system being deployed.
2.3.4 Memory is part of delegated intelligence
Chapter 3 established memory as a defining property of delegated intelligence rather than a convenience feature, and the previous chapter surveyed the products that supply it. What this chapter adds is the instruction-layer view: when an orchestrator delegates intelligence they are not asking a model for one more output; they are teaching a bounded system how to work on a class of problems, and memory is part of the teaching.
Modern agentic systems generally include some form of this persistence. The research literature defines that capability more precisely than “storing notes”: memory in an LLM-based agent is the set of mechanisms that acquires, stores, retains, and retrieves knowledge and experience to support the agent’s actions — four operations, and storage is only one of them.23
Two clarifications keep this from going wrong. First, memory is not the same as context: context is what the runtime sends to the model for this run, while memory is information retained somewhere so it can be selected for a later run — and the retrieval step is where things break, because a system may hold thousands of remembered facts and include only a small and possibly badly chosen subset in any given context. Second, memory is not automatically authoritative. Maya’s system may remember that a particular support tag was unreliable last quarter, which is useful context and not permission to ignore new evidence. A temporary exception recorded during an incident should not quietly harden into a permanent rule, and a statement copied out of an untrusted customer message should not become organizational memory merely because an agent summarized it neatly.
So memory needs provenance, ownership, freshness, and a route to correction or removal. The orchestrator is not only teaching a system what to remember. They are deciding what qualifies as knowledge and what has to stay a hypothesis, a warning, or a discarded observation.
And it is more complicated than storage, because what a system stores is not what it keeps. Memory has to be distilled. Raw episodes get summarized into shorter records, recurring patterns consolidate into durable facts, and stale entries age out — the memory research calls this consolidation, moving detail from episodic form into semantic form. Some systems run that maintenance on a schedule: Chapter 4 called this dreaming, and it is becoming standard practice, because a system that keeps everything eventually drowns in its own history. The distillation step is where both the value and the danger live. What survives the long run is a design decision, which is why Chapter 4 put the curation step at the center of memory management rather than treating memory as a bucket.
Agent Skills, “Specification,” https://agentskills.io/specification. The specification defines
SKILL.md, metadata, optional scripts/references/assets, progressive disclosure, and validation.↩︎OpenAI, “Skills & Plugins” and “Build skills,” https://learn.chatgpt.com/docs/skills-and-plugins and https://learn.chatgpt.com/docs/build-skills. The documentation describes skills as reusable workflows with supporting resources, invocation behavior, and plugin packaging.↩︎
OpenClaw, “Overview,” https://docs.openclaw.ai/tools. The documentation distinguishes tools as callable actions, skills as workflow instructions, and plugins as runtime extensions.↩︎
OpenClaw, “Plugins,” https://docs.openclaw.ai/tools/plugin. The documentation describes plugins that add tools, skills, channels, model providers, hooks, and other runtime capabilities, along with install policy and live-runtime verification.↩︎
OpenAI, “Skills & Plugins” and “Package your plugin,” https://learn.chatgpt.com/docs/skills-and-plugins and https://developers.openai.com/plugins/build/plugins. The documentation describes plugins as installable bundles that can contain skills, connectors, MCP servers, hooks, and presentation assets.↩︎
Claude Code, “Hooks reference,” https://code.claude.com/docs/en/hooks. The documentation describes lifecycle-triggered commands, HTTP endpoints, MCP tools, and prompts that can add context or return decisions around agent events.↩︎
Zeyu Zhang, Quanyu Dai, Xiaohe Bo, Chen Ma, Rui Li, Xu Chen, Jieming Zhu, Zhenhua Dong, and Ji-Rong Wen, “A Survey on the Memory Mechanism of Large Language Model-based Agents,” ACM Transactions on Information Systems 43, no. 1 (2025): 1–47, https://dl.acm.org/doi/10.1145/3748302 (preprint: arXiv:2404.13501). The survey defines agent memory by its operations — acquiring, storing, retaining, and retrieving information to support actions — and distinguishes episodic, semantic, and procedural memory forms, including the consolidation of recurring episodic detail into durable semantic knowledge.↩︎