1.10 Task agents

Book 2 · The Delegation ContractChapter 1 · section 10 of 14

Task agents combine an inference engine with context, tools, state, and a loop — observe, interpret, choose, act, inspect the result, and then continue, hand off, request approval, or stop. A coding agent that reads a repository, edits a branch, runs tests, and opens a pull request is a task agent, and so is a research agent that gathers sources and returns an evidence packet. A classifier may choose a label with no tools at all, and an evaluator may inspect a result while being forbidden to change it. Because the word agent covers all of that, the contract matters considerably more than the label — which is the subject of the next section.

A note on the name, because the reader deserves the truth about it: “task agent” is this book’s label, not the market’s. The industry calls every one of these things an “agent” and is done with it, and there is no product category called “task agents” anywhere.40 The plain term is useful anyway, because the market’s word is the problem the next section describes. What it points at is real: a bounded unit of delegated work, running inside some environment, against a contract. And to answer the question directly: a task agent is not a harness, and a harness is not a task agent. Hermes Agent, Goose, Claude Code, Cursor — those are harnesses and runtimes, the environments agents run inside (Section 5.11). The task agent is one bounded unit of work running inside one of them, the way a job is not the operating system it runs on.

1.10.1 A task agent, concretely

Strip the abstraction and a task agent has six parts: an objective (what result, for whom), the context it may read (files, records, search results), the tools it may call, the permissions bounding those calls, an output format, and stop conditions — what makes it finish, hand off, or ask a person. Those six parts are the contract, and the contract is what an orchestrator actually designs. Capability comes and goes with model releases; the contract is yours.

One unit of work, start to finish, as an illustrative Devin run. A developer assigns Devin, Cognition’s cloud coding agent, a failing GitHub issue: the checkout page crashes when the cart is empty. Devin clones the repository, reproduces the crash, plans the fix, edits the files, runs the tests, and delivers a pull request with the fix — then stops.41 Nothing about that run is mysterious. The objective was the issue. The context was the repository. The tools were file edits and the test runner. The boundary — that it could open a pull request but not merge one — depends on repository permissions and explicit configuration. The output was the diff. The stop condition was the deliverable in hand; whether the tests actually pass is reported, not guaranteed. Every product in this section is a variation on that loop with different parts emphasized, and the orchestrator’s job is knowing, for each one, where the six parts are written down.

1.10.2 The products

The current field, by what the agent’s contract emphasizes:

  • Devin (Cognition). The purest delegation: assign work through a web interface or Slack and get back a pull request. It runs in the cloud on managed infrastructure, billed by compute units rather than seats, because the product really is the agent rather than the tooling around it.
  • Claude Code (Anthropic). The terminal-borne version: an agent that reads the repository, edits files, runs commands, and opens pull requests from the command line.
  • Codex (OpenAI). Cloud coding tasks, multiple in parallel, each in its own environment, each delivering a branch and a pull request for review.
  • GitHub Copilot’s coding agent. Issue-to-pull-request inside GitHub itself: assign an issue and the agent opens a draft PR with a working branch — the shortest path from ticket to code in the set, because it never leaves the platform the ticket lives on.
  • Cursor’s agents. Editor and cloud agents, including background runs that produce merge-ready pull requests; the platform drift around them is the next section’s subject.
  • Deep research agents (OpenAI, Google, Perplexity). The non-coding shape of the same idea: given a question, the agent searches, reads, cross-checks, and returns a cited report — an evidence packet instead of a diff, and a reminder that task agents are not a coding phenomenon.
  • OpenClaw. The personal always-on version: wakes on a schedule, checks the things it was told to check, escalates what matters (Section 5.3.2 covered the economics of exactly this pattern).
  • The minimal end. A classifier that labels a support ticket is one model call with a schema — also a task agent, with a tiny contract. An evaluator that critiques another system’s output while being forbidden to change it — likewise. Small contracts are still contracts, and these are the two shapes most of an orchestrated system’s volume is made of.

1.10.3 Where task agents show up for an orchestrator

Task agents are the units an orchestrator actually deploys, evaluates, and stops, so the encounter is the job itself. The recurring decisions: writing the contract, which is the task agent’s real definition, not its label; evaluating — building the set of cases a task agent must pass before it runs unattended, and re-running that set when the model, the tools, or the instructions change; and supervising — the monitoring that notices a stuck loop, a budget overrun, or an agent quietly doing something adjacent to its task.

The ladder of Section 5.2.1 gives the placement rule: every task agent sits on one or more rungs, and the orchestrator’s question is which rung this agent actually needs. Chapter 7 makes this the architecture question — which responsibilities become agents and which stay as tools — and Chapter 9 collects the failures. Every deployment still comes down to two questions: what is this agent’s contract, and can the system prove it was honored?


  1. The industry term is simply “agent” or “AI agent” — see IBM, “What Are AI Agents?” https://www.ibm.com/think/topics/ai-agents, and Wikipedia, “AI agent,” https://en.wikipedia.org/wiki/AI_agent, which notes there is no universally agreed definition; NIST’s AI Agent Standards Initiative (announced August 2026) standardizes on “agent” as well. “Task agent” is this book’s neutral label for a bounded unit of delegated work, chosen precisely because the market’s word covers everything and therefore nothing.↩︎

  2. Product documentation: Devin, https://devin.ai; Claude Agent SDK, https://www.anthropic.com/news/agent-sdk; OpenClaw, https://docs.openclaw.ai; and the OpenAI and GitHub product documentation cited in the text. Comparative surveys: daily.dev, https://daily.dev/blog/best-ai-coding-agents-comparison, and Braintrust, https://www.braintrust.dev/articles/best-ai-coding-tools-2026. Capabilities are vendor claims; the six-part contract framing is this book’s.↩︎