1.6 Tool interfaces and actuation

Book 2 · The Delegation ContractChapter 1 · section 6 of 14

Tools are how a system touches the world outside its context window: reading a repository, querying a database, driving a browser, opening a ticket, running a test, sending a message. “Tools” is a generic word for a layer that in 2026 has acquired real structure.

1.6.1 The connection format

The connection format has consolidated. Function calling — the model emitting a structured call that the host executes — is available across the major vendors, and structured outputs that constrain a model’s response to a JSON schema generally use JSON Schema or schema-like definitions, but the APIs and supported keywords are not uniform: a portable contract needs a lowest-common-denominator schema plus provider-specific validation tests.26 The practical consequence for orchestration is still substantial: a tool contract can be written once and enforced deterministically — the model proposes, the schema validates, the host decides whether to run it. Above that sits the Model Context Protocol, an open standard for connecting AI applications to data sources, tools, and workflows — now with an enterprise-managed authorization extension that lets an organization’s identity provider, rather than each individual employee, decide which MCP servers may be reached and under what conditions.27 By August 2026 the public directories tracked on the order of twenty-two thousand catalog listings, with the ecosystem adding roughly a thousand a month — from the official registry through aggregators like PulseMCP and Smithery. Catalog listings are not distinct servers: after de-duplication across registries, one mid-2026 estimate put the count of unique servers closer to nine thousand four hundred.28

MCP is more than a bundle of callable functions. The protocol defines three primitives at the same level, and the distinction between them matters for orchestration. Tools are executable actions: the model decides when to call them, and a call can change something in the world. Resources are application-controlled context — files, database rows, documents that the host application reads and injects. They are often consumed without model-invoked side effects, like GET endpoints in an API answering “what is there,” but the specification does not guarantee that a given server’s implementation is read-only. Prompts are reusable templates that a person deliberately selects, not anything the model chooses on its own.29 In practice, tools dominate the ecosystem — most servers expose tools and nothing else — but the resource primitive is the one an orchestrator should watch, because it separates “what this system can read” from “what this system can do,” and that is the same read/act split the permission envelope is built on.

For a concrete sense of the ecosystem, the servers that sit at the top of every ranking are Microsoft’s Playwright server — browser automation, and the standard way an agent verifies its own web work — and the official GitHub server, which turns a coding agent into a repository citizen.30

A tool is not intelligence; it is a controlled connection to something outside the model, and its permissions decide what the system can read, change, or trigger. Which is why the same model with read-only repository access and the same model with permission to merge and deploy are two different delegated systems.

1.6.2 Actuation

Actuation — where the “agent” actually does something to a system that was never designed for agents — has split into three product families:

  • Browser automation. The open-source Playwright MCP server — browser automation exposed as MCP tools, driven through accessibility-tree snapshots rather than screenshots — has become a default install for coding agents, with tens of thousands of stars and wide client support; Browser Use, an open-source Python library for LLM-driven browsing, has become one of the most-starred agent projects on GitHub; Browserbase runs production browser fleets for agents with its Stagehand SDK; and Anthropic’s computer-use toolset and OpenAI’s computer-use tool expose full desktop control, with prompt-injection classifiers watching screenshots for hostile instructions.
  • Code execution. Sandboxed environments for running model-generated code — E2B with its Firecracker microVMs, Modal for GPU-heavy workloads, Daytona for container-based sandboxes, Fly.io for persistent stateful machines — give an agent a place to compute without giving it your laptop. The use cases now have names of their own: data analysis, vibe-coding runtimes, reinforcement-learning evaluation farms.
  • Repositories, tickets, and messaging. The MCP ecosystem’s long tail: GitHub, GitLab, Jira, Linear, Slack, Notion, database connectors, deployment systems — each one a tool server that turns a verb an agent could only describe into a verb it can actually perform.

1.6.3 What to ask of a tool layer

For the orchestrator, tooling is the layer where the permission envelope becomes concrete, because every tool is an authority grant. The survey questions follow directly: what does this tool expose, what identity does it run under, what does it log, who can revoke it — and is the tool a deterministic connection with a schema, or a probabilistic agent that decides what to do? Confusing the two is how a system ends up with an “automation” that improvises.

And the more powerful the tool, the more specific the grant has to be. When a system can execute commands — a terminal, a code-execution sandbox, a browser-automation layer like Playwright — the permission decisions are concrete ones: which commands it may run, which directories it may touch, which URLs it may visit, whether the grant lasts one call, one session, or indefinitely. The best tool layers offer that granularity; when a connector does not, enforce the constraints in a gateway, a sandbox, a network policy, or the downstream API instead. The failure mode is not only the tool lacking permission controls; it is just as often the operator granting more than the task needs, because scoping permissions is work and granting everything is fast.

Everyone approves everything. The developer tools in this chapter all ask permission, and the asking has a standard shape: Claude Code surfaces “Allow this shell command?” with the option to allow once, allow for the session, or add the command to the permanent allowlist; Cursor has the same prompt, plus a switch that stops the asking entirely. What developers actually do with those prompts is documented. Anthropic has said that Claude Code users approve 93 percent of the prompts they are shown, and a large share of the community skips the prompts altogether by running the flag whose own name is the warning — --dangerously-skip-permissions — or by turning on Cursor’s YOLO mode. The vendors know this is happening. Anthropic’s answer is auto mode, a classifier that reviews each proposed action and approves the safe-looking ones without asking — and Anthropic’s own evaluation, run on fifty-two curated real overeager actions, found a 17 percent false-negative rate: roughly one approved action in six in that evaluation set was actually dangerous. The number describes that curated set, not one in six of everything auto mode approves. Independent testing reported that Cursor’s safeguards were “easily bypassed” in YOLO mode. None of this is stupidity. It is fatigue — analogous to documented clinical alert fatigue, in which frequent low-value alarms get ignored, delayed, or switched off: after the fortieth prompt in an hour, a developer grants global permissions to make the prompts stop.31 In development the habit feels survivable, but development is not automatically a small blast radius: a local coding agent may inherit shell access, credentials, network reach, package registries, and cloud tooling, and code review can catch a bad diff without undoing a disclosed secret or an external side effect. And the habit being trained is exactly wrong for production. An orchestrated system that operates in production, that makes decisions affecting money, infrastructure, or human lives, cannot inherit the click-through culture; its permissions have to be encoded deliberately — least privilege, scoped grants, expiry, revocation — by the one person whose job it is to think about what the system should never be allowed to do. The lesson the industry is going to learn the expensive way, over the next several years, is that the approval prompt everyone kept clicking through was the governance model — and almost nobody noticed when it went away.

1.6.4 Where tools show up for an orchestrator

Tools are where an orchestrated system stops being talk and starts touching things, so this is the layer where the orchestrator’s authority decisions become literal configuration. The connection-format consolidation matters here for a specific reason: because function calling and MCP have standardized, tool choice has become composable. A tool written once as an MCP server can often work across the MCP-compatible runtimes surveyed in this chapter — subject to supported protocol versions, transports, authorization, and optional extensions — which means the orchestrator’s tool inventory can be evaluated, replaced, and revoked as a portfolio — the same lifecycle management Chapter 7 gives delegation itself. For every tool this system can call, an orchestrator should be able to answer three things without looking them up: who granted it, under what conditions, and what revoking it would break.


  1. Browser tools: Playwright MCP, https://github.com/microsoft/playwright-mcp; Browser Use, https://github.com/browser-use/browser-use; Stagehand, https://www.browserbase.com/stagehand; Anthropic computer use, https://platform.claude.com/docs/en/agents-and-tools/tool-use/computer-use-tool; and OpenAI computer use, https://developers.openai.com/api/docs/guides/tools-computer-use. Sandboxed execution providers include E2B, Modal, Daytona, and Fly.io. OpenAI, Anthropic, and Google each document schema-constrained structured output; their supported subsets differ.↩︎

  2. Model Context Protocol, https://modelcontextprotocol.io/. Open standard by Anthropic November 2024, donated to the Linux Foundation’s Agentic AI Foundation December 2025. Latest specification release July 28, 2026 (https://blog.modelcontextprotocol.io/posts/2026-07-28/); enterprise-managed authorization extension, https://modelcontextprotocol.io/extensions/auth/enterprise-managed-authorization, verified August 29, 2026. The extension routes MCP server access decisions through an organization’s identity provider with centralized revocation; documented capability, client support varies and is opt-in.↩︎

  3. MCP server ecosystem scale as of August 2026: PulseMCP statistics, https://www.pulsemcp.com/statistics (22,070 catalog listings on August 14, 2026; roughly 1,000 new servers indexed per month in early-mid 2026); Smithery (~7,300 servers as of May 2026); the official MCP registry (~9,600 servers); a de-duplicated cross-registry estimate of ~9,400 distinct servers as of May 2026 per Digital Applied, “MCP Server Ecosystem Tracker,” https://www.digitalapplied.com/blog/mcp-server-ecosystem-tracker-50-servers-cataloged-2026. The widely repeated “78% enterprise adoption” figure could not be traced to a primary source and is deliberately not used here.↩︎

  4. MCP’s three first-class primitives are defined in the specification: tools (model-controlled executable actions), resources (application-controlled contextual data — read-only by convention rather than by guarantee), and prompts (user-controlled reusable templates) — Model Context Protocol specification, https://modelcontextprotocol.io/specification/2026-07-28, sections on Tools, Resources, and Prompts (verified August 29, 2026). The observation that tools dominate in practice — most servers expose tools and nothing else — is documented in Philipp Schmid, “Model Context Protocol (MCP) an overview,” https://www.philschmid.de/mcp-introduction, and Sean Goedecke, “Model Context Protocol explained as simply as possible,” https://www.seangoedecke.com/model-context-protocol, who found that none of the public MCP server codebases he reviewed exposed prompts or resources, only tools.↩︎

  5. Popular-server rankings as of August 2026 converge on the same names. Playwright (Microsoft, https://github.com/microsoft/playwright-mcp, roughly 36K stars, the most-tracked server in the ranked list at https://github.com/tolkonepiu/best-of-mcp-servers and “frequently cited as the most popular” per Gamut, https://www.gamut.so/blog/best-mcp-servers); Filesystem and Fetch (reference servers in the MCP project, which Anthropic originally introduced; https://github.com/modelcontextprotocol/servers); GitHub (official, https://github.com/github/github-mcp-server); Context7 (https://github.com/upstash/context7); Postgres and Supabase (https://github.com/supabase/supabase-mcp); Sequential Thinking (modelcontextprotocol/servers). Ranking sources: Gamut, https://www.gamut.so/blog/best-mcp-servers; AY Automate, “Best MCP Servers 2026,” https://www.ayautomate.com/blog/best-mcp-servers-2026; Totalum, “Best MCP Servers in 2026,” https://www.totalum.app/blog/best-mcp-servers-2026. Characterizations of what people use them for are this book’s synthesis of those surveys.↩︎

  6. Anthropic, “How Claude Code auto mode was built,” https://www.anthropic.com/engineering/claude-code-auto-mode: users approved 93 percent of prompts; its curated 52-case evaluation measured a 17 percent false-negative rate, not a general production rate. Cursor’s YOLO-mode bypasses were reported by The Register, https://www.theregister.com/2025/07/21/cursor_ai_safeguards_easily_bypassed/. The comparison with clinical alert fatigue draws on AHRQ, https://psnet.ahrq.gov/primer/alert-fatigue; applying it to permission prompts is this book’s analysis.↩︎