6.5 Establishing boundaries

Book 2 · The Delegation ContractChapter 6 · section 6 of 7

Boundaries are set before the system runs, not after the bill lands.

  • Spend limits per task, per day, per environment. A hard ceiling per task catches runaway loops — the agent that retries forever, the fan-out that spawned two hundred sub-agents. A daily ceiling catches drift. A lower ceiling in the test environment than in production makes the staging system structurally incapable of the production bill.
  • Model routing as a standing decision. Cheap steps run on cheap models: classification, routing, summarization, heartbeat checks. Judgment runs where judgment is needed. Chapter 5’s routing example is the pattern — five percent frontier, the rest cheap, a 94 percent reduction. The routing table is where the architecture becomes an economic decision, and the orchestrator owns the routing table.
  • Kill switches tied to budgets, not just errors. Every orchestrated system has an error kill switch. The budget kill switch is the one that matters more: the rule that pauses the workload — and pages a person, not a model — when spend crosses the line. Cost, like failure, is a condition the system must be able to detect about itself.
  • The boundary the orchestrator does not delegate. The system optimizes its task; the orchestrator optimizes the spend. A system told to answer all employee questions will answer all of them at whatever quality the routing allows, and it has no reason to notice that the answering costs four hundred dollars a month rather than twenty-seven. Cost pressure is not delegable, because no agent in the system holds the budget — the budget is held by the organization, and the orchestrator is the standing interface between the two.

6.5.1 Budget as a permission

Chapter 6 drew the line this book keeps returning to — text is guidance, credentials are rules — and cost belongs on the credentials side of it. A system-prompt sentence such as do not spend more than five dollars on this task is advice to a component that cannot see its own bill. A budget the harness or the account enforces stops the run whether or not the model agrees. As of September 2026 the enforcement points look like this; I checked each against its documentation.

At the run level, the harnesses count. OpenAI’s Agents SDK takes a max_turns argument on every Runner entry point and raises MaxTurnsExceeded when the loop exceeds it; None disables the limit, and an error handler can turn the exception into a controlled final output.5 LangGraph counts super-steps against a recursion_limit passed in the run config and raises GraphRecursionError at the ceiling; since version 1.0.6 the default is a thousand steps.6 Claude Code’s non-interactive mode has --max-turns, which has no default limit and exits with an error when it is reached, and — the one that counts in dollars — --max-budget-usd, which stops the run when API spend reaches the amount, counts sub-agent spend against the same cap, refuses to spawn another sub-agent once the cap is hit, and stops background sub-agents still running.7 Two of the three count turns, a proxy for money that breaks the moment one turn re-reads a million-token context. One counts dollars. All three stop one run; none knows about the other four hundred the scheduler started the same hour.

At the account level, the providers count in dollars, monthly, and with different teeth. OpenAI distinguishes a spend alert, which notifies while traffic continues, from a hard spend limit, which makes affected requests fail with a 429 and an organization_spend_limit_exceeded or project_spend_limit_exceeded code; the limit is monthly, per organization or per project, and enforcement is not instantaneous, so recorded spend can slightly exceed the cap.8 Anthropic’s API tiers carry monthly spend caps of $500, $1,000, and $200,000, with none on custom tiers; an organization can set its own limit below the cap, per organization or per workspace, and requests past it return a 400 beginning “You have reached your specified API usage limits.”9 OpenRouter, the routing service Chapter 5 used to read the market, enforces an account balance and an optional per-key credit cap, readable as limit, limit_remaining, and limit_reset, and returns 402 when either is exhausted.10 The per-key cap is the most useful of the lot, because a key is something you can hand to one workload and nothing else.

At the workflow level, durable execution engines can bound time. Temporal’s Workflow Execution Timeout caps how long a workflow may stay open including retries; it defaults to infinite, and Temporal’s documentation says it generally does not recommend setting it, because workflows are designed to be long-running.11 That is a fair position for an order-fulfillment workflow and a dangerous default for an agent loop, and knowing which one is running is the orchestrator’s job.

Now the complication, because every one of those controls is blunt in a way that matters. A monthly cap is a fence around the wrong interval: a loop that starts Friday night can spend the month by Monday and stay inside the rule. A turn count is a fence around the wrong unit. A hard limit at the organization level stops the good traffic with the bad — which is why OpenAI’s documentation warns that hard limits “can interrupt production traffic,” and why Anthropic’s guidance for its Enterprise plan recommends group and per-user limits over an organization-wide one, “without the risk of cutting off your entire org if a limit is hit.”12 None of that argues against the controls. It argues for layering them: a per-run dollar cap in the harness for the loop, a per-key cap at the gateway for the workload, a per-project or per-workspace cap at the provider for the team, and the organization cap as the fence you hope never to reach. Each layer catches what the one inside it cannot see.

What the layering buys is that the budget kill switch above becomes a real object with a real owner. The budget_id in the cost record points at one of these limits; the budget_owner is the person who set it, who can raise it, and who gets paged when a run reports stopped_by_limit. That is the shape of a credential: granted by a named person, enforced by a system that does not read prose, revocable without consulting the component. A budget in a prompt has none of those properties. An orchestrator who has written one and stopped there has documented an intention, not built a boundary.


  1. OpenAI Agents SDK documentation, “Running agents,” https://openai.github.io/openai-agents-python/running_agents/ (accessed September 2026): the loop raises MaxTurnsExceeded when max_turns is exceeded; max_turns=None disables it; error_handlers keyed on "max_turns" can return a controlled final output instead. Counts agent-loop turns (LLM calls), not tokens or dollars.↩︎

  2. LangGraph documentation, “Graph API,” https://docs.langchain.com/oss/python/langgraph/graph-api (accessed September 2026): recursion_limit is a standalone config key, counts super-steps, raises GraphRecursionError when exceeded, and defaults to 1,000 steps from version 1.0.6. A step limit, not a spend limit.↩︎

  3. Anthropic, Claude Code documentation, “CLI reference,” https://code.claude.com/docs/en/cli-reference (accessed September 2026). --max-turns: print mode only, no default limit, exits with an error when reached. --max-budget-usd: print mode only; sub-agent spend counts toward the cap; at the cap, spawning a sub-agent fails with Budget limit reached and background sub-agents are stopped; requires v2.1.217 or later. No equivalent is documented for interactive sessions.↩︎

  4. OpenAI, “Spend limits,” https://developers.openai.com/api/docs/guides/spend-limits (accessed September 2026): spend alerts notify only; hard limits return 429 with organization_spend_limit_exceeded or project_spend_limit_exceeded; “Enforcement is not instantaneous, so recorded spend can slightly exceed the configured amount”; hard limits “can interrupt production traffic.” A separate OpenAI-assigned monthly usage limit by tier ($100 to $200,000) is at https://platform.openai.com/docs/guides/rate-limits. Vendor documentation; the behavior has changed before.↩︎

  5. Anthropic, Claude Platform documentation, “Rate limits,” https://platform.claude.com/docs/en/api/rate-limits (accessed September 2026): tier spend caps (Start $500, Build $1,000, Scale $200,000; Custom uncapped); at the tier cap, 429 with enforced_spend_limit_reached and no retry-after; self-set limits at organization or workspace level return 400 invalid_request_error with the quoted message. The API is prepaid (https://support.claude.com/en/articles/8977456), so an exhausted balance is itself a stop unless auto-reload is on. Vendor documentation.↩︎

  6. OpenRouter documentation, “Limits,” https://openrouter.ai/docs/api-reference/limits (accessed September 2026): account balance plus optional per-key credit caps; GET /api/v1/key returns limit, limit_reset, limit_remaining, and usage; exhaustion returns 402. One routing service; other gateways expose similar controls under other names.↩︎

  7. Temporal documentation, “Detecting Workflow failures,” https://docs.temporal.io/encyclopedia/detecting-workflow-failures (accessed September 2026): Workflow Execution Timeout (default infinite, includes retries and Continue-As-New), Workflow Run Timeout (single run), and the statement that Temporal “generally do[es] not recommend setting Workflow Timeouts, because Workflows are designed to be long-running and resilient.” Chapter 16 covers Temporal’s event history; this note is only about the timeouts.↩︎

  8. Anthropic, Claude Help Center, “Claude Enterprise consumption guide,” https://support.claude.com/en/articles/14782391-claude-enterprise-consumption-guide (accessed September 2026). Concerns the Claude Enterprise seat product, not the API; cited for the design reasoning, which transfers, not as a description of API controls.↩︎