6.5 A worked month: the four-agent topology, itemized
The percentages above need a bill. Here is one, itemized, for the delegated topology Chapter 3 defined and Maya’s team in Chapter 1 would recognize: four agents doing the ongoing work of a mid-sized internal operation — triage and classification, retrieval and drafting, review support, and a heartbeat monitor — plus one human-reviewed decision step. The arithmetic uses the September 2026 list prices Chapter 5 footnotes, before caching and batch discounts, so it overstates the real bill; treat it as a ceiling, not an estimate of best practice.
The workload: 12,000 classification calls, 8,000 retrieval-and-draft cycles, 4,000 review-support summaries, and 30,000 heartbeat checks a month. The model mix is the routing pattern this book keeps recommending, made explicit: 95 percent of tokens on a cheap model at $10 per million input and $10 per million output, 5 percent routed to a frontier model at $50/$75 for the steps where judgment pays for itself. Assume a 4 percent retry rate from malformed outputs and a 60 percent input-cache hit rate on the retrieval cycle, which is what a memory-backed system earns after its first weeks.
The arithmetic, in one pass: classification is 12,000 calls averaging 1,500 input and 300 output tokens, cheap-model only — about $270. Retrieval and drafting is 8,000 cycles averaging 6,000 input and 2,000 output tokens; with 60 percent cache hits on the input side and 95 percent of them on the cheap model, about $1,150. Review support is 4,000 summaries averaging 3,000 input and 1,500 output, half routed frontier — about $600. The heartbeat is the trap the chapter already computed: 1,000 agents checking in daily at $3.10 per call is $74,000 a month, which is why the heartbeat arithmetic gets its own paragraph in Chapter 5 and why no always-on fleet ships without routing. Retries add 4 percent across the metered categories — about $80. The monthly bill at the routed mix is roughly $2,100. The same workload with every call on the frontier model is about $28,000. The routing decision is worth $25,900 a month, and it is one table in one config file that one person owns.
Three honest footnotes to the exercise. First, the variance is the story: the same topology with a careless heartbeat (all agents, five-minute intervals, frontier model) exceeds the routed bill by an order of magnitude, which is why the boundary section below treats the interval and the model as budget decisions, not configuration details. Second, the numbers are list-rate arithmetic — real deployments negotiate, cache better than assumed, or burn more retries than planned, and the workbook in Chapter 5 is where those assumptions live. Third, the FinOps figure this chapter opened with — 98 percent of organizations now managing AI spend, up from 31 percent two years ago — describes organizations discovering this arithmetic under deadline. The orchestrator who has already run it once has the only version of the number that survives contact with a budget committee: the itemized one, with the routing column explicit.
- 6.1 What the work costs
- 6.2 Two surprise invoices
- 6.3 Assessing workloads
- 6.4 Measuring, not estimating
- 6.5 A worked month: the four-agent topology, itemized
- 6.5 Establishing boundaries
- 6.6 The constant pressure