6.4 Measuring, not estimating
The cost report is an orchestration artifact, the same way the evidence packet is: something the orchestrator produces as part of operating the system, not a response to an accusation from finance. The people who will ask for it already exist. The FinOps Foundation’s 2026 survey of 1,192 practitioners — the people who manage cloud bills for a living — found 98 percent now managing AI spend, up from 31 percent two years earlier, named AI cost management the skill their teams most want to add, and recorded many being asked to self-fund AI investment out of savings found elsewhere.3 The survey does not say how many of those teams can attribute a token to a workload; my experience is that most cannot. The report has four properties:
- Per-workload attribution. Every run is tagged to the workload that caused it — the same tagging discipline as the trace, because unattributed spend is unmanageable spend.
- Cost per completed task, not per attempt. Attempts are cheap to count and meaningless to budget. The unit of account is the finished thing: the answered question, the triaged ticket, the proposed change. Chapter 5’s worked example — the internal documentation agent whose month costs about five dollars on a cheap model and four hundred fifty on a flagship, and about twenty-seven when five percent of the questions route to the frontier and the rest route cheap — is the whole method in one line.4
- The multiplication in the footnote. Any worked cost example shows its arithmetic: the token volume, the price, the product. Chapter 5’s heartbeat math is the model. A 300,000-token read at $10 per million is $3.00; a 2,000-token reply at $50 per million is ten cents; $3.10 a call, times 24 calls a day, is $74.40 — call it seventy-five — and times 1,000 agents is about $74,000 a day. Every step is on the page, because a cost claim a reader cannot re-derive is a number the reader cannot trust.
- A trend, not a snapshot. One month’s bill says nothing. The report that matters shows the cost per completed task moving — because the routing got smarter, because memory cut the re-reads, or because the workload grew and the bill grew with it. Which of those is happening is the orchestrator’s reading of the system’s health.
6.4.1 The cost record
A report needs a record underneath it, and the record is per run. If the bill is an architectural document, this is the document: the minimum I would ask a system to emit for every run, whether or not the run finished.
| Field | What it holds | Why it is there |
|---|---|---|
run_id |
One attempt by one agent | The unit everything else hangs off |
trace_id |
The customer request this run served | Joins the cost to Chapter 16’s evidence pack, which uses the same key |
task_id, workload |
The task-system item and the workload it belongs to | Attribution — unattributed spend is unmanageable spend |
model, effort |
Model identifier and reasoning-effort setting | The two dials that set the price range |
input_tokens, cached_input_tokens, output_tokens |
Metered volume, cached reads broken out | Cached reads are the cheapest tokens you buy |
tool_calls, subagent_runs |
Tool invocations made and child runs spawned | Fan-out is where budgets disappear |
retry_of |
The run_id this run is repeating, if any |
A retry is a full-price run and must not look like new work |
wall_time_ms |
Start to finish | The always-on tax is a function of time, not tokens |
rate_card, cost_usd |
The per-million prices in force at run time, and the product | Prices move; store the rate you paid |
outcome |
completed, failed, stopped_by_limit |
Only completed runs count toward cost per completed task |
budget_id, budget_owner |
The spend limit this run drew against and who authorized it | A credential has an issuer |
What the record decides is narrow and useful. It decides attribution, because every row belongs to a workload. It decides the unit of account: cost per completed task is the sum of cost_usd over a workload divided by the count of rows whose outcome is completed, with the retries and failures in the numerator where they belong. It shows what share of the bill was cache hits, fan-out, or a loop the harness had to stop; the stopped_by_limit rows are how the budget kill switch in the next section becomes visible in the data rather than in an incident channel. And because it shares trace_id with Chapter 16’s join keys and outcome with the operating record Chapter 8 defined, the cost of a run can sit beside the authority the run used and the human decision that followed, which is the only arrangement in which “was that worth it” has an answer.
What the record does not decide matters as much. It does not know whether a completed task was correct; that is the operating record’s business, and a workload with a beautiful cost per completed task and a thirty percent error rate is a cheap way to produce mistakes. It does not know whether the workload was worth delegating at all; that was the triage in the previous section. And it does not know what the price will be next quarter, which is why rate_card is a field and not an assumption. The record is an instrument for reading the system. Reading it is still the orchestrator’s job.
FinOps Foundation, State of FinOps 2026, February 2026, https://data.finops.org/ — sixth annual survey, 1,192 respondents, more than $83 billion in annual cloud spend; 98 percent manage AI spend (63 percent in 2025, 31 percent in 2024); AI cost management the most-desired skillset; many asked to self-fund AI through optimization savings. A self-selected practitioner survey; it measures the scope of the FinOps remit, not the size or frequency of AI cost overruns.↩︎
The documentation-agent arithmetic, the five-percent routing example, and the heartbeat arithmetic are from Chapter 5’s pricing sections, which footnote the September 2026 list prices in full. Heartbeat: 300,000 × $10/M = $3.00, plus 2,000 × $50/M = $0.10, so $3.10 per call; × 24 = $74.40 per day; × 1,000 agents ≈ $74,000. Chapter 5 rounds the daily figure to “about seventy-five dollars.” List rates, before caching and batch discounts.↩︎