1.2.1 One measured case, before the argument continues
Everything after this section builds on the shape of Maya’s morning, so it is worth pausing on the one place in this book’s opening stretch where the shape is measured rather than illustrative. Klarna, the Swedish payments company, put an AI assistant on its customer-service front line and announced the results in a press release: in its first month the assistant handled 2.3 million conversations — two-thirds of all chats — resolved errands in under two minutes against eleven previously, cut repeat inquiries by 25 percent, matched human agents on satisfaction, and did work the company said was equivalent to about 700 full-time agents, on track for $40 million in profit improvement in 2024.2
Read those numbers the way the rest of this book reads numbers. They are company-reported, aggregate, and announced on launch day by a company selling a story about itself; no segmented breakdown by issue type was ever published, and the satisfaction-parity claim is the vendor’s own scorecard. But they are also the largest measured deployment of delegated conversational work in a consumer-facing function up to that date, and the direction of every number is the direction this chapter described: the work moved from humans to a system, at scale, with the results counted.
What happened next is the part that matters here. Fifteen months later, Klarna’s chief executive told Bloomberg that cost had been “a too predominant evaluation factor” in how the operation was organized, that the result was “lower quality,” and that the company was hiring people again so a customer could always reach a human.3 The assistant kept the high-volume tier — two-thirds of chats, 82 percent faster responses — and the humans came back for the tier where measured parity had not held: the complex, emotional, expensive conversations.
Hold the two halves together, because the V-shaped story is the finding. The assistant absorbed the volume, exactly as advertised — that half of the announcement survived a year of scrutiny. What failed was not the technology but the boundary: nobody had designed where volume stops being enough, who owns the moment a customer needs something the assistant cannot give, and what evidence would show the difference. The company measured resolution time and got exactly what it measured. The orchestrator’s job — the job this book is about — is the part Klarna had to add back by hand: deciding what the system may do alone, where the human sits, and what counts as done well enough to keep. Maya’s morning, with the numbers attached and the walk-back included.4
Klarna, “Klarna AI assistant handles two-thirds of customer service chats in its first month,” press release, February 27, 2024, https://www.klarna.com/international/press/klarna-ai-assistant-handles-two-thirds-of-customer-service-chats-in-its-first-month (fetched and verified September 20, 2026). Cited for the first-month figures exactly as listed: 2.3 million conversations (two-thirds of chats), 700-FTE-equivalent capacity, satisfaction parity, 25% drop in repeat inquiries, under-2-minute resolution against 11 previously, 23 markets, and the $40M estimated 2024 profit improvement. All figures are company-reported launch metrics; aggregate only, with no segmented breakdown by issue type or cohort released.↩︎
Klarna CEO Sebastian Siemiatkowski, interview with Bloomberg, May 8, 2025, “Klarna Slows AI-Driven Job Cuts With Call for Real People,” https://www.bloomberg.com/news/articles/2025-05-08/klarna-turns-from-ai-to-real-person-customer-service; quotes (“a too predominant evaluation factor,” “lower quality”) as reported by CX Dive, https://www.customerexperiencedive.com/news/klarna-reinvests-human-talent-customer-service-AI-chatbot/747586 (both verified September 20, 2026). CX Dive’s same-period reporting supplies the retained-assistant figures: two-thirds of inquiries still automated, 82% faster response times since launch, 25% fewer repeat issues. Cited for the reversal and its stated reason; the efficiency numbers are still the company’s own.↩︎
This book returns to Klarna three times — the escalation-design lesson in Chapter 7, the failure-pattern reading in Chapter 9, and the summary here — with the same care each time: the 700-agent figure was announced task-equivalent capacity, not a headcount reduction, and the 2025 correction was a scope rebalance, not an abandonment of the assistant. Cited here because it is the only consumer-scale deployment with a first-month measured announcement AND a public executive walk-back — the pair this book’s evidence discipline requires.↩︎