2.4 Start with an honest assessment
Before changing any titles or buying an agent platform, map the work the organization already does. For each department, record:
- the routine tasks people perform;
- the decisions that require judgment;
- the systems and data those tasks depend on;
- the permissions required to take action;
- the common failure modes;
- the people who know how to recover from those failures;
- the work that customers or other departments receive; and
- the evidence used to decide whether the work was successful.
Do not hand that inventory only to managers, because the most valuable knowledge usually sits elsewhere: with the support agent who knows which customers are being routed incorrectly, the database administrator who knows which report is built on an undocumented table, the security analyst who knows a permission is documented but not enforced, or the operator who knows which alert can safely be ignored and which one cannot.
Then sort the work into three categories:
- Routine execution: work that follows a known pattern and may be delegated to a system.
- Expert judgment: work that depends on domain knowledge, unusual conditions, or consequences that are difficult to reverse.
- System responsibility: work that defines goals, permissions, evaluation, handoffs, monitoring, and accountability.
Orchestration reduces how much routine execution people have to perform, and it removes neither expert judgment nor system responsibility. In most organizations those two become more important precisely because so much more routine work is now being produced by systems.
2.4.1 How to identify talent
The assessment also identifies people. The inventory that sorts the work shows who already behaves like an orchestrator in miniature, and the markers are observable rather than mystical:
- They ask what would change. When a new tool or request arrives, they want to know the effect on the workflow, the permissions, and the records, not only the ticket in front of them.
- They trace work end to end. They follow a request from the form to the database instead of treating the system as a vending machine, and when something breaks they already know which of the four teams involved is having the bad day.
- They have ideas about how the work should change. They keep a list, written or mental, of pointless handoffs, repeated reconciliations, and decisions everyone makes the same way every week — exactly the routine execution this section said moves into systems.
- They want the review queue. When a system starts producing candidate work, they are the first to read it and the first to say the third one is wrong.
The research on who gains from these tools is better than folklore, and it cuts two ways. In the largest field experiment to date, 758 consultants at Boston Consulting Group were randomized to GPT-4 access or no access across eighteen realistic tasks. The group with AI completed 12.2 percent more tasks, 25.1 percent faster, at roughly 40 percent higher rated quality — and the gains were largest for below-average performers, 43 percent versus about 17 percent for the strongest. On routine work inside the model’s capability frontier, the tools are a leveler: they raise the floor more than the ceiling. That is an argument for giving them to everyone, and it is specifically not an argument for promoting the most fluent user. The skill that leveled is producing work. The skill this section is staffing for is judging it — and the same experiment found the advantage disappeared on tasks just outside the frontier, where the model confidently produced wrong answers that its users repeated.6
A second finding matters just as much for selection: people are poor judges of their own gains. In a 2025 randomized controlled trial, experienced open-source developers given AI coding tools believed they were about 20 percent faster while measured times showed them about 19 percent slower, across 246 tasks. The study’s authors suggest the developers were partly choosing the tools because the work felt easier, not because it went faster.7 Enthusiasm is real, and it is not evidence.
So run a try-out rather than an interview performance. Give the candidate a real workflow from their own area with read-only access and a week, and ask for the map — inputs, permissions, handoffs, failure modes — plus one proposed improvement and the evidence behind it. Have a specialist review the consequences. Then staff on the artifact: discount confident self-assessments, and leave your strongest performers on the tools where they are already strong — pulling the best operator out of the environment they are good at, to do coordination full time, is a staffing decision that should require evidence too.
Fabrizio Dell’Acqua et al., “Navigating the Jagged Technological Frontier: Cognitive Tasks, Problem-Solving, and the Effect of AI,” SSRN 4573321 (2023), https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4573321. Randomized field experiment with 758 BCG consultants on 18 realistic tasks; inside the AI’s capability frontier, consultants with GPT-4 access completed 12.2 percent more tasks, 25.1 percent faster, at roughly 40 percent higher rated quality, with below-median performers improving 43 percent versus about 17 percent for the strongest. The jagged frontier of the title is the finding that adjacent-seeming tasks fall on different sides of the capability line, and outside it, AI users repeated the model’s confident errors.↩︎
Joel Becker et al. (METR), “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity,” arXiv:2507.09089 (July 2025), https://arxiv.org/abs/2507.09089. Randomized controlled trial, 246 tasks by experienced open-source maintainers of mature repositories; developers forecast a 24 percent speedup and afterwards estimated a 20 percent speedup, while measured completion times were about 19 percent slower with AI tools. The authors suggest participants may have traded productivity for ease. METR’s February 24, 2026 update (https://metr.org/blog/2026-02-24-uplift-update/) marks the 2025 result as out of date for current tools and calls its follow-up data unreliable, while repeating that self-reported speedups are unreliable; Chapter 13 carries the update. Results are from one toolset and population, not a claim about every workflow.↩︎
- 2.1 Block: hierarchy vs. intelligence
- 2.2 JPMorgan Chase: connectivity as the moat
- 2.3 Titles: what the market is doing right now
- 2.4 Start with an honest assessment
- 2.5 Help developers become orchestrators
- 2.6 Start new projects differently
- 2.7 Design the team: composition, size, and specialists
- 2.8 Recognize the impedance mismatch
- 2.9 Plan the transition in ninety days
- 2.10 The human cost is part of the plan