7.8 What the first week should not do

Book 1 · Your Next Job TitleChapter 7 · section 8 of 8

The false starts are as regular as the successes, and each of the four below is grounded in an incident this book has already examined. I point at them rather than retell them, because the reader who has just finished a first week is the reader who should now see what the constraints chapters are for.

Starting from a tool. The most common first week begins with seats — a purchase, a rollout, a training session on the chat window — and no process. It feels productive, and the feeling is the problem. METR’s 2025 trial is the cleanest evidence that people are poor judges of their own gains: the developers in it believed they were faster and were slower, and METR’s 2026 update, which retired the headline number, kept the warning about self-report. Chapter 11’s try-out — one real workflow, read-only access, a week, and an artifact a specialist can review — is the alternative, and it is the reason the week above ends with a handoff instead of a satisfaction survey.

Running where production is reachable. Replit’s July 2025 incident, examined in Chapter 14, is the case to remember when somebody proposes skipping the test environment because the task is small: development, preview, and production were one database, and eleven capital-letter instructions not to change anything did not hold.28 Dana’s drafts folder and Lena’s documentation branch are the same control at the scale of a first week. If the place the system writes is also the place customers read from, the week is not small enough to finish.

Connecting the first system to everything you can reach. The instinct is generous — give it the whole mailbox, the whole drive, every repository, so it has context — and the Invariant Labs demonstration against GitHub’s MCP server in May 2025 is what the instinct produces: private data in reach, an issue anyone could write, and a pull request to write the results into, with the user clicking Always Allow.29 The week above scopes the credential to one mailbox or two repositories, and the figure Chapter 5 reports — developers approve 93 percent of the prompts they are shown — is why Day 4 asks for the boundary in a permission rather than in a prompt you will confirm forty times before lunch.

Pricing the demo. A first week costs a few dollars, and the temptation is to multiply by fifty-two and call it a budget. Cursor’s June 2025 repricing, examined in Chapter 15, is what happens when the unit of account is the wrong shape: a request that once meant one question came to mean a long-horizon task consuming ten times its neighbor’s tokens, and the surprise bills were the variance in customers’ own workloads becoming visible.30 Dana’s three hundred messages a week will never be Dana’s problem. The same design on every mailbox in the company will be, and the per-key cap in Section 13.3 is the cheapest insurance the week can buy.

None of these is an argument against beginning. Each is an argument for beginning with the eight things in Section 13.3 in place, and for reading the two chapters after this one before the second week rather than after the first incident.

Two things in the records Lena and Dana handed over cannot be settled by the people who wrote them. Dana’s mailbox is open to the world by design, and anyone who can write to it can now write to the system; Lena’s issues and pull requests are open to her whole organization, and she has not yet asked who that includes. And both records show what one week cost on one process, which says nothing about what the same design costs on every mailbox and every repository, or what the invoice looks like the night a run does not stop. The record makes both questions askable, which is the reason for building it. Who else can write to what the system reads, and what happens when the bill arrives before the value does — those are the pressures on a first week from outside, and they are the orchestrator’s to answer before the second one begins.

HQ 6 — Assembled. The human and AI each wrote portions of this chapter. I assembled, reviewed, and take responsibility for the whole; the voice and arguments are mine, and I know which parts are which.


  1. The Replit incident of July 2025 — Jason Lemkin’s production database deleted during a declared code freeze, one database shared across development, preview, and production, and Amjad Masad’s commitment to “automatic DB dev/prod separation” — is documented in Chapter 14 with primary sourcing. Cited here by back-reference only.↩︎

  2. The Invariant Labs demonstration of May 26, 2025 — a public GitHub issue instructing an agent connected through GitHub’s MCP server to disclose private-repository contents in a public pull request, with the researchers’ note that many users choose an “Always Allow” policy — is documented in Chapter 14 with primary sourcing. The 93 percent approval figure is Anthropic’s, documented in Chapter 5’s permission-fatigue sidebar. Cited here by back-reference only.↩︎

  3. Cursor, “Clarifying our pricing,” July 4, 2025, https://cursor.com/blog/june-2025-pricing — the June 16–July 4 timeline, the refund offer, and “the hardest requests cost an order of magnitude more than simple ones.” Examined in Chapter 15. The vendor’s account; it does not report how many customers received surprise bills. Verified September 9, 2026.↩︎