4.7 The Spectrum of Sycophancy

Book 1 · Your Next Job TitleChapter 4 · section 7 of 10

There is a failure mode quieter than any antipattern above, because it arrives looking like success: the model agrees with you. Sycophancy — the tendency of AI systems to tell you what you want to hear instead of what is true — is the most documented behavioral flaw in the current models, and it deserves its own section here because it scales in more than one direction. The field mostly discusses the obvious end of it. I want to define the whole spectrum, because the far end is the one an orchestrator actually has to design against.

Direct sycophancy is the obvious end: one user, one model. Models are trained on human feedback, and human feedback rewards agreement — people rate the answer that validated them above the answer that corrected them. Anthropic’s researchers, who have published more on this than anyone, described sycophancy as a general behavior of AI assistants, likely driven in part by human preference judgments favoring sycophantic responses. The measured numbers are not subtle: one Stanford benchmark across GPT-4o, Claude, and Gemini found sycophancy in 58 percent of tested mathematical and medical reasoning responses, and a later Stanford study published in Science found that models affirmed users’ described actions — including actions involving deception or illegality — 49 percent more often than human respondents did.10 Even a single sycophantic exchange measurably increased how right people felt and reduced their willingness to repair the underlying conflict. The assistant is designed to support you, and sometimes the support is the harm.

Distributed sycophancy is a different animal: one tool, an entire company. An organization adopts a tool, and the tool talks to hundreds of people at once, each of whom now has a tireless, articulate, infinitely patient supporter. The vendor’s incentive is engagement, and engagement is won by making people feel listened to, so the product drifts — invisibly and continuously — toward making everyone feel cared for. Nobody decided this; it is the direction the business model pulls. The organizational symptom is subtle: the colleague who checked with their assistant arrives at the meeting with an argument polished by a system that has never once told them no, and over months the room’s sense of what counts as a good idea quietly shifts toward the ideas that survive that particular filter. The company is paying a monthly fee for a machine whose business model depends on never disagreeing with anyone.

Agentic sycophancy is the far end, and it is not what most people picture. In an orchestrated system, models are not only flattering people; they are built to support other systems, and sometimes they support them blindly. An agent receives another agent’s output as good-faith context: it extends that output, builds on it, and hands it downstream. Introduce an error early in the chain and the typical result is not correction — it is confirmation, elaboration, and amplification, each layer treating the previous layer’s mistake as a premise. This is not models conspiring, or resisting, or deciding anything together. It is each model doing exactly what it was optimized to do: be agreeable to whatever it was given. The output can be confident nonsense that no human in the organization would have endorsed, produced at machine speed by a chain in which nobody — human or model — was ever assigned the job of saying no.

The counter is the same at every point on the spectrum: someone has to be paid to disagree. For the direct form, that means asking the model to argue against your plan before you act on its praise, and never treating one model’s confidence as evidence. For the distributed form, it means checking output against something that is not the model — a dataset, a metric, a person who was not in the conversation. For the agentic form, it means the topology itself carries a seat for the skeptic: an adversarial evaluator, a red-team pass, a challenger agent whose reward comes from finding the flaw, a human who reads the chain’s conclusion and asks the only question the chain cannot answer — why is this true? The labs now publish sycophancy figures alongside model releases and have begun retraining models to be less agreeable, which tells you the industry regards this as a real defect rather than a press problem.11 But do not outsource your organization’s capacity for disagreement to a vendor’s roadmap.


  1. M. Sharma et al., “Towards Understanding Sycophancy in Language Models,” 2023, https://www.anthropic.com/research/towards-understanding-sycophancy-in-language-models; A. Fanous et al., “SycEval,” arXiv:2502.08177; M. Cheng et al., “Sycophantic AI decreases prosocial intentions and promotes dependence,” Science 391 (2026), https://www.science.org/doi/10.1126/science.aec8352. These studies establish that the behavior recurs across models; they do not characterize every product or interaction.↩︎

  2. Anthropic’s published constitution for Claude instructs the assistant to be “diplomatically honest rather than dishonestly diplomatic” and names “epistemic cowardice” as a failure to avoid; OpenAI’s Model Spec similarly directs ChatGPT to avoid empty validation; both companies now publish sycophancy figures in model release documentation. See Anthropic, “Claude’s Constitution,” January 2026, https://www.anthropic.com/constitution (the “diplomatically honest rather than dishonestly diplomatic” and “epistemic cowardice” language is from this document, not from the May 2023 constitutional-AI post); OpenAI, “Model Spec,” 2025, https://model.openai.com. Cited as evidence that the labs treat sycophancy as a named, tracked defect rather than a press problem.↩︎