3.2 Contracts and grants
3.2.1 The contract before the conversation
Before any agent in the home-goods runbook opened a prompt, the contract already existed as a typed document in the team’s repository. That is the whole point: a contract is not a paragraph the orchestrator typed into a chat box but a file that CI loads, that the agent reads at startup, that the operating record attaches to every action, and that a reviewer can get through in three minutes.
# delegation/contract.yaml
name: order-service-isolation-routing
version: 2026.02.04
objective: |
Identify affected stores and order cohorts, then propose routing them
to manual review during a Sev-2 incident, before any human is paged.
inputs:
- read: order-service schema (prod, read-replica)
- read: store and fulfillment-route map (read-only)
- read: incident page metadata (correlation id, severity, route)
- read: routing-rule inventory (read-only)
authority:
- may: query schema, fulfillment map, inventory, and recent deploy log
- may: produce a structured isolation and routing proposal
- may_not: stop all order processing
- may_not: reroute orders outside the affected stores or cohorts
- may_not: page a human without a structured isolation proposal
- may_not: read payment-service data
personality: |
You are a cautious incident-readiness agent. Prefer returning
"insufficient evidence" over guessing. Treat any schema field you
were not granted as absent. Do not infer the affected store or order cohort from
deploy notes alone.
evaluation:
held_out_set: incidents/routing/2025-Q3-Q4.jsonl
metrics:
- precision_on_store_and_cohort_identification >= 0.97
- false_isolation_rate <= 0.01
- median_proposal_latency_seconds <= 30
on_failure: revoke_routing_authority_and_escalate: order-service-on-call
evidence:
must_return:
- correlation_id
- affected_stores (list, must match fulfillment map)
- affected_order_cohort (string, must match order records)
- confidence (0.0-1.0)
- contradicting_signals (list)
stop_conditions:
- affected stores or order cohort cannot be identified from approved records
- affected stores do not match the fulfillment map
- contradicting_signals is non-empty
- medical, temperature-sensitive, or regulated shipments are in scope
- confidence < 0.85
reviewer: order-service on-call engineer
reversal:
- structured proposals are not actions; only the approving agent may change routing
- if the approving agent changes routing, the rollback command is infra/runbooks/order-routing-restore
- the approving agent must replay the last 5 minutes of traffic against staging before restoring normal routing
One contract is not one strictness. The envelope above is tight because the consequences it bounds are contained: a mis-identified store routes orders to a human queue, and a rollback command reverses the action. Not every delegation has that shape, and the same organization will need different envelope strength for different classes of consequence. A claim denial moves money and can be reversed on appeal, so it earns a tighter evidence section than the routing agent and a named human in the loop before the denial is final. A production deployment is not reversible by appeal — customers have it — so it earns the tightest envelope this book describes: a deterministic gate outside the model, a blast-radius limit, and a stop that no agent can overrule. A claim-drafting step that only prepares text for review barely needs an envelope at all, because nothing it does is final. The design error to avoid is borrowing another system’s envelope wholesale: a contract calibrated for reversible work, deployed around irreversible work, is the Replit incident in miniature. Size the envelope to the consequence class — reversibility, blast radius, and whether a customer sees the result — and write the sizing into the contract itself.
Two contracts that share the word “isolate” and differ on every line describe two different systems, and an orchestrator who treats them as the same system will eventually pay for it. Note also which half of the document does the heavy lifting: a contract that names what the system may not do is worth more than one that names what it may.
None of that contract is invented from nothing. Every field in it has a shipping counterpart, and the strongest example comes from AWS, which open-sourced Cedar — a language for writing authorization as policy code — and in March 2026 made generally available a service called Policy in Amazon Bedrock AgentCore. A policy engine holds the Cedar policies and attaches to a gateway that sits between agents and their tools; the gateway intercepts every agent-to-tool request and evaluates it against the policy before allowing or denying the call. The evaluation happens outside the agent’s code and outside the model’s reasoning, nothing is permitted unless a policy permits it, and the service validates policies against the gateway’s actual tool schemas, so a rule naming a tool that does not exist is rejected before deployment. Teams can also write a policy in plain language and let the service convert it to Cedar.11
AWS publishes a reference implementation for the harder case — delegation chains — which runs three policy checks at every hop: whether the calling agent has the trust level and namespace to invoke the tool, whether the delegated task falls within the receiving agent’s registered capabilities and how many hops the chain has already taken, and whether the human who started the chain holds the role and the multi-factor authentication the request requires. The originating user’s identity travels through the chain in a signed envelope, so downstream agents cannot alter it.12
In Cedar, the grant reads almost like a sentence. The reference implementation permits the finance agent to invoke the payment tool when the agent’s trust level is at least three, when it belongs to the payments namespace, and when it is in the production stage:
// L1-001: Finance agent can invoke payment tools
permit(
principal == AgentAuthz::Agent::"finance-agent",
action == AgentAuthz::Action::"invoke_tool",
resource == AgentAuthz::Tool::"process_payment"
) when {
principal.trust_level >= 3 &&
principal.namespace == "payments" &&
principal.lifecycle_stage == "production"
};
The condition is the contract: if any one of the three clauses fails, the call is denied, and the evaluation happens outside the agent’s code, so it does not depend on the model remembering its instructions.
The same pieces are visible across the rest of the stack. Amazon Bedrock agents declare the operations they may call as OpenAPI schemas that define what the agent can invoke and with which parameters.13 Claude Code reads allow, ask, and deny rules from a settings file kept in the repository, and a matching deny rule blocks the call no matter what an allow rule says.14 Open Policy Agent has spent a decade making policy-as-code ordinary for infrastructure — declarative rules, kept in version control, enforced at admission, reviewed and tested like code.15 And the A2A protocol publishes each agent’s capabilities as a JSON manifest at a well-known address, so a caller can check what an agent does before sending it work.16 What none of these ships is the assembled document: inputs, authority, personality, evaluation, stop conditions, and reversal in one versioned file that the operating record attaches to every action. The pieces exist, and they are serious. The assembly is still done by hand.
The strongest objection to policy-driven delegation in orchestrated systems is that some of these systems are dynamic. An inference engine can learn, and the set of tools and skills it understands can change over time. The most dynamic pattern in the field — install a minimal agent like Pi and let it add new skills as the work demands, since Pi’s design deliberately leaves tools and skills as installable extensions17 — is the opposite of a strictly defined policy written before deployment. There is a give-and-take here that no vendor has resolved: it may not be possible to enumerate everything an agent will ever try to do, and a policy that has to predict the future is hard to write. But the conclusion runs one way for anything that matters. For critical resources — source code control, financial accounts, anything that touches human beings — what the system can and cannot do has to be locked down with a formal policy engine, enforced together with credentials that strictly limit access. If the system’s capabilities change at runtime, that change is a change to the delegation contract: it goes through review, or it does not happen. That is the one constraint in this chapter that does not bend.
This book is not suggesting that the systems being orchestrated are forbidden from asking for more. An inference engine that hits the edge of its envelope and wants an additional permission, a new tool, or broader access to resources should say so — requesting more authority is one of the most useful signals a delegated system can produce. What it cannot do is grant the request itself. And the orchestrator’s job is not to write a policy once and defend it forever. The job is to read the policy alongside the execution records — what the agents requested, where they hit stop conditions, where the envelope proved too tight or too loose, what the incident reviews found — and then decide how the policy needs to change to meet new requirements and new operations. A delegation contract that never changes is not a sign of good design. It is a sign that nobody is doing this part of the job.
3.2.2 Two agents with different evidence
Useful disagreement comes from independent evidence. Useless disagreement is the same model, given the same inputs, producing three slightly different paragraphs — which is why more agent voices do not automatically produce more insight.
Research separates the useful and useless cases, and the dividing line is what makes the runs differ. Sampling many reasoning paths from one model and keeping the most consistent answer measurably improves reasoning — but only because the sampled paths genuinely differ.18 Running several instances of the same model as a debate improves it further, because each round changes what the next agent sees: in the original study, arithmetic accuracy rose from 67 to 82 percent and grade-school math from 77 to 85 with three agents and two rounds of debate.19 Same model, same question. The mechanism forced the disagreement to carry information, which is exactly what identical inputs on identical prompts do not do.
Two caveats belong with those numbers. First, the study is from 2023 and reports the raw error rates of 2023 models, and those numbers are already conservative. When OpenAI released o1, a model trained to reason longer before answering, its accuracy on competition mathematics went from the twelve percent its predecessor managed to seventy-four percent on a single attempt — and to eighty-three percent when the model compared its own attempts and kept the consensus answer, which is self-consistency again, paid for with compute.20 Second, the failure being measured is arithmetic, and arithmetic is the most addressable error an inference engine has. Give the agent a code tool so it computes instead of guessing and the errors collapse: the program-aided-language-models work found that offloading arithmetic to a Python interpreter beat much larger models reasoning in plain text, holding above sixty percent on problem sets where text-only reasoning fell to roughly twenty.21 So the lesson for an orchestrator is not that debate is mandatory. It is that reasoning errors can be bought down — with stronger reasoning, multiple iterations, and a quality gate in front of the answer — and the orchestrator chooses how much of each to pay for.
The home-goods company tested that distinction during a flash-sale review, asking whether the order service could absorb a 30 percent surge without harming customers.
The first agent read ninety days of order-service latency, concurrency during earlier promotions, the database connection-pool ceiling, and the autoscaling group’s historical lag. It concluded that the service could absorb roughly 28 percent more traffic before the latency budget broke, and recommended a 25 percent guardrail.
The second agent received a different evidence set. It read the same service logs, but also checked two weeks of support tags, the customer cohort in a recent checkout experiment, and incident records from the previous six months. It found that a week-47 configuration change had made the partial-refund path fail under concurrency. Its recommendation was narrower: protect the partial-refund path instead of throttling the entire order path.
That disagreement was worth having because the two agents had looked at different parts of the operation: one saw capacity, and the other saw a customer-impacting failure hidden inside one particular order path. The orchestrator did not pick the answer with the more confident tone, but compared the evidence, the assumptions, and the exact action each recommendation would authorize. The decision record captured the useful part:
| Field | Value |
|---|---|
| Alternatives considered | Throttle entire order path; throttle partial-refund path; do nothing |
| Evidence that ruled them out | Plan A: latency headroom. Plan B: week-47 partial-refund regression under concurrency. |
| Objections raised | Plan A flagged the partial-refund path as in scope; Plan B flagged it as the load-bearing assumption. |
| Resolution | Plan B adopted with a narrow throttle on the partial-refund path only. |
| Date and context | model versions, prompt revisions, deploy ids, schema version |
The real boundary in that case was the week-47 config change. Both plans had been given identical authority to read internal data, and only Plan B had been told to look outside the team’s usual priors. An orchestrator who treats “read access” as a single dial gets one opinion printed three times; an orchestrator who treats context, authority, personality, and evaluation as four separate grants gets disagreement that survives scrutiny.
3.2.3 Context, authority, personality, evaluation are four separate grants
Most early orchestration work collapses those four grants into one thing and calls it “access,” and the result then misbehaves in four different places.
Context is what the system can read, authority is what the system can change, personality is the operating instruction it is given, and evaluation is the held-out set, the metrics, and the action that a failure triggers. Each grant can be tight or loose independently, and the combination matters far more than any individual setting. A model with broad context and narrow authority makes a good security reviewer, since it sees the architecture without being allowed to touch it. A model with narrow context and broad authority is a different shape entirely, capable of changing a payment service without ever seeing the fraud patterns the change might break. On a diagram the two look nearly identical, and they fail in completely different places.
Evaluation is the most consistently underweighted of the four. It is not a score but a theory of what matters, encoded as a held-out set plus the action that follows when a metric crosses a line. An evaluation set built from the team’s existing intuition will measure the team’s existing intuition, while one built from real incidents — including the ones the team got wrong — measures something closer to the actual work. If the agent proposing order isolation is graded only on how fast it responds, the organization has declared that response time matters more than correctly identifying the affected stores and customers. The metric did not merely measure the system. It chose the system’s values.
An inference engine adapts to its metrics exactly as a model adapts to its incentives — Strathern’s phrasing of Goodhart’s law applies without translation: “when a measure becomes a target, it ceases to be a good measure,”22 so if you reward speed you get speed and have said nothing about correctness. The orchestrator does not get to skip this choice. The metrics are the choice.
Vendors are starting to ship that half of the contract as product. Amazon Bedrock AgentCore Evaluations, generally available since March 2026, scores agents with thirteen built-in evaluators covering response quality, safety, task completion, and tool usage, and it runs in two modes: continuous scoring of sampled production traffic, and on-demand runs wired into CI/CD pipelines for regression testing, with the same evaluators in both places — what a pipeline tested is what production monitors.23
Before promoting a delegation contract, a useful checklist:
- the context grant lists each schema, repository, and data source by exact name;
- the authority grant lists each verb (may / may_not) by exact effect;
- the personality grant is short enough to read aloud and contains one operating instruction, not a manifesto;
- the evaluation set is held out from training and includes cases the team got wrong;
- the on_failure action is reversible from a routing rule and a known rollback command.
That checklist is not the architecture on its own, but it is the fastest test of whether a contract is ready: if a team cannot answer those five questions in three minutes, the contract is not ready for production.
One question sits above all four grants, because it decides whether the system needs another agent at all. The design question is never how many agents the system can run. It is where nondeterministic judgment earns its coordination cost — and for every proposed agent: what unique capability, context boundary, or authority boundary does it own; why can that responsibility not be a tool or a prompt inside an existing agent; how will the system know the agent is done, correct enough, over budget, or unsafe; what state and evidence cross the boundary, and what must not; and who owns the final result after a handoff. Ask it before the contracts are drawn, because a topology that never asks it adds agents one at a time, each defensible on its own, until nobody can say which agent the business actually depends on or which one it could safely remove.
3.2.5 Structured outputs make disagreement usable
When every delegated system returns a paragraph, the orchestrator ends up comparing prose, and prose is exactly where assumptions go to hide. Prefer structured outputs wherever the task permits them: claims, sources, confidence, counterexamples, unresolved questions, proposed action, and reason to stop. The two researcher agents in the flash-sale example did not return paragraphs. They returned objects sharing one schema:
{
"correlation_id": "fs-2026-q1-001",
"thesis": "string",
"affected_paths": ["string"],
"recommended_throttle": {"path": "string", "percent": 0.0},
"confidence": 0.0,
"contradicting_signals": ["string"],
"stop_condition_triggered": false
}
The orchestrator compared the two objects field by field, and the disagreement turned out to live in recommended_throttle.path and in contradicting_signals. A paragraph reporting that “the surge looks absorbable, you should probably throttle” would have concealed both.
Structured output does not make an answer true; it makes an answer inspectable, and it lets a later agent challenge one specific field instead of praising or rewriting an entire essay.
Typed output also makes the checking automatable, which matters more as volume grows. When every proposal shares a schema, a small deterministic program can compare two agents’ proposals field by field and flag every disagreement without a person reading either one; it can flag continuity breaks, missing evidence, and confidence scores that do not match the citations. If every output is a paragraph, none of that is possible — the only quality check left is a human reading everything the system produces, and that stops working the first week the system writes a hundred paragraphs a day.
The capability is not aspirational. OpenAI’s Structured Outputs feature constrains decoding so that a model’s output cannot deviate from the developer-supplied JSON schema; on the company’s own evaluations of complex schema-following, a model with the feature scored a perfect 100 percent where an earlier model without it scored under 40.24
3.2.6 What the agent can actually do
“Access” is too blunt a word to design with. A security review that ends with “the agent has access to the order service” has not said much, because seven different questions are hiding inside that sentence, and each deserves its own answer.
One: can the system read the data — customer records, order histories, payment logs? Two: can it change state, and which state exactly? Three: can it propose a change for someone else to approve, or only draft one? Four: can it finalize a decision without a human gate? Five: can it contact anyone outside the system — a customer, a vendor, a partner — and on what terms? Six: can it spend money, up to what limit? Seven: can it expand its own budget, grant itself more compute, more storage, more of anything?
An agent that can do all seven and an agent that can do one are both “agents with access to the order service,” and they are completely different systems. The one that can do all seven can read a customer’s records, change them, approve its own change, email the customer about it, pay a vendor to speed the work up, and raise its own spending limit when the invoice runs over. The one that can only read can do none of that. Until those seven questions are answered separately, nobody — not the reviewer, not the on-call engineer, not the CEO — knows which system they actually have.
The most common agent token is “everything.” Most people installing their first agent — OpenClaw at home, Hermes or Claude Code on a development workstation — have never had to think about privileged access before. When they grant the system a credential, they face a wall of fifty-odd privilege scopes and levels, and clicking through them is tedious. The common move is to skip the review and grant everything, exactly as the most common user password for years was the word “password.”25 The developer-tools world made the same trade for a decade: a classic GitHub personal access token with the
reposcope hands an agent “broad access to all data in private repositories the user has access to, in perpetuity,”26 which is why GitHub now ships fine-grained tokens — specific repositories, specific permissions, with organization approval policies — and recommends them “instead of personal access tokens (classic) whenever possible.”27 Granting everything is fine inside a sandbox; it is fast, and failure is cheap. In production it is not a recipe for disaster, because a disaster at least has a cause. It is a recipe for the unknown: agent systems get creative in ways that are tough to anticipate, and if you grant something unlimited permission to do potentially destructive things, bank on it eventually doing one of them. The habit that has to change on the way to production is precisely this one — from the token that can do everything to the token that can do the job.
The home-goods runbook answers all seven questions for every agent in it, which is what makes the runbook checkable in one pass. The routing agent’s answers are the contract in 7.2.1. The containment agent may change one routing rule for one approved store cohort, and that is the whole extent of what it may change. The on-call engineer is the only participant who may communicate with the outside world, and the only one whose approval is final. Write “access: order service” instead, and nobody can say in advance which dangerous thing just became possible.
None of those seven questions is a new idea; the rest of infrastructure has been answering them precisely for years. Kubernetes role-based access control grants permissions as named verbs — get, list, watch, create, update, patch, delete — bound to specific resources, and a rule can pin a single named instance, so the right to read one config map and nothing else is one declarative statement.28 Agent frameworks are converging on the same specificity: LangGraph lets a team configure approval per tool — require human approval for one tool, allow only approve-or-reject on another, and leave a read-only tool alone.29
Naming authority stage by stage is what makes a contract useful. Research may get broad reading access and no external communication. Implementation may get write access in a sandbox and no customer data. Review may see the evidence and the proposed action without being able to silently change the result. Promotion may require a human whose responsibility is on the record. This is the same instinct that makes good architecture separate services, since separation creates places where a decision can be inspected before it travels any further, and delegation ought to create the same places. An orchestrator who grants “can write code” or “can change routing” without asking which actions that includes has not defined a permission — whatever the grant accidentally allows is the real policy, and the eventual audit will find it.
3.2.7 Permissions are conditional on the situation
Every grant in the lifecycle table already says when it is valid — “during Sev-2 incidents while the contract and health checks are current” is a condition, not a decoration. The point worth making explicit is that permissions can be conditional on the operating situation, and that the situation can change what the system is allowed to do. Healthcare wrote this pattern first: an Epic user opening a restricted patient record must type a reason before the record opens, and every break-the-glass event is logged and reviewed later.30 The rest of infrastructure has a standard name for the general pattern — attribute-based access control, in which a request is evaluated against “attributes associated with the subject, object, requested operations, and, in some cases, environment conditions,” and because the inputs include the environment, the same policy “can perform dynamic authorization for data or applications that grants or revokes access in real time.”31
In an orchestrated system the environmental attribute that matters most is the health of the critical systems the arrangement depends on. Take the medical-shipment runbook that the design review in 7.5.1 comes back to. In normal operations the routing agent may propose reroutes for routine shipments, the containment agent may adjust one routing rule for an approved store cohort, and the on-call engineer approves at the gate. Now suppose the warehouse-management system loses power — the same system whose data the routing agent uses to know what is moving, what is temperature-sensitive, and what is controlled. The permissions cannot stay the same, because the environment that justified them is gone. The policy should say so in advance: when a critical dependency is unavailable, the routing agent’s grant to propose routing narrows to refer everything to human oversight; the containment agent’s single-rule grant is suspended; the on-call engineer’s authority widens from approving proposals to operating the process manually. An outage changes the situation, and the situation changes the permissions. An orchestrator who only writes the sunny-day policy has written half a policy.
That is also why the enforcement point matters. A conditional policy that lives in the prompt is a suggestion; a conditional policy enforced by the policy engine — the same engine that checks every tool call — is a rule. Cedar policies can already express conditions on the request context (when { ... } blocks evaluate request attributes at decision time),32 so the machinery for situation-dependent authority is shipping today; what ships is not a judgment about which situations should tighten a grant. That is the orchestrator’s to write.
3.2.8 Break-glass: operating without the AI
The dependency argument has a harder version, and it is one of the most important requirements in this chapter: the arrangement must be able to keep operating when the AI cannot. Infrastructure named this pattern long before agents existed. “Break glass” — the phrase borrowed from the fire-alarm case — is the “pre-planned, highly secure, and temporary method for granting elevated, often administrative, access to a small, authorized group of individuals during a major incident.”33 HIPAA requires covered entities to establish “procedures for emergency access” to electronic health records, and Yale’s implementation guide describes pre-staged emergency accounts, “created in advance to allow careful thought to go into the access controls and audit trails associated with them.”34 Two properties are invariant wherever the pattern is implemented: the emergency path is defined before the emergency, and every use of it is logged and reviewed.
What break-glass means for an orchestrated system is that the loss of AI is a policy-engine problem and a permissions problem, not just an operations problem. The provider can go down — the June 12, 2026 export-control directive forced Anthropic to disable its two most capable models, Claude Fable 5 and Claude Mythos 5, for every customer worldwide, with no advance notice.35 Credentials can expire. A model can become unsafe or unreliable. A company can lose access to a vendor. When any of those happens, the system must still operate — and the way it still operates is that a human jumps in and becomes responsible. The policy engine has to support that transition by name: a break-glass mode, declared in the policy, in which the agents’ grants are suspended, the affected workflows hand their current state and evidence to a named human, and that person’s actions are logged under an emergency authority that expires when the incident closes. Zapier’s 2026 enterprise survey put the stakes on the record: 81% of enterprise leaders are concerned about dependency on specific AI vendors, and for 47% “at least one key business function would stop working” if their primary vendor went down — which is exactly the number that break-glass exists to reduce.36 The drift to notice is the inverse of the sunny-day drift: in normal operation the system earns autonomy, and in an emergency the system hands autonomy back. An arrangement that cannot hand it back has no floor under it.
The design review should ask one blunt question about this mode: who is the human, and does the procedure work if that person is asleep? A break-glass procedure that has never been rehearsed is an assumption, not a control — infrastructure practice is explicit that “unpracticed emergency access is dangerous” and the override “should be tested regularly.”37 The same is true one level up: an AI-unavailable plan that has never been run with the models turned off is a document, not a capability.
3.2.9 Certified model dependencies
There is a contingency the policy engine has to name directly: which models are certified to handle this system’s inference. A system can grow comfortable depending on a single proprietary frontier model — take a hypothetical medical-logistics arrangement whose every agent has been built, tuned, and evaluated against Anthropic’s Opus 4.8. The dependency compounds quietly: prompts that read best against one model’s habits, tool schemas the other vendors do not quite support, evaluation sets whose thresholds were calibrated on that model’s error profile, skills whose instruction files have absorbed its conventions. Switch to a different vendor’s model — OpenAI’s GPT-5.6, say — and the system does not degrade gracefully. It behaves differently in ways nobody has certified, and “differently” in a consequential system means unreviewed.
The policy response is to make the dependency explicit in the contract, the same way other dependencies are named. The delegation contract should carry a certified-models field: which models this arrangement was evaluated against, which are approved to act in production, and what happens when a model is removed from the list — the same reconfirmation trigger the lifecycle table already names (“after a contract change, model change, schema change, or incident review”). The June 2026 export-control episode is the standing reminder that a frontier model can become unavailable for reasons no engineering team controls, and the survey numbers say most enterprises are not ready: only 6% believe they could switch their primary AI provider “without material operational disruption.”
The certified-models field also disciplines the other direction — the temptation to keep the dependency but quietly assume any model will do. If a team cannot say which models are certified for this system, they also cannot say what they would lose by switching, which means the dependency is unmeasured and the switching cost is unknown until the day it is paid. Naming the certified models costs an afternoon; measuring an unmeasured dependency after it breaks costs the incident.
AWS, “Policy in Amazon Bedrock AgentCore is now generally available,” March 3, 2026, https://aws.amazon.com/about-aws/whats-new/2026/03/policy-amazon-bedrock-agentcore-generally-available/, and the developer guide at https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/policy.html. Cedar policies are held in a policy engine attached to an AgentCore Gateway, which intercepts agent-tool traffic and evaluates each request before allowing or denying access; policies can be authored in natural language that converts to Cedar. Cedar is default-deny by design: see the Cedar policy language documentation, https://docs.cedarpolicy.com/.↩︎
Dhananjay Karanjkar, “Enforce least-privilege authorization in multi-agent AI chains using Cedar,” AWS Security Blog, July 2026, https://aws.amazon.com/blogs/security/enforce-least-privilege-authorization-in-multi-agent-ai-chains-using-cedar/, with the reference implementation at https://github.com/aws-samples/sample-cedar-agentic-ai-authorization. Documents three-layer policy evaluation for delegation chains (agent trust level and namespace; delegation hop limits and registered-capability subset checks; originating-user role and MFA) and HMAC-signed user-context envelopes.↩︎
Amazon Bedrock User Guide, “Define OpenAPI schemas for your agent’s action groups,” https://docs.aws.amazon.com/bedrock/latest/userguide/agents-api-schema.html — action-group operations and parameters are declared as OpenAPI schemas that define what the agent can invoke.↩︎
Claude Code documentation, “Configure permissions,” https://code.claude.com/docs/en/permissions — allow, ask, and deny rules in settings files, evaluated deny-first, so a matching deny rule blocks a call regardless of allow rules.↩︎
Open Policy Agent, https://www.openpolicyagent.org/ — CNCF project providing policy-as-code (Rego) enforced at control points such as Kubernetes admission, API gateways, and CI/CD pipelines.↩︎
Agent2Agent (A2A) protocol, Linux Foundation / a2aproject, https://github.com/a2aproject/A2A and specification materials at https://github.com/a2aproject/A2A/blob/main/docs/specification.md; Google Developers, “Developer’s Guide to AI Agent Protocols,” https://developers.googleblog.com/developers-guide-to-ai-agent-protocols/. Cited for agent discovery (Agent Cards), task-oriented collaboration across frameworks, and the architectural split between MCP (tools/data) and A2A (agent-to-agent). Spec and adoption are early; cited as evidence that horizontal topology is becoming a protocol problem.↩︎
Pi, the open-source coding agent by Mario Zechner (badlogic), ships a small core — loop, tools, context, sessions — and deliberately leaves MCP, sub-agents, plan mode, and permission systems as extensions, installed on demand as packages (https://pi.dev/). Cited as evidence that runtime skill acquisition is a shipping design pattern, and that an agent’s capability set changes over its life.↩︎
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou, “Self-Consistency Improves Chain of Thought Reasoning in Language Models,” ICLR 2023, arXiv:2203.11171, https://arxiv.org/abs/2203.11171. The method “first samples a diverse set of reasoning paths instead of only taking the greedy one, and then selects the most consistent answer by marginalizing out the sampled reasoning paths.”↩︎
Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch, “Improving Factuality and Reasoning in Language Models through Multiagent Debate,” ICML 2024, arXiv:2305.14325, https://arxiv.org/abs/2305.14325. Table 1: arithmetic 67.0 to 81.8 percent and grade-school math 77.0 to 85.0 percent, single agent versus three-agent debate over two rounds.↩︎
OpenAI, “Learning to Reason with LLMs,” September 12, 2024, https://openai.com/index/learning-to-reason-with-llms/. On the 2024 AIME exams, “GPT-4o only solved on average 12% (1.8/15) of problems. o1 averaged 74% (11.1/15) with a single sample per problem, 83% (12.5/15) with consensus.” Cited for reasoning accuracy improving with reasoning effort and multiple attempts.↩︎
Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig, “PAL: Program-aided Language Models,” arXiv:2211.10435 (ICML 2023), https://reasonwithpal.com. “Using a Python interpreter leads to more accurate results than much larger models”: on GSM8K it beat PaLM-540B with chain-of-thought prompting by eight points top-1, and on GSM-Hard — the same problems with larger numbers — it held above 60 percent where text-only reasoning fell to roughly 20 percent. Cited for arithmetic errors being addressable by delegating computation to a code tool.↩︎
Marilyn Strathern, “‘Improving Ratings’: Audit in the British University System,” European Review 5(3), 1997, pp. 305-321, restating Goodhart’s law — after Charles Goodhart’s 1975 observation on monetary policy and Keith Hoskins’s 1996 formulation — as “when a measure becomes a target, it ceases to be a good measure.”↩︎
AWS, “Amazon Bedrock AgentCore Evaluations is now generally available,” March 2026, https://aws.amazon.com/about-aws/whats-new/2026/03/agentcore-evaluations-generally-available. Thirteen built-in evaluators cover response quality, safety, task completion, and tool usage; online evaluation samples and scores live production traces, and on-demand evaluation supports regression testing in CI/CD pipelines.↩︎
OpenAI, “Introducing Structured Outputs in the API,” August 6, 2024, https://openai.com/index/introducing-structured-outputs-in-the-api/. “On our evals of complex JSON schema following, our new model gpt-4o-2024-08-06 with Structured Outputs scores a perfect 100%. In comparison, gpt-4-0613 scores less than 40%.”↩︎
NordPass, “Top 200 Most Common Passwords,” annual research; in the 2024 US list, “123456” and “password” hold the top ranks alongside “secret” and “qwerty” variants, https://nordsecurity.com/press-area/americans-most-common-password-is-secret- and https://nordpass.com/most-common-passwords-list. The word “password” and its close variants have appeared in the top ranks every year the research has run.↩︎
GitHub Blog, “Introducing fine-grained personal access tokens for GitHub,” https://github.blog/security/application-security/introducing-fine-grained-personal-access-tokens-for-github: classic tokens’
reposcope “provides broad access to all data in private repositories the user has access to, in perpetuity.”↩︎GitHub documentation, fine-grained personal access tokens: per-repository, per-permission grants with organization approval policies and maximum lifetimes; GitHub “recommends that you use fine-grained personal access tokens instead of personal access tokens (classic) whenever possible,” https://docs.github.com/en/authentication/keeping-your-account-and-data-secure/creating-a-fine-grained-personal-access-token and https://github.blog/security/application-security/introducing-fine-grained-personal-access-tokens-for-github/.↩︎
Kubernetes documentation, “Using RBAC Authorization,” https://kubernetes.io/docs/reference/access-authn-authz/rbac/. Rules grant verbs (get, list, watch, create, update, patch, delete) on resources, optionally pinned to specific resource names.↩︎
LangChain/LangGraph documentation, “Human-in-the-loop,” https://docs.langchain.com/oss/python/langchain/human-in-the-loop. Per-tool approval is configured through the interrupt_on setting, including the allowed_decisions option (for example, approve and reject only, with no edit).↩︎
Epic “break-the-glass” access: emergency override of restricted patient records requiring reason entry, with full logging and later review; ComplyDome, “HIPAA Emergency Access,” https://www.complydome.com/compliance-resources/hipaa-emergency-access-how-to-create-a-compliant-procedure-for-your-ehr-system-45-cfr-164312a2ii (“Epic uses ‘break-the-glass’ alerts, requiring reason entry”). The healthcare pattern is the ancestor of the general break-glass procedure.↩︎
Vincent C. Hu, David Ferraiolo, Rick Kuhn, et al., Guide to Attribute Based Access Control (ABAC) Definition and Considerations, NIST Special Publication 800-162 (January 2014; updated August 2019), https://csrc.nist.gov/pubs/sp/800/162/upd2/final. NIST defines ABAC as an access control method in which authorization “is determined by evaluating attributes associated with the subject, object, requested operations, and, in some cases, environment conditions against policy, rules, or relationships that describe the allowable operations for a given set of attributes.” AWS IAM documentation, “Attribute-based access control,” https://docs.aws.amazon.com/IAM/latest/UserGuide/introduction_attribute-based-access-control.html (“ABAC’s attribute system that provides both high user context and granular access control… Because ABAC is attribute-based, it can perform dynamic authorization for data or applications that grants or revokes access in real time”).↩︎
AWS Cedar policy conditions:
when { ... }(andunless { ... }) blocks evaluate attributes of the request at decision time, so a policy can grant an action only under specified conditions — the shipping mechanism for situation-dependent authority. AWS Cedar documentation, https://docs.aws.amazon.com/cedar/; see also the Cedar policy example in 7.2.1.↩︎Break-glass procedure: “a pre-planned, highly secure, and temporary method for granting elevated, often administrative, access to a small, authorized group of individuals during a major incident” (Cloudanix, “Break-Glass Procedure: Emergency Access in the Cloud,” https://www.cloudanix.com/learn/break-glass-procedure-emergency-access-for-critical-resources). Under the NIST Cybersecurity Framework, break-glass access falls under the Respond and Recover functions and must be pre-approved and documented.↩︎
HIPAA Security Rule, 45 CFR § 164.312(a)(2)(ii): covered entities must “establish and implement procedures for emergency access” to electronic protected health information. Yale’s implementation guide describes break-glass as “based upon pre-staged ‘emergency’ user accounts… created in advance to allow careful thought to go into the access controls and audit trails associated with them,” https://hipaa.yale.edu/security/break-glass-procedure-granting-emergency-access-critical-ephi-systems.↩︎
Anthropic, “Statement on the US government directive to suspend access to Fable 5 and Mythos 5,” June 12, 2026, https://www.anthropic.com/news/fable-mythos-access. Anthropic said a Commerce Department “is informed” letter required blocking foreign nationals; because nationality could not be filtered at the API layer, it disabled the models worldwide. The directive was not a published BIS rule. Reuters reporting: https://www.aljazeera.com/news/2026/6/13/us-orders-anthropic-to-disable-ai-models-for-all-foreign-nationals.↩︎
Zapier 2026 enterprise AI survey (542 U.S. executives with active AI vendor contracts): “81% say they’re at least a little concerned about their organization’s dependency on specific AI vendors”; “for nearly half (47%), the impact would be more serious — at least one key business function would stop working correctly” if their primary AI vendor went down; only 6% believe they could switch providers without material operational disruption. https://zapier.com/blog/ai-vendor-lock-in-survey.↩︎
Privileged access management practice on rehearsing emergency access: “The override should be tested regularly — unpracticed emergency access is dangerous” (hoop.dev, “Secure Break-Glass Access in Privileged Access Management,” https://hoop.dev/blog/secure-break-glass-access-in-privileged-access-management).↩︎