1.8 Policy and permission systems
If identity answers who, policy answers under what rules — and the two are separate categories with separate products, even though the industry sells them as one “access” problem. A policy engine is a program that evaluates a request against rules and returns a verdict, independent of any single application, and the defining property of this layer is operational: every protected action is mediated by deterministic code the model cannot bypass. The decision point may run as a sidecar or gateway, a hosted service, an embedded library, or a WebAssembly module — what matters is the mediation, not the process boundary. The root of an orchestrated system is an inference engine — a large language model, a probabilistic component that interprets. A rule written into that component’s instructions is a directive, and directives have real uses: guidance, priorities, how to behave when the situation is ambiguous. But a directive cannot enforce. The model can weigh the rule against other instructions, misapply it, or decide this case is different, and nothing in the mechanism stops it. If the intention is to enforce a rule — never, under any circumstances — then the rule has to be evaluated outside the model’s reach, by deterministic code the model cannot talk its way past. That is the difference an orchestrator has to hold: policy implemented by an inference engine is a suggestion to a probabilistic interpreter; policy implemented by a policy engine is a deterministic verdict from code that behaves the same way on every request. Chapter 3 separated directives from credentials for exactly this reason; here that separation becomes running architecture.
1.8.1 The policy engine field
The 2026 field divides by lineage:
- Rego and OPA — Open Policy Agent, the CNCF-graduated general-purpose policy engine, evaluates fine-grained rules over which tools an agent can call, with what parameters, and under what conditions, when it is placed in the execution path; its Rego language is widely used for policy-as-code, embedded in gateways that check agent tool calls at request time.
- Cedar and its descendants — an AWS-backed general-purpose authorization language, in production in Amazon services such as Verified Permissions and in Bedrock AgentCore’s agent-to-tool authorization, where a policy engine attached to a gateway intercepts and evaluates every request an agent makes of a tool; the open-source Cedarling embeds the same engine locally for edge and sidecar use.
- Relationship-based access control — SpiceDB, AuthZed, OpenFGA, Auth0 FGA, the family descended from Google’s Zanzibar paper, which stores authorization as relationship data (who is connected to what) rather than as rules, and answers permission questions by graph traversal.
- Gateway enforcement — the pattern where none of the above runs inside the agent at all: an identity or policy gateway sits in the path between agent and resource, applying token exchange, scoping, and verdicts per tool call — the architecture Microsoft, Strata (now Rubrik), and the MCP enterprise-authorization extension have all converged on.
- Kernel-isolated runtimes with verifiable policy — NVIDIA’s OpenShell, open-sourced under Apache 2.0, runs each agent in a sandbox and compiles the operator’s file, network, tool, process, and credential limits into a policy the runtime checks before the agent starts and enforces as it works; the companion Sentry design extends enforcement to BlueField-4 DPUs watching out of band. Chapter 14 evaluates the stack against real incidents.35
1.8.2 A policy engine, concretely
The field list above stays abstract until you see one engine up close, and this is the one layer in the survey where the concrete mechanics are worth learning now, because policy engines are new to most agent builders.
What a policy contains. In Cedar, the policy language AWS built for general-purpose authorization, a policy is a short statement with an effect (permit or forbid), a scope (who, doing what, to which resource), and optional conditions. A complete permission envelope for one class of agent looks like this:
permit (
principal in Agent::"doc-agents",
action == Action::"read",
resource in Document::"contracts"
);
forbid (
principal in Agent::"doc-agents",
action == Action::"merge",
resource
);
In words: agents in the doc-agents group may read anything in the contracts collection, and no other policy can turn that into merge rights, because a forbid cannot be overridden.37 A policy set is a collection of these statements, and on every request the engine evaluates all of them and returns a verdict. The recommended authorization contract is allow or deny, with no fuzzy middle — the same answer no matter which model is asking. (The engine can also return structured JSON, or an undefined result when no policy matches and no default is defined, which is why a policy set should always define its default.)
How policies are configured. Policies are written as code, in the repository, next to the application that depends on them: reviewed in pull requests, versioned in git, and tested. Both engines ship test frameworks, so the author writes the requests that should be allowed and the requests that should be denied, and the suite fails if the policy gets either one wrong.
How policies are distributed. The engine runs as its own process, typically a sidecar next to the application or a gateway sitting between the agent and its tools — and that placement is the whole point: the deterministic component shares a process boundary with neither the agent nor the model, so there is no instruction the model can give that changes what the policy engine does. Policies reach the engine without redeploying the application. In OPA the unit of distribution is the bundle: the policy set, plus any data it evaluates against, is packed into a compressed archive and published to a server, and bundle signing is a supported option — when it is configured, every running OPA instance polls that server on a configured interval, verifies the signature, and swaps in the new policy set when one arrives. A rule change ships in minutes and rolls back by publishing an earlier version, and the application binary is never touched. Cedar deployments follow the same shape: the policy set is attached to the gateway — AWS’s Bedrock AgentCore, for instance, evaluates every agent-to-tool request against Cedar policies attached at the gateway, outside the agent entirely.
How policies are managed. Because policies are code, the lifecycle is the software lifecycle: propose in a branch, review, test, publish a signed bundle. And the engines can report back what they are running — when configured to do so. OPA can emit decision logs — which request, which policies matched, what the verdict was — as a configuration option, not a default. Cedar itself evaluates policy; the integrating application or managed service must record the decision, and in AWS Verified Permissions, authorization data events require explicit CloudTrail configuration. Where those logs exist, they are the audit half of the layer. When a regulator asks why an action was allowed on a given afternoon, the answer is not a screenshot of a configuration screen; it is the policy version that was live at the time and the logged verdict it produced.
That is the whole loop: written like code, reviewed like code, deployed like configuration, enforced outside the model, logged on every request. For learning it hands-on, the two starting points are the Cedar policy language reference (docs.cedarpolicy.com) and the OPA documentation on policy distribution and bundles (openpolicyagent.org/docs/management-bundles). Both are readable in an afternoon, and building one small policy set with its own test suite is the fastest way to make this layer concrete.
1.8.3 What a policy system can see
For orchestration, the useful question about a policy system is what its verdicts can see. Can the policy condition on the combination of identities in a run — the human delegator and the agent — or only on whichever single principal presented the token? Can it condition on task context, spend so far, time of day, or the history of the session? Can it escalate rather than merely allow or deny? AWS’s reference implementation for multi-agent delegation chains — Cedar policies layered by originating user, agent-to-agent delegation, and cross-cutting constraints — is the most complete public example, and even it decomposes the authority matrix into single-principal checks rather than recording the matrix itself. The gap named in the previous section shows up here as a product requirement: the policy layer can evaluate who may act; no product yet evaluates and records what combination of granted authorities produced this decision.
1.8.4 Where policy shows up for an orchestrator
Policy is the layer where the orchestrator’s authority decisions become executable. The encounters are concrete and recurring: writing the rules — translating “this agent may read the repository but never merge” into a policy the engine evaluates on every call; placing the engine — deciding whether the check runs in a gateway between agent and tool (verifiable, adds a hop) or inside the agent’s runtime (fast, trusts the thing it constrains); and testing the rules — running the scenarios that should pass, and especially the ones that should not. A policy that has never been tested for denial is not a policy; it is a wish. Testing refusal matters as much as defining the rule: the tests prove the request is actually refused when the policy is violated, and that the people responsible for running the system are notified — the system itself receiving a deny verdict is not the same as anyone finding out.
The connection to the rest of the book is direct. Chapter 3’s permission envelope becomes a Cedar or Rego policy; Chapter 7’s delegation architecture becomes the layering of policies the AWS reference implementation sketches (originating user, agent-to-agent, cross-cutting constraints); Chapter 9’s failure modes include the policy that was correct at design time and wrong at run time because the world moved and the rules did not. Policy expressed as code and evaluated outside the model makes violations deterministically blockable when every execution path is mediated — not impossible, since coverage gaps, stale policy or data, bugs, and bypass paths remain — while policy expressed as instructions to the model makes them merely unlikely.
NVIDIA OpenShell, open source under Apache 2.0, https://github.com/NVIDIA/OpenShell; announcement and reference design: “NVIDIA Launches Open Agent Safety Platform to Secure Agents From Testing to Deployment,” September 28, 2026, https://nvidianews.nvidia.com/news/open-agent-safety-platform, and NVIDIA Technical Blog, “NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring,” https://developer.nvidia.com/blog/nvidia-open-agent-safety-platform-a-reference-for-continuous-in-silicon-agent-monitoring/. Announced September 28, 2026 — days before this chapter’s September 2026 survey cutoff. Cited for the runtime-layer pattern, not as evidence of maturity. Verified September 29, 2026.↩︎
Policy systems as of August 29, 2026: Open Policy Agent, https://openpolicyagent.org/docs/integration; Cedar, https://docs.cedarpolicy.com; Cedarling, https://janssenproject.github.io/developer-docs/cedarling/cedarling/index.html; SpiceDB, https://authzed.com; and OpenFGA/Auth0 FGA. AWS’s layered Cedar example for multi-agent delegation is at https://github.com/aws-samples/sample-cedar-agentic-ai-authorization. Zanzibar-lineage relationship systems and general policy engines solve related but distinct problems.↩︎
Cedar syntax and deny semantics: https://docs.cedarpolicy.com/policies/syntax-policy.html and https://docs.cedarpolicy.com/policies/policy-examples.html; the example in the text is adapted from those patterns. OPA bundle distribution and signing: https://openpolicyagent.org/docs/management-bundles; deployment options: https://openpolicyagent.org/docs/integration. OPA and Cedar both provide testing support. AWS Bedrock AgentCore Policy is discussed with the policy-engine survey above.↩︎
- 1.1 The technology under discussion
- 1.2 The survey: connecting the taxonomy to the market
- 1.3 Inference engines and models
- 1.4 Retrieval systems
- 1.5 Memory systems
- 1.6 Tool interfaces and actuation
- 1.7 Identity and access control
- 1.8 Policy and permission systems
- 1.9 Durable orchestration engines
- 1.10 Task agents
- 1.11 Everything is called an agent
- 1.12 Multi-agent coding systems
- 1.13 The architecture of participation
- 1.14 The chapter in one picture