Download Cosmonic Desktop (Beta)

Policy enforcement points for AI agents:
every action gets checked

Agents use six categories of resources. All of them need to be checked.

Anthropic, October 2025. A Claude Code deny rule said: never read this file. A symlink, a shortcut whose name points somewhere else, got past it, because the rule checked the name and not the file it led to1.

Google, November 2025. Antigravity's file guard refused to read .env. So the prompt-injected agent ran cat .env in the terminal, and the secrets left through webhook.site, a site on the default allowlist.

OpenAI, September 2026. A training agent's web proxy blocked its requests. So it hid its questions inside DNS names2 and got a reply from a public chatbot: "The capital of France is Paris." Nobody was checking the DNS resolver.

In all three cases, each guard did its narrow job: the deny rule matched, the file tool refused, the proxy blocked. The agent simply took another path, because the resource itself had no guard. A policy enforcement point (PEP) is a guard every request must pass, on every path. To sandbox an agent properly, put one in front of each resource the agent can touch, not at one edge of the network. That is what zero trust asks for.

Policy enforcement points are one of The 6 Principles of Secure Agent Platforms, building on our guide to choosing an agent sandbox. They build on deny by default and least authority, because, in the series' building analogy, a guard needs a door: those two decide which doors exist. Then every door gets a guard who checks each trip through it, and checks which room you actually enter, not what the sign on the door says.

Policy enforcement points around an AI agent: one guard each for files, network, commands, secrets, config and the host control plane.

What is a policy enforcement point?

A policy enforcement point (PEP) is the component that sits in front of a resource and allows or blocks each request, based on a decision made by policy. NIST's glossary, quoting SP 800-162, says a PEP "enforces policy decisions in response to a request from a subject requesting access to a protected object."

The terms PEP and policy decision point (PDP) come from policy-based networking (RFC 2753, 2000), and NIST SP 800-207, Zero Trust Architecture3 (August 2020) made them the vocabulary of zero trust. For an AI agent, the PEP is whatever stands between the agent and a file, a host, a program or a key: a kernel hook, an egress proxy, or a Wasm host. It checks the real resource (the resolved file, the real destination), not its name.The idea is almost fifty years older.

One guard, every door: where the PEP idea comes from

In 1972, James P. Anderson wrote a computer security planning study that described one guard standing between every program and everything it can touch: the security reference monitor. Operating system designers still build guardrails in this model. A security reference monitor is a guard that checks every access to every object, cannot be tampered with, and is small enough to verify.

Picture one guard at every door who never takes a break, can't be bribed, and works from a rulebook you can read in an evening. Anderson's three properties:

  1. Always invoked. Every access goes through it, every time.
  2. Tamperproof. The program it watches can't change it or switch it off.
  3. Small enough to verify. You can test it, or even prove it, completely.

You already run versions of it. Windows has a kernel component called the Security Reference Monitor that checks every object access. Linux has LSM hooks, where SELinux, AppArmor and Landlock plug in. In Wasm, the host checks every import call a component makes.

What is complete mediation?

Complete mediation is the security principle that every access to every resource is checked for authority, every time, not just the first time. Saltzer and Schroeder4 named it in 1975 among their eight design principles. The OWASP Developer Guide still teaches it: authority must not be "circumvented in subsequent requests," so the check runs "upon every request for the object."

PDP vs PEP in zero trust: who decides and who enforces

The policy decision point (PDP) decides whether a request is allowed; the policy enforcement point (PEP) carries out that decision at the door. NIST SP 800-207 splits them so the enforcing part stays simple, and no request is trusted because of where it sits on the network.

RoleQuestion it answersIn an agent sandbox
Policy decision point (PDP)Is this request allowed by the policy?The policy file, the allowlist, a human answering an approval prompt
Policy enforcement point (PEP)Did this exact request get a yes?A kernel hook (Landlock, Seatbelt), an egress proxy, the Wasm host

In NIST's model, the PDP pairs a policy engine (it makes the call) with a policy administrator (it tells the PEP to open or close the path). Access control frameworks add a policy information point (PIP) that supplies facts and a policy administration point (PAP) where people write rules.

An AI model can be neither the PDP nor the PEP. A model reads untrusted text, and untrusted text can become instructions5. If the model decides, a poisoned README decides. If the model enforces, the guard can be talked out of its job. So the harness, the OS or the runtime enforces, and a human or a policy file the agent can't edit decides.

PDP vs PEP: the policy decision point decides, the policy enforcement point enforces, and the AI model is neither.

Three rules for an enforcement point

An enforcement point has to meet three rules, which are Anderson's properties restated for agents:

  1. Every access crosses it. No first-time-only checks, no cached yes.
  2. Nothing goes around it. No second path, no other tool, no off switch the agent can reach.
  3. It is small enough to verify. You can read it, test it and trust it.

The DNS escape broke rule 2: DNS went around the proxy. OpenAI wrote that its "safety case assumed that the model could not access the live internet," then paused training, evaluation, and inference with tool use for its most capable models.

Five strategies follow: a PEP in front of every resource (zero trust); decision (PDP) kept apart from enforcement (PEP); a guard modeled on the security reference monitor; a minimal trusted base; and canonicalize before checking, which means turning a name into the one real thing it points to, then checking that.

TOCTOU and name-vs-object bugs

TOCTOU (time of check to time of use) is a race where a resource changes between the moment a guard checks it and the moment it is used. Its close cousin is the name-vs-object bug: the guard checks a name, but the program uses whatever the name points to. Garfinkel 2003 found both in early sandboxes. Today, we find agent tools repeat them:

  • Claude Code's allowed-folder check compared text prefixes, so access to claude_code also allowed claude_code_evil (CVE-2025-54794, Aug 2025).
  • Its deny rules checked a symlink's name, not its target (CVE-2025-59829, Oct 2025).
  • Symlinks and look-alike folder names escaped Anthropic's Filesystem MCP server (EscapeRoute, CVE-2025-53109 and -53110, Jul 2025).

The fix is to resolve and check in one step, inside the kernel or runtime, so nothing can change in between. On Linux, openat2(2) with RESOLVE_BENEATH refuses any path that climbs out of a starting directory, and RESOLVE_NO_SYMLINKS refuses symlinks. A WASI preopened directory does the same in Wasm. seccomp can't do it alone: its BPF filters can't dereference pointers, so they never see the path a syscall is asked to open6.

Capabilities close the same gap from the other side, because a capability is unforgeable: a program can't guess one, build one from a name or string, or copy one it was never given7. So the check happens at the moment the capability is granted, and a lookalike name can't trick the runtime later. A name-based check has to resolve a path or hostname on every call and get it right every time; a Wasm component can call only the imports the host linked, and the runtime checks every resource handle it presents.

Complete mediation checks the resolved file, not its name, which stops symlink and prefix bypasses.

Small enough to verify

Every real guard bends rule 3 to a certain extent, which is an argument for defense in depth. On Linux the guard is the kernel: hundreds of system calls, tens of millions of lines. Windows trusts its whole kernel; macOS trusts all of XNU. A VM trusts its hypervisor and device emulation, which in QEMU is large; a microVM such as Firecracker cuts that to a handful of virtio devices. gVisor puts a smaller guard in front of a container's shared kernel. A Wasm runtime's guard is small, but its compiler is part of what you trust.

At the far end, seL4 (Klein et al. 2009) is a kernel small enough to prove correct. You can't shrink every guard that far, so you stack different ones: defense in depth.

The six resources an AI agent sandbox must guard

An AI agent may touch many types of resource and we'll focus on the six most common kinds of resource where each needs its own enforcement point: files, the network, commands, secrets, config files that become code, and the host's control plane.

Here is the part most teams miss. Operating system sandboxes support guards for all six. Linux has Landlock, seccomp, bubblewrap and cgroups. macOS has Seatbelt and a Network Extension filter. Windows has AppContainer, Job Objects and App Control. But none of them is configured for a program by default.

A plain process starts with all six open: it can read your files and keys, call any host, run any program, write your git hooks and reach your Docker socket. You add each guard yourself, tool by tool. Every guard you forget is an open door, and that is where most of the incidents below come from.

These rules are hard to get right.

Why not write the perfect rules up front? Because you can't. Rice's theorem8 says no program can decide, in general, what another program will do, so nobody can predict every file, host or syscall an arbitrary program will need. It is why few apps ship a seccomp profile, even though seccomp has existed since 2005 and its filter mode since 2012. Agents make it harder, because their behavior changes with every prompt.

So stop predicting. Start from nothing (deny by default), grant each resource on purpose, and put the guard on the resource itself, where it sees every access from every tool.

ResourceWhat its PEP must checkRated Yes in the matrix
4a FilesystemThe resolved path, below every toolLinux sandbox, macOS sandbox, VM, microVM, Wasm
4b Network egressDestination, method, path and credentialWasm
4c CommandsThe real program, at exec timeLinux sandbox, macOS sandbox, Wasm
4d SecretsWhether the sandbox ever sees the keyWasm
4e Config and hooksWrites to files that run laterLinux sandbox, macOS sandbox
4f Host control planeSockets, daemons and localhostLinux sandbox, Windows sandbox, microVM, Wasm

A plain process rates No on all six. The OS sandboxes are rated as configured: they earn a Yes only after you set them up and enforce them, which, for general programs, is very hard to do.

Filesystem (4a)

Which files the agent can read, write or execute. Check the resolved path (after symlinks and ..), below the agent's tools. This is harder to get right than it may appear: Codex CLI accepted a working directory chosen by the model as its writable root, so the AI could write outside the project (CVE-2025-59532, Sep 2025). Google Antigravity guarded its read tool, not the file (PromptArmor, Nov 2025; see below).

Controls: Landlock and bubblewrap bind mounts (Linux); Seatbelt file rules and TCC protected folders (macOS); AppContainer ACLs and brokered file pickers (Windows); WASI preopened directories9. At the agent layer, allow writes only in the workspace and hide .env, ~/.ssh and ~/.aws.

Network egress allowlist (4b)

Where the sandbox can send data. Allow specific destinations, and also check method, path and credential, because an allowed domain is not an allowed destination. In "Claude Pirate" (Oct 2025), the allowed api.anthropic.com carried a user's data to the attacker's own account. Only Wasm rates Yes, and only with a host that enforces each component's allowed hosts, as CNCF wasmCloud does: an import grants the whole interface, so the destination check lives in the host. OS sandboxes, containers and microVMs rate Partial, because the kernel filters by IP or port and a domain allowlist needs a proxy.

Egress has the most incidents: the DNS escape, Antigravity via webhook.site (Nov 2025), sandbox-runtime's empty allowlist (CVE-2025-66479, Dec 2025, low severity), a null-byte hostname bypass (May 2026, researcher report; vendor fix not independently confirmed), Claude Code DNS exfiltration (CVE-2025-55284, Jun 2025), Copilot Studio SSRF (CVE-2024-38206, Aug 2024) and a Custom GPT Actions SSRF (Nov 2025).

Commands and execution (4c)

Which programs the sandbox can start. Parse the real command and check it when it runs. Gemini CLI checked only the first word, so once grep was approved, grep … && curl ran silently (Jul 2025). Claude Code auto-approved echo, and a crafted echo ran any command (CVE-2025-54795, Aug 2025). Codex CLI listed git show as safe, but --output wrote into .git/config (GitPwned, Feb 2026). Cursor trusted built-ins like export even with an empty allowlist (CVE-2026-22708, Jan 2026).

Controls: Landlock execute rights, AppArmor exec rules and noexec mounts (Linux); Seatbelt process-exec rules (macOS); App Control for Business and Smart App Control (Windows). WASI has no process spawning at all; tools are typed host functions.

Secrets and credentials (4d)

Who can read keys and tokens. The best secret is one the sandbox never sees: the host attaches it to an approved request on the way out. A repo's settings file pointed Claude Code's API traffic at an attacker's server, leaking the user's key before the trust prompt (CVE-2026-21852, Jan 2026). At Smithery (Oct 2025), a path-traversal bug exposed an admin token covering 3,000+ hosted MCP servers and their users' keys.

Controls: a proxy that injects credentials only for allowed hosts (Claude Code's sandbox10 supports credential masking at its proxy, with TLS termination); a separate user account for the agent, so it can't decrypt the credentials you have stored (Windows); Keychain item accessibility (macOS); kernel keyrings (Linux); workload identity and the Secrets Store CSI Driver (Kubernetes); host-held, per-component wasmCloud secrets (Wasm).

Built-in secret stores (Keychain, Credential Manager, GNOME Keyring) keep secrets from other users, not from programs running as you. An agent running as you can usually read them. A sandbox helps only if it closes the path to the store: no D-Bus socket inside bubblewrap, a Seatbelt rule that denies the Keychain service, a separate Windows user, or a seccomp filter on the kernel keyring (Docker's default profile11 already blocks it).

Config and hooks: files that become code (4e)

Files that something outside the sandbox runs later: git hooks, editor settings, agent hooks, shell rc files. A write to one is an escape that waits for the next launch. Code inside Claude Code's bubblewrap sandbox wrote .claude/settings.json hooks that later ran outside it (CVE-2026-25725, Feb 2026). A prompt-injected Cursor agent wrote git hooks that git later ran on the host (CVE-2026-26268, Feb 2026), and another wrote a Stop hook that Cursor ran when the turn ended (CVE-2026-48124, Jul 2026).

Controls: bubblewrap --ro-bind over .git, .vscode and agent config (Linux); Seatbelt rules denying writes to config paths (macOS); App Control script enforcement (Windows); a Wasm preopen that leaves out .git. Load no project config before trust, re-prompt when its hash changes, and know where git looks ($GIT_DIR/hooks, core.hooksPath, per githooks).

Host control plane: sockets, daemons and localhost (4f)

Management APIs reachable from inside the sandbox: the Docker socket, local dev servers, metadata endpoints. Reaching one usually means full control. Any container could call Docker Desktop's engine API without auth and mount the host's C: drive (CVE-2025-9074, Aug 2025). Sandboxed Codex, Cursor and Gemini CLI agents could still call Docker Desktop and mount the user's home folder (Pillar Security, Jul 2026). Anthropic's MCP Inspector ran an unauthenticated local server, so any website the developer visited could run commands (CVE-2025-49596, Jun 2025).

Controls: a network namespace and Landlock's abstract Unix socket scoping (Linux); Seatbelt rules denying localhost, Unix sockets and mach-lookup (macOS); AppContainer loopback isolation, on by default (Windows); never mount docker.sock12; no wasi:sockets import (Wasm). Local MCP servers need auth too; see MCP security.

Guard the resource, not the tool

The most common failure in the incident record is a check on one tool while another tool reaches the same resource. Antigravity's guard stopped the read tool from opening .env, so the agent ran cat .env in the terminal. Gemini CLI checked a command's first word, so the second command ran free.

The vendors document the same gap in their own designs. Claude Code's sandboxing docs say: "The sandbox covers shell commands only. Claude's file tools, MCP servers, and hooks run outside it." Those run under permission rules, which is where the prefix and symlink bugs lived. On native Windows, Claude Code runs commands unsandboxed. OpenAI's Codex docs describe Seatbelt on macOS, bubblewrap and seccomp on Linux, a native Windows sandbox, and three modes: read-only, workspace-write and danger-full-access. Both guard the shell well. Neither guards every resource.

Escape hatches are another way around a guard. Claude Code can retry a blocked command outside the sandbox (dangerouslyDisableSandbox), and its excludedCommands run with no filesystem rules and no network proxy. Docker has --security-opt seccomp=unconfined and --privileged. When a build breaks, people reach for these, and rule 2 is gone. The agent asks too: blocked from npx by bubblewrap, Claude Code asked to turn its sandbox off (Mar 2026).

The fix is to mediate in the OS or runtime, below every tool, so the read tool, cat, an MCP server and a hook all meet the same guard.

Policy enforcement points across the sandbox types

Every sandbox type has a guard, and every one trusts something large behind it, so all eight sandbox types in the matrix rate Partial on this principle; a plain process rates No.

As elsewhere, we call these "sandbox types" after the conventional usage, though containers and VMs are not sandboxes on their own.

ProcessNoNoNoNoNoNo
Its guard
The kernel checks every access13
What it trusts
A huge base; policy is per user, not per program
Controls to add
Run under an OS sandbox
Linux sandboxYesPartialYesPartialYesYes
Its guard
LSM hooks and seccomp: LSM hooks check each access to an object, and seccomp filters each system call
What it trusts
The kernel
Controls to add
Landlock; seccomp-bpf; AppArmor or SELinux
macOS sandboxYesPartialYesPartialYesPartial
Its guard
The sandbox kernel extension checks every operation14
What it trusts
XNU; Seatbelt is deprecated
Controls to add
Seatbelt; SIP; TCC
Windows sandboxPartialPartialPartialPartialPartialYes
Its guard
The Security Reference Monitor checks every object access
What it trusts
A huge kernel base
Controls to add
AppContainer; App Control; VBS / HVCI
ContainerPartialPartialPartialPartialPartialPartial
Its guard
seccomp and LSM profiles on by default (Docker)
What it trusts
The shared host kernel
Controls to add
Custom seccomp profile; gVisor as a smaller guard
gVisorPartialPartialPartialPartialPartialPartial
Its guard
A user-space kernel (the Sentry) handles each system call before the host kernel sees it
What it trusts
The Sentry, under a tight seccomp filter
Controls to add
Run runsc as the OCI runtime; keep it patched
VMYesNoNoPartialPartialPartial
Its guard
Hardware traps privileged instructions and device access
What it trusts
The guest OS inside; large device emulation15
Controls to add
Minimal virtual hardware; sandbox the VMM
microVMYesPartialNoPartialPartialYes
Its guard
Hardware traps through a small VMM with a few virtio devices
What it trusts
The guest kernel inside
Controls to add
Jailer: seccomp, cgroups, chroot
WasmYesYesYesYesPartialYes
Its guard
The host checks every import call, and the runtime bounds-checks every memory access (Wasmtime)
What it trusts
The compiler
Controls to add
Small host interface; fuzzed compiler; stay patched

A plain process guards none of the six resources. A configured Linux or macOS sandbox guards files, commands and config, and Linux adds the host control plane. Wasm guards five of six by default and leaves config and hooks to you. For every principle side by side, see the platforms vs principles matrix.

How to put enforcement points around an agent

  1. List the resources the agent touches. Walk all six, 4a to 4f.
  2. Pick one PEP per resource, below the tools. A kernel rule, a proxy or a runtime import, never a check inside one tool.
  3. Resolve names before checking. Check the real file and the real destination, in one step.
  4. Keep the decision apart and read-only. Policy files and approvals live where the agent can't write (see 4e).
  5. Deny by default, then grant on purpose. Start from nothing (deny by default).
  6. Log every decision. Allowed and denied, with the resource it touched.

On Cosmonic's platforms, tools run as WebAssembly components, so each file, network or secret access is an import the host checks, below every tool. Cosmonic Desktop, a free app for macOS, Windows and Linux, shows what a component can reach before it runs, and its MCP server lets coding agents such as Claude Code and Codex deploy tools into that deny-by-default sandbox. Cosmonic Control runs the same components on Kubernetes, with admission controller integration and RBAC, egress denied by default with DNS controls, secrets from Kubernetes Secrets and External Secrets, and OpenTelemetry for the decision log. It installs on premises and in air-gapped networks. Wasm has limits: it can't run arbitrary Linux binaries, so the agent harness itself still needs an OS sandbox or VM around it.

To see how this principle fits the other five, go back to The 6 Principles of Secure Agent Platforms, or use the main guide's six-step method to choose a sandbox.

Put the enforcement point outside the agent

Public beta

Run tools as Wasm components in Cosmonic Desktop, where every file, network and secret access is a declared import the host checks. Cosmonic Control adds Kubernetes admission controller integration, RBAC and egress denied by default.

Cosmonic Desktop's Inspect view: a component's declared interfaces and network egress drawn out, with everything else denied by default.

Frequently asked questions

What is a policy enforcement point?
A policy enforcement point (PEP) is the checkpoint in front of a resource that allows or blocks each request, carrying out decisions made by a policy decision point. NIST SP 800-207 defines it for zero trust. In an AI agent sandbox, it is a kernel hook, an egress proxy or a Wasm host.
What is the difference between a PDP and a PEP?
The policy decision point (PDP) decides whether a request is allowed, and the policy enforcement point (PEP) carries out that decision. Keeping them apart keeps the enforcer simple enough to trust. For AI agents, the model is neither: a policy file or a human decides, and the OS or runtime enforces.
What is complete mediation in cyber security?
Complete mediation is the security principle that every access to every resource is checked for authority, every time, with no cached answer and no bypass path. Saltzer and Schroeder named it in 1975. For an AI agent, it means every file read, network call and command crosses a guard.
What is a security reference monitor?
A security reference monitor is a guard that checks every access to every object, cannot be tampered with, and is small enough to verify. James P. Anderson described it in 1972. Real operating system examples include the Windows Security Reference Monitor, Linux LSM hooks and the import checks a Wasm host makes.
What is a TOCTOU vulnerability?
A TOCTOU (time of check to time of use) vulnerability is a race where a resource changes between being checked and being used. In AI agents it shows up as symlink and path-prefix bypasses. The fix: resolve and check in one step, in the kernel or runtime.
Is a policy enforcement point part of zero trust?
Yes. NIST SP 800-207 puts a policy enforcement point in front of every resource, has a policy decision point decide each request, and trusts nothing because of its network location. For an AI agent, that means one guard per resource, not one firewall at the edge.

Further reading

  1. CVE-2025-59829: Claude Code deny rule bypassed through a symlink. GitHub security advisory. github.com ↩
  2. OpenAI, “An agent used DNS to reach an external chatbot”, September 2026. alignment.openai.com ↩
  3. NIST SP 800-207 (2020), Zero Trust Architecture. csrc.nist.gov ↩
  4. Saltzer and Schroeder (1975), “The Protection of Information in Computer Systems”, Proceedings of the IEEE. doi.org ↩
  5. Greshake et al. (2023), “Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection”. doi.org ↩
  6. Linux kernel documentation, Seccomp BPF (seccomp_filter). docs.kernel.org ↩
  7. Dennis and Van Horn (1966), “Programming Semantics for Multiprogrammed Computations”. doi.org ↩
  8. Rice (1953), “Classes of Recursively Enumerable Sets and Their Decision Problems”. doi.org ↩
  9. Wasmtime documentation, security. docs.wasmtime.dev ↩
  10. Anthropic, Claude Code sandboxing documentation. code.claude.com ↩
  11. Docker, seccomp security profiles. docs.docker.com ↩
  12. Docker, “Protect the Docker daemon socket”. docs.docker.com ↩
  13. Linux manual page, credentials(7). man7.org ↩
  14. macOS manual page, sandbox-exec(1). keith.github.io ↩
  15. QEMU documentation, security. qemu.org ↩
  16. NIST SP 800-207 (2020), Zero Trust Architecture. doi.org ↩