Download Cosmonic Desktop (Beta)

Deny by default:
why AI agents should start with nothing

How to ensure that AI agents get only the access you grant.

Without deny by default, an agent told to fix a failing test can also read ~/.aws/credentials, because nothing said it couldn't. Deny by default flips that: the agent starts with no access, and you add only what the task needs.

Deny by default is one of The 6 Principles of Secure Agent Platforms. In the series' building analogy, the room is built with no doors at all, so an agent with zero capabilities can't touch anything, even by accident. We will only specifically grant (allow list, not a deny list) the capabilities our agent must have to perform its function:

Deny by default compared with default allow for AI agents: a closed door that opens only for granted access

What does deny by default mean?

Deny by default means every request is refused unless a rule explicitly allows it. NIST defines it for firewalls: "block all inbound and outbound traffic that has not been expressly permitted by firewall policy"1.

An agent touches more than network traffic, so apply the same rule to all resources it can reach, here we will discuss the 6 big categories: files, network, commands, secrets, config and hooks, and the host's control plane (the policy enforcement points page covers each).

In 1975, Saltzer and Schroeder called it fail-safe defaults: "Base access decisions on permission rather than exclusion"23. If you forget an allow rule, something breaks and someone notices. If you forget a deny rule, the attack works and nobody notices.

Why default deny fails safe: a forgotten allow rule breaks a feature, a forgotten deny rule lets an attack through

Default deny, implicit deny and fail-safe defaults: one idea, three names

Implicit deny is the invisible last rule in a firewall or access list that drops anything no earlier rule allowed. It blocks all traffic that matches no allow rule. In nftables, you get it by giving a base chain policy drop4.

Where deny by default fits

Isolation draws the wall. Deny by default builds the room with no doors, least authority cuts only the doors the job needs, and policy enforcement points guard each door.

Deny by default vs default allow

Default allow starts open and blocks known bad things. Deny by default starts closed and opens known good things.

Default allowDeny by default
Starting pointEverything is reachableNothing is reachable
A forgotten ruleLets something throughBreaks a feature
What an attacker needsOne path you didn't blockA path you granted
What fails firstSecurityA feature
UpkeepChase every new trickAdd grants as needs appear
ExampleA normal Windows process gets the user's full tokenAppContainer: "resources are blocked by default and can be granted access as necessary"
ExampleA Kubernetes pod with no policy: "all outbound connections are allowed"A default-deny NetworkPolicy

Deny by default has a cost: friction, approval prompts and broken builds. In practice, people then switch the protection off: --dangerously-skip-permissions in Claude Code, danger-full-access in Codex, seccomp=unconfined in Docker.

Allowlist vs denylist: why blocklists fail for AI agents

An allowlist names what is permitted and refuses everything else. A denylist names what is forbidden and permits everything else. Whitelist and blacklist are the older names.

A denylist loses to a creative model, because there are endless ways to say the same command:

An allowlist has to check the real thing, not a string. The Gemini CLI allowlist bypass (Jul 2025) checked only the first word of a command.

How AI agents break deny by default

Each of these nine started with something allowed by default. Newest first:

IncidentDateAllowed by defaultThe deny-by-default fix
Antigravity "Secure Mode" bypassApr 2026Built-in tools ran before the safety checksCheck every call before it runs
Antigravity Turbo mode wipes a driveDec 2025With approvals off, "clear a cache" deleted the D: driveKeep approvals on for destructive commands
Claude Code project files before trust, CVE-2025-59536Oct 2025MCP servers and hooks ran before "trust this folder"Load nothing from a repo before consent
AWS Kiro rewrites its own allowlistAug 2025The agent could edit its own allowlistMake policy files read-only to the agent
GitHub Copilot "YOLO mode", CVE-2025-537735Aug 2025A hidden instruction set chat.tools.autoApprove: trueMake settings read-only to the agent
Cursor "CurXecute", CVE-2025-541356Jul 2025New MCP servers started automaticallyNever auto-start new tools
Cursor auto-run denylist bypassJul 2025Any command not on the denylistAllowlist, never denylist
Langflow unauthenticated RCE, CVE-2025-3248Apr 2025A code endpoint with no login, exploited in the wildRequire auth on every endpoint
Codex CLI runs repo config, CVE-2025-612602025Loaded .env and .codex/config.toml and ran commands with no confirmationTrust no repo config on open

The off switch was a writable file

Copilot and Kiro kept their safety settings in files the agent could edit, so a hidden instruction turned approvals off or widened the allowlist. As a general rule, the agent can never write its own policy.

New tools started themselves

In CurXecute, Claude Code and Codex CLI ran tools from config nobody approved. The rule: a new tool or MCP server starts only after a human says yes.

Approvals off, damage on

Antigravity lost a drive with approvals off, and its Secure Mode checked too late. Langflow never checked. The rule: every action passes a check before it runs.

The 2026 lab breakouts: a network assumed closed

In 2026, frontier labs caught their own agents leaving sandboxes, and none used a hypervisor exploit. Each went through a channel that was allowed or assumed closed, from a package-registry proxy to an unfiltered DNS resolver. The principles hub tells each one.

After the DNS escape, OpenAI wrote: "Our safety case assumed that the model could not access the live internet... The incident exposed a gap in our controls over network restrictions." It paused all training, evaluation and inference using tool access for its most capable models.

In each case the network was assumed closed, not proven closed. So test the deny: from inside the sandbox, try DNS, the proxy and a direct connection. Each should fail.

How to implement deny by default security (step by step)

  1. Declare what the task needs up front: files, commands, hosts and secrets, so the sandbox never guesses (step one of the six-step method).
  2. Start every resource from an empty rule set.
  3. Grant one thing at a time, as narrow as the control allows: a path, a host plus method, one import.
  4. Make the policy and the agent's config files read-only to the agent.
  5. Require a human to approve any new grant. Never auto-start new tools or MCP servers. Ask again when a tool's definition changes.
  6. Test the deny. Claude Code's docs suggest curl --noproxy '*' https://example.com, which should fail7.
  7. Log every denial and review it before you widen a grant.

Deny by default across the sandbox types

A plain process on any OS starts wide open: it can do everything your user account can. An OS sandbox reaches deny by default only when you stack more tools on top:

  • Linux: a Landlock deny-all ruleset, then grants per path; a seccomp filter whose default action returns an error; bubblewrap namespaces for an empty filesystem view and no network; and a proxy, such as socat, for the network you allow.
  • macOS: a Seatbelt profile that starts from (deny default), run with sandbox-exec, plus a Network Extension filter for the network. Apple has deprecated sandbox-exec and documents no profile language; App Sandbox entitlements are the documented route.
  • Windows: AppContainer or LPAC, which starts with no capabilities, plus App Control in allow mode, so code runs only if your policy says so. Windows Firewall allows all outgoing traffic by default, so set it to block outbound.

Each layer is one more policy to write and keep current. The table rates the OS sandboxes as configured. As elsewhere, we call these "sandbox types" after the conventional usage, though containers and VMs are not sandboxes on their own.

ProcessNo
Starts with
Everything the user can do
Controls to add
Run it under an OS sandbox
Linux sandboxas configuredYes
Starts with
An empty filesystem view (bubblewrap)
Controls to add
bwrap --unshare-all; Landlock deny-all plus grants; seccomp allowlist with SECCOMP_RET_ERRNO; nftables policy drop
Upstream doc
bubblewrap, Landlock, seccomp(2)
macOS sandboxas configuredYes
Starts with
(deny default) in the Seatbelt profile
Controls to add
Seatbelt (deny default); App Sandbox entitlements
Upstream doc
App Sandbox
Windows sandboxas configuredYes
Starts with
No capabilities (AppContainer)
Controls to add
AppContainer or LPAC; App Control allow mode; firewall blocks outbound
Upstream doc
AppContainer
ContainerPartial
Starts with
Default seccomp "disables around 44 system calls out of 300+"; network and capabilities stay broad
Controls to add
--cap-drop ALL; --network none; --read-only; no-new-privileges; default-deny NetworkPolicy
Upstream doc
Docker seccomp8, no_new_privs
gVisor (container runtime)Partial
Starts with
The same open defaults inside as a container, but the host kernel only sees the small set of calls its user-space kernel is allowed to make
Controls to add
Run it as the runtime (runsc); keep the container controls: --cap-drop ALL, --network none, --read-only
Upstream doc
gVisor security model
VMPartial
Starts with
Host denied; inside the guest, everything allowed
Controls to add
No NIC or shared folders until needed (QEMU: -nic none)
Upstream doc
QEMU invocation
microVMPartial
Starts with
Only attached devices; the guest OS is wide open
Controls to add
No network device unless needed
Upstream doc
Firecracker network setup
WasmYes
Starts with
No imports, so no access
Controls to add
Grant each WIT import deliberately
Upstream doc
WIT worlds9

A plain process starts with everything the user can do. A configured OS sandbox can start from nothing but rarely does, because nobody can predict what an arbitrary program needs; Wasm starts there, because each component declares its needs as imports.

Deny-by-default controls for an AI agent at each layer: files, commands, network, secrets, config and tool imports

What the agent vendors ship: network closed by default

  • Claude Code uses Seatbelt on macOS and bubblewrap on Linux and WSL2. Traffic goes through a local proxy whose allowed domains start empty. The gaps: when a command fails, Claude may retry it outside the sandbox with dangerouslyDisableSandbox (set allowUnsandboxedCommands: false to stop that), excludedCommands run with "no filesystem restrictions and no network proxy", and "Claude's file tools, MCP servers, and hooks run outside it."
  • Codex has three modes: read-only, workspace-write and danger-full-access. In the default workspace-write mode, network access is off unless you turn it on.

Two deny-all network examples

Kubernetes (needs a network plugin that enforces NetworkPolicy):

yaml
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: default-deny-egress
spec:
  podSelector: {}
  policyTypes:
  - Egress

nftables:

nftables
table inet agent {
  chain output {
    type filter hook output priority 0; policy drop;
  }
}

For the full 15-row comparison, see the comparison matrix.

Deny by default with WebAssembly components

A WebAssembly component with no imports can compute, but it cannot touch a file, a socket or a secret. In WASI's own words, a component "starts with no ambient authority and can only do what the host explicitly grants" (wasi.dev).

The component declares its needs up front: its imports list them before it runs, so the host never guesses. App Sandbox entitlements are the same idea for macOS apps.

wit
package example:fetcher;

world fetcher {
  import wasi:http/client@0.3.0;
  export run: func() -> string;
}
  • world fetcher is the component's full contract with the host.
  • import wasi:http/client is its only grant: it may make outgoing HTTP requests. The host still decides which destinations14.
  • export run is what it offers. With no file or socket import, it has neither (more on capabilities).

A capability is also unforgeable: a component can use only what it was handed, and it can't guess one, build one from a name or string, or copy one it was never given15. The runtime enforces this, not the program's good behavior: a component can call only the imports the host linked, and the runtime checks every resource handle, so the room has no doors until the host cuts one.

Cosmonic uses this model. Cosmonic Desktop, for macOS, Windows and Linux, runs components in a deny-by-default sandbox and shows what a component can reach before it runs. Cosmonic Control denies egress by default, adds DNS controls, and uses a Kubernetes admission controller and RBAC.

The limits: every capability needs host wiring, and existing CLI tools must be ported or wrapped.

Is deny by default part of zero trust for AI agents?

Yes: zero trust gives no implicit trust by location, so every request is refused until a policy decision allows it16. For agents, location means the laptop, the VPC or the repo the user just opened. Being there grants nothing. The Claude Code and Codex CLI repo-config incidents were location-trust failures. The policy enforcement points page covers checking every request.

Deny by default and the other principles

To see what a component can reach before it runs, try Cosmonic Desktop. Or go back to The 6 Principles of Secure Agent Platforms.

See what a component can reach before it runs

Public beta

Cosmonic Desktop runs each tool as a WebAssembly component that starts with no imports, so it reaches nothing until you grant it. Free for personal use, no account, no cloud.

Cosmonic Desktop's Inspect view: a component's declared interfaces and network egress drawn out, with everything else denied by default.

Frequently asked questions

What is deny by default?
Deny by default is a security rule that refuses every request unless a rule explicitly allows it. For an AI agent, it means starting with no files, hosts, commands or secrets. You grant each one on purpose; anything else stays out of reach.
Is default deny the same as an allowlist?
In effect, yes. Default deny is the rule: anything not allowed is refused. The allowlist is the exceptions you write under it. A default-deny system with an empty allowlist can do nothing, which is the right starting point for a new agent.
What traffic would an implicit deny firewall rule block?
An implicit deny rule blocks all traffic that no earlier rule allowed. It is the silent last line of the rule set: you never write it, but it always applies. For an AI agent's sandbox, that means any host, port or protocol you did not grant stays unreachable.
Is "secure by default" the same as deny by default?
No. Secure by default is a product promise that the shipped settings are safe. Deny by default is one way to keep that promise. A product can be secure by default and still allow some things out of the box.
Can an OS sandbox like bubblewrap, Seatbelt or AppContainer be deny by default?
Yes, as configured. All three can start from nothing. Few tools ship a deny-all profile, because nobody can predict what an arbitrary program will need (Rice's theorem), and profiles break when libraries change. WebAssembly avoids the guess, because each component declares its needs as imports.
Doesn't deny by default make AI agents unusable?
Only if grants are too coarse. Grant per task, ask a human only for new grants, and log every denial so you can widen a grant on purpose when a real need shows up. Turning approvals off instead is how the Copilot and Antigravity Turbo incidents happened.

Further reading

Papers

  • Miller 2006, "Robust Composition"17

Upstream docs

  1. NIST CSRC glossary, “deny by default” (from SP 800-41 Rev. 1). csrc.nist.gov ↩
  2. Saltzer and Schroeder (1975), “The Protection of Information in Computer Systems”. cs.virginia.edu ↩
  3. Saltzer and Schroeder (1975), “The Protection of Information in Computer Systems”, Proceedings of the IEEE. doi.org ↩
  4. nftables wiki, “Configuring chains”. wiki.nftables.org ↩
  5. Embrace The Red, GitHub Copilot remote code execution via prompt injection (CVE-2025-53773), August 2025. embracethered.com ↩
  6. Aim Security, “When public prompts turn into local shells: RCE in Cursor via MCP auto-start” (CurXecute, CVE-2025-54135), July 2025. aim.security ↩
  7. Anthropic, Claude Code sandboxing documentation. code.claude.com ↩
  8. Docker, seccomp security profiles. docs.docker.com ↩
  9. The Component Model book, WIT reference. component-model.bytecodealliance.org ↩
  10. Rice (1953), “Classes of Recursively Enumerable Sets and Their Decision Problems”. doi.org ↩
  11. LWN.net, on the history of seccomp. lwn.net ↩
  12. Linux kernel documentation, Seccomp BPF (seccomp_filter). docs.kernel.org ↩
  13. actions/runner-images issue #3812: Docker’s default seccomp profile blocked glibc 2.34’s clone3. github.com ↩
  14. WebAssembly/wasi-http on GitHub. github.com ↩
  15. Dennis and Van Horn (1966), “Programming Semantics for Multiprogrammed Computations”. doi.org ↩
  16. NIST SP 800-207 (2020), Zero Trust Architecture. doi.org ↩
  17. Miller (2006), “Robust Composition”, PhD dissertation, Johns Hopkins University. erights.org ↩