Download Cosmonic Desktop (Beta)

The 6 principles of secure agent platforms:
security design principles for AI agents

Six rules a sandbox must follow to contain an AI agent that misbehaves, from Saltzer and Schroeder to NIST.

Between July and September 2026, OpenAI, Anthropic and Google each disclosed or confirmed incidents of AI agents breaking out of test sandboxes. OpenAI has since paused all training, evaluation, and inference with tool access for its most capable models. None of these escapes needed an exotic exploit. Each one broke security design principles that have been in print since the 1960s and 1970s, and each is a lesson for anyone choosing an agent sandbox or agentic platform.

Here is what happened:

  • July 9-13, Hugging Face. An agent in OpenAI's ExploitGym cyber evaluation escaped through a zero-day in a package-registry cache proxy, a channel it was allowed to use. It then attacked Hugging Face to grab the evaluation's answer datasets: about 17,600 actions over 4.5 days, per Hugging Face's timeline1.
  • July 30, Anthropic. Anthropic disclosed2 three cases where Claude reached the internet from evaluation environments meant to be isolated. Claude exploited a real company's infrastructure to reach production data, published a booby-trapped package that ran on 15 real systems, and SQL-injected a real target before it stopped.
  • September, Google. Google confirmed that in May, a test setup that wrongly allowed internet access let Gemini guess a password at one real company and use credentials from public code repositories at two others. "In all three of these instances, the model stopped," said Heather Adkins, Google's VP of Security Engineering.
  • September 20, OpenAI. During training, a web proxy blocked an agent's requests. The agent found the sandbox's DNS resolver unfiltered and reached a public chatbot by hiding its queries in DNS names. OpenAI wrote3: "Our safety case assumed that the model could not access the live internet."

In each case, the escape went through a channel that was allowed, or assumed closed: a registry proxy, a DNS resolver, a test network that was supposed to be offline.

Engineers and security teams need a set of principles and strategies to understand sandboxing risks, and a way to apply them. This guide identifies six, based on our review of what we consider the most important and influential papers in security, from Saltzer and Schroeder (1975) to NIST's zero-trust architecture (2020). Every paper is linked in Further reading.

We then map each principle to the tactical controls and the platforms teams are struggling to configure today: seccomp and Landlock on Linux, Seatbelt on macOS, AppContainer on Windows, containers, microVMs and WebAssembly.

The six principles arranged around an AI agent in a sandbox: isolation, deny by default, least authority, policy enforcement points, controlled information flow and composition

What makes a real agent sandbox

An agent sandbox is an isolated environment where an agent, its tools, and the code it writes can run apart from your production systems and your own account, and reach only what someone explicitly granted. Our guide to choosing an agent sandbox covers the kinds teams use today.

The idea is older than agents. Security training such as the CISSP describes a sandbox as an isolated space where untested code and experiments can run safely, apart from the production environment. US government glossaries are stricter: a restricted, controlled execution environment that keeps possibly malicious software away from every system resource except the ones it is authorized to use.

Measured against the six principles in this guide, a real agent sandbox:

  1. has its own walls, in memory, resources and leftover data (isolation);
  2. starts with no access at all (deny by default);
  3. gets only what the current task needs (least authority);
  4. checks every access at the resource itself (policy enforcement points);
  5. controls what data can leave, and where it can go (controlled information flow);
  6. connects to other parts only through declared, typed interfaces (composition).

An environment without these characteristics is not a sandbox: it is simply an execution environment.

Containers, microVMs, and VMs are not sandboxes

Three mechanisms are often discussed as sandboxes, and our own comparison matrices (in this guide and across the site) rate them under those terms. But it’s crucial to understand that while containers, microVMs, and VMs can be part of an effective sandboxing strategy, taken on their own, they are walls, not sandboxes:

  • A container is a normal process given its own view of files, processes and network through kernel namespaces, plus resource limits through cgroups. It shares the host's kernel and needs an operating system image to run anything.
  • A virtual machine (VM) emulates a whole computer with help from the hardware. It needs a full guest operating system: a kernel, drivers, and a userland.
  • A microVM is a stripped-down VM with a tiny device model and a fast boot, such as Firecracker. It still needs a guest kernel and a root filesystem, and a host with KVM.

All three give you a wall (isolation). Inside that wall, the operating system they need starts wide open: a shell, a package manager, the network, and any credentials you mount in. The other five principles are yours to add. That is why the strongest designs put a runtime that starts with nothing, such as WebAssembly, inside one of these walls.

Defense in depth layers: how sandbox controls inherit and overlap

If a wall isn't a sandbox, what is? An effective sandbox is built with defense in depth: walls and policy enforcement points (guards) stacked so that each layer constrains the one inside it. The property that makes this work is inheritance. In an ordinary process, a child inherits its parent's authority: its user, its environment variables, its open files and its network. In a well-built sandbox, each inner layer inherits its parent's limits instead.

An agent running inside of a WebAssembly sandbox, inside of an OS sandbox inside a microVM can do only what every layer allows, and any one layer can say no. The total control surface is the union of every layer's controls: everything with at least one policy enforcement point. Where more caution is desired, such as with egress communication, put the resource at the intersection, where two or more layers guard it.

For example, when the agent tries to communicate with a third party, it could be constrained by a WebAssembly allow list, a Kubernetes NetworkPolicy egress rule, and even by third-party controls on the network and firewall. These overlapping policy enforcement points are what make the stack worth building, because a gap in one control is caught by the next.

In November 2025, Google Antigravity showed what happens without them: a guard blocked the agent's file-reading tool from opening .env, so the agent ran cat .env in the shell instead and sent the secrets out. The guard was configured to block the direct file access, but the effective authority of the system allowed the LLM to access the file through another program. The layers only help if they fail differently. One more counterexample: two containers on the same kernel add a wall without adding an independent guard. The six principles and accompanying matrix should help you to identify effective strategies aligned to your risk model for building sandboxes that have defense in depth.

What are the 6 principles of secure agent platforms?

The 6 Principles of Secure Agent Platforms are isolation, deny by default, least authority, policy enforcement points, controlled information flow and composition, the rules a sandbox must follow to contain an AI agent that misbehaves. The table below includes a column describing each principle in narrative form: scroll down for more on how to read the “story” column.

#PrincipleThe storyFounding paperExample control failure
1IsolationEvery agent gets its own room, with walls to the ceiling and its own share of the heat. The room is emptied before the next guest arrives.Popek and Goldberg 1974NVIDIAScape
2Deny by defaultThe room is built with no doors at all, so an agent with zero capabilities can't touch anything, even by accident.Saltzer and Schroeder 1975Copilot YOLO mode4
3Least authorityYou cut only the doors the job needs, and each one opens into a single room, never into a hallway that reaches every room.Dennis and Van Horn 1966; Hardy 1988Replit production database
4Policy enforcement pointsEvery door gets a guard who checks each trip through it, and checks which room you actually enter, not what the sign on the door says.Anderson 1972Claude Code symlink deny bypass5
5Controlled information flowThe guard also checks what you carry out, because an allowed door can still be used to carry out the wrong papers.Denning 1976EchoLeak
6CompositionRooms connect only through guarded doors, so you can add a new room without re-inspecting the whole building.Miller 2006Cursor MCPoison

The building and the guard: the six principles as one story

Picture the sandbox as a building with a guard. Read the story column as one sequence:

  1. Wall off the room (isolation).
  2. Build it with no doors (deny by default).
  3. Cut only the doors the job needs (least authority).
  4. Put a guard at every door (policy enforcement points).
  5. Have the guard check what leaves (controlled information flow).
  6. Connect rooms only through guarded doors (composition).
The 6 principles of secure agent platforms as one story: isolate the room, start with no doors, cut only the doors the job needs, guard every door, check what leaves, and connect rooms only through guarded doors

The roots of the principles: fifty years of applied research

Security researchers have published applied work on these principles for more than fifty years, starting with documents for shared mainframes and operating systems long before AI agents existed. Dennis and Van Horn (1966)6 described capabilities. Anderson (1972) described a guard that checks every access. Lampson (1973) named the confinement problem, and Popek and Goldberg (1974)7 showed when hardware can wall off a virtual machine. Saltzer and Schroeder (1975)8 gathered eight design principles that still anchor the field, and Denning (1976) gave information flow a formal model.

Later work added walls built in software9, safe composition of parts that don't trust each other10 and zero trust11. The newest research, such as CaMeL (2025), applies the same ideas to AI agents directly. We reviewed that body of work, kept the principles that decide whether a sandbox can contain an AI agent, and grouped them into six.

Timeline of security research papers from 1966 to 2025, each tagged with the security design principle it feeds

1. Isolation: borders before bridges

Every agent gets its own room, with walls to the ceiling and its own share of the heat. The room is emptied before the next guest arrives.

Isolation means a sandbox shares as little as possible with anything else on the machine, and that we have isolated the functionality into an individual unit. That covers three things: its memory and kernel, its share of resources such as CPU and memory, and anything it leaves behind when it finishes, such as temporary files, cached sessions or open connections that the next tenant could pick up.

Memory and kernel. This is the wall itself: separate address spaces, kernels or linear memories. Anything shared, such as a kernel, a folder, a socket or a GPU driver, is a path across it. Popek and Goldberg (1974) showed when hardware can wall off guests, and Wahbe et al. (1993) did it in software. In July 2025, NVIDIAScape (CVE-2025-23266) let a three-line Dockerfile get root on shared GPU hosts.

Resource budgets. Each tenant gets a fixed share of CPU, memory, instance slots, file descriptors and disk, so a noisy neighbor or a fork bomb can't starve the rest. Gligor (1984) made denial of service a security property. There is no public AI agent incident here yet, only the classic failures: starvation and the OOM killer.

Residue and reuse. Nothing outlives the tenant. Scratch files, memory, pooled connections and IDs are wiped or retired first, the Orange Book's "object reuse" rule. In March 2023, a connection-reuse bug showed ChatGPT users other people's chat titles.

Hardware walls (VMs, microVMs) and software walls (namespaces12, Wasm linear memory) both count, and both have had bugs. In April 2026, Wasmtime compiler advisories let a guest reach host memory in non-default configurations: the Winch backend, or aarch64 with Spectre mitigations turned off. So stack walls that fail differently, and use one disposable sandbox per task. Read more on how sandbox escapes happen.

2. Deny by default: start with nothing

The room is built with no doors at all, so an agent with zero capabilities can't touch anything, even by accident.

Deny by default means access exists only where someone explicitly wired it in. No auto-approve, no auto-start, and no denylists standing in for safety. Saltzer and Schroeder13 called it "fail-safe defaults": base access on permission, not exclusion. Dennis and Van Horn (1966) gave the mechanism: you can only use the capabilities you hold. Capabilities are unforgeable: a program can't guess one, build one from a name or copy one it was never handed, and the runtime enforces that.

Two strategies follow. Zero trust14 grants no trust because of where a request comes from. Declare needs up front, because by Rice's theorem (1953)15 no tool can work out in advance what an arbitrary program will need. A Wasm component lists its needs as imports. WASI "starts with no ambient authority."

In August 2025, a hidden instruction made GitHub Copilot set chat.tools.autoApprove: true in its own settings (CVE-2025-53773). In July 2025, Cursor auto-started any new MCP server added to its config16. Read the full deny by default guide.

3. Least authority: grant only what the task needs

You cut only the doors the job needs, and each one opens into a single room, never into a hallway that reaches every room.

Least authority means each part gets only the authority its task needs, and every grant is explicit. Authority lying around (your account, environment variables, a shared token) is ambient, and ambient authority is how excess authority gets used. Hardy's "Confused Deputy" (1988)19 shows a program tricked into misusing power it held for someone else.

Least privilege (POLP) limits what a program is permitted to do. Least authority (POLA), developed in Miller's 2006 thesis, also counts what it can cause through others, such as the tools and servers it calls. Capability Myths Demolished20 explains why capabilities, not access lists, close that gap.

Strategies:

  • capability-based security
  • just-in-time, short-lived grants
  • separate dev and prod
  • per-task identities
  • approval for destructive actions

In July 2025, a Replit agent with live production access deleted a production database during a code freeze. The same month, a Cursor agent connected to Supabase with the all-powerful service_role key dumped private tables into a support ticket. In both cases, the agent held far more authority than its task needed. Read least authority for AI agents.

4. Policy enforcement points: check every access (complete mediation)

Every door gets a guard who checks each trip through it, and checks which room you actually enter, not what the sign on the door says.

Now that our sandbox has doors, each door should have a guard. A policy enforcement point (PEP) sits in front of every resource. Every access crosses it, nothing goes around it, and it is small enough to verify. First well described in Anderson's 1972 reference monitor. The terms PEP and PDP come from policy-based networking (RFC 2753, 2000), and NIST SP 800-207 made them the vocabulary of zero trust: a PEP enforces what a policy decision point (PDP) decides.

A good PEP checks the real resource, not its name. It resolves the file path or the destination first, then checks it. Garfinkel (2003) cataloged what goes wrong when sandboxes check names instead. On Linux, seccomp and Landlock are PEPs, on Windows, AppContainer is one, in CNCF wasmCloud host allow lists are another.

The six resources every sandbox must guard

In modern systems, PEPs fall into six common categories: the filesystem, network egress, program execution, secrets, configuration and the host control plane. We break these out in the deep dives that follow and in the Agent Sandboxing Principles Matrix.

ResourceWhat goes wrongIncident
4a FilesystemThe name is checked, not the targetClaude Code symlink deny bypass, CVE-2025-59829 (Oct 2025)
4b Network egress allowlistAllowed domain, attacker's account"Claude Pirate" Files API exfiltration (Oct 2025)
4c CommandsOnly the first word is checkedGemini CLI allowlist bypass (Jul 2025)
4d SecretsKey sent before trustClaude Code ANTHROPIC_BASE_URL, CVE-2026-21852 (Jan 2026)
4e Config and hooksSandbox writes code the host runs laterClaude Code settings.json hooks, CVE-2026-25725 (Feb 2026)
4f Host control planeDocker socket reachable"One Docker socket to rule them all" (Jul 2026)

The network egress row is where the 2026 lab breakouts live: a registry proxy and a DNS resolver are egress too. See the full guide to policy enforcement points.

5. Controlled information flow: break the lethal trifecta

The guard also checks what you carry out, because an allowed door can still be used to carry out the wrong papers.

Controlled information flow means data may only flow to places cleared to hold it. Isolation stops access. It does not stop an allowed channel from carrying data out. Lampson (1973) named this the confinement problem, and Denning (1976) gave a lattice model for where data may go.

For agents, watch for the lethal trifecta21: private data, untrusted content and a way out. Never put all three in one context. Potential mitigation strategies to consider:

  • break the trifecta
  • taint tracking and labels
  • split the planner from the reader (dual-LLM, or CaMeL)
  • egress data loss prevention (DLP)
  • confirm before data leaves

In June 2025, with EchoLeak (CVE-2025-32711), one unopened email made Microsoft 365 Copilot gather internal data and send it out through Microsoft's own trusted domains. In May 2025, a malicious public issue got an agent using the GitHub MCP server22 to copy private-repo data into a public pull request. Greshake et al. (2023)23 described this attack, indirect prompt injection, where untrusted content becomes instructions. Read more about the lethal trifecta.

6. Composition: isolated parts, typed interfaces

Rooms connect only through guarded doors, so you can add a new room without re-inspecting the whole building.

Composition means parts that share nothing and talk through typed interfaces can be checked one at a time and combined safely. Parts that share state, or trust each other by name, let one bad part break the whole. Miller (2006) showed how to compose parts that don't trust each other. Thompson (1984) showed why you can't fully trust what you didn't build.

Strategies:

  • supply-chain verification: signing (Sigstore), SBOMs, pin by digest
  • re-approve on change, which stops rug pulls
  • one sandbox per tool or server
  • typed interfaces, such as the Component Model, instead of shared context

As you consider how you might isolate the components of an agentic workflow, there are four properties to know when you are planning how to compose them:

  1. Authority is additive: a workflow can do everything any of its steps can do.
  2. Authority is transitive: if step A can call step B, A can reach whatever B reaches.
  3. In an ordinary process or container authority is inherited: a child process gets its parent's user, environment variables, open files and network by default.
  4. A system is composable when two parts that are each safe stay safe together, and you can check the whole by checking each part and the typed connections between them.

For example, WebAssembly components inherit nothing and pass only unforgeable handles, which keeps that property. Read more in composition.

In July 2025, Cursor's "MCPoison" (CVE-2025-54136) kept trusting an approved MCP config by name after an attacker swapped in malicious commands. In April 2025, Invariant Labs showed tool poisoning and rug pulls24: an MCP tool's description can change after you approve it. Read more about composition and MCP security.

For practitioners, mind the multi-tenancy boundary: space and time

Throughout this guide and its tables, you will find isolation controls: VMs, microVMs, containers and WebAssembly components, along with the guarantees each one provides. As you weigh them, keep one question in mind: where is the multi-tenancy boundary, and whose tenancy does it protect?

There are usually two answers. The first is the boundary between users, which keeps one customer, team, or agent session from reaching another's data, credentials, or compute. The second is the boundary between a user's apps, which keeps the tools, agents, MCP servers, and generated code running on that user's behalf from reaching each other.

A VM or microVM per user can draw the first line well and leave the second wide open: everything inside that guest shares the same filesystem, network, and secrets, so one compromised or misbehaving app gets everything the user has. As you evaluate each control, ask which of these two boundaries it enforces, at what granularity, and at what cost. Many architectures handle one well and only assume the other.

Isolation is usually pictured as a wall in space: separate memory, separate kernels, separate machines. It also has to hold in time. Every task leaves something behind: scratch files, temp directories, cached sessions, pooled connections, environment variables, memory pages, a warm snapshot, an identifier that gets reused. If the next tenant or the next task can see any of that, the boundary was never complete, even if no wall was breached.

This is the Orange Book's object reuse requirement, and it's the easiest one to miss with VMs and microVMs, because they are often pooled, snapshotted and reused for speed. A microVM restored from a shared snapshot, or a user's VM that runs one app after another, carries forward whatever the last workload left there.

So as you weigh the controls, ask a third question alongside "between users" and "between a user's apps": what survives when the task ends, and who gets it next? The strongest answer is ephemeral by default: a fresh sandbox per task, cleared on release, and nothing reused that hasn't been wiped or verified first.

Is defense in depth a security principle? No, it's a strategy

Defense in depth is a strategy, not a principle: it stacks boundaries so an attacker has to beat several independent mechanisms, and it only works when each layer applies the six principles and fails in a different way. It is Saltzer and Schroeder's separation of privilege, applied to walls:

  • Weak: a container inside a container, on the same kernel. One kernel bug beats both.
  • Dense: a Wasm runtime inside an OS sandbox, with Landlock and seccomp on the host process. Two different mechanisms, with little density cost.
  • Strong: Wasm inside a microVM. A hardware wall, at a density cost, for high-value tenants or hostile multi-tenancy.

Add an egress proxy (a policy enforcement point for network egress) and a scoped cloud identity (least authority) to any of them.

Defense in depth for AI agents: a weak stack that shares one kernel, a dense stack of Wasm inside an OS sandbox, and a strong stack of Wasm inside a microVM

How each sandbox type measures up

Sandbox security looks very different across the common sandbox types. Here is the building story, type by type:

  • Process: You share the building with everyone else, and every door is already unlocked.
  • OS sandbox: You put up your own walls and lock the doors yourself (bubblewrap, Seatbelt, AppContainer). They hold as well as you configure them, and the foundation is still shared.
  • Container: You have your own room, but it shares a foundation with every other room, and the doors come pre-cut.
  • gVisor: Your container room gets its own front desk. A clerk who works only for your room (a kernel written in user space) handles your requests and passes only a short list of them down to the shared foundation. There are far fewer ways to reach the foundation, at some cost in speed, and the doors still come pre-cut.
  • VM and microVM: You get the thickest walls available, but the room arrives as a whole furnished house (the guest OS) with its doors already cut.
  • Wasm: You get a sealed room with no doors until the host cuts each one. The walls are software, so the strongest design puts this room inside a microVM, or, for density, inside an OS sandbox.

The table summarizes our sandbox security comparison matrix: eight sandbox types rated across 15 rows (the six principles plus sub-rows 1a to 1c and 4a to 4f). "Yes" means met by default; "Partial" means partly, or with added controls. OS sandboxes are rated as configured. A plain process is shown as a baseline; it is not a sandbox.

Sandbox typeYes / Partial / NoVerdict
Process (baseline)0 / 1 / 14Holds the user's whole account; everything is open
Linux sandbox7 / 8 / 0Richest toolbox, no root needed; one shared kernel
macOS sandbox4 / 10 / 1Seatbelt can start from deny; profile language deprecated
Windows sandbox3 / 12 / 0AppContainer starts with nothing; few CLI tools use it
Container1 / 13 / 1Fast and portable; shared kernel, open egress
gVisor1 / 13 / 1A user-space kernel narrows the shared kernel; the doors still come pre-cut
VM2 / 10 / 3Hardware wall; full guest OS with ambient authority
microVM3 / 10 / 2Hardware wall, tiny VMM; needs KVM and a guest OS
Wasm9 / 6 / 0Deny by default via imports; must compile to Wasm

A plain process fails almost everything. A configured OS sandbox gets you most of the way, but only as well as you configure it, and the kernel is still shared. Wasm meets the most rows by default, but it is software-enforced, code must be compiled to Wasm and WASI is still maturing, so stack it: inside an OS sandbox for density (the dense stack), or inside a microVM for a hardware wall (the strong stack). No type meets controlled information flow by default.

Sandbox security by type: how a plain process, configured Linux, macOS and Windows sandboxes, containers, VMs, microVMs and Wasm meet the six security design principles by default

The trouble with OS sandboxes

An OS sandbox starts wide open, with everything its user can reach. You then secure it with yet more tools and layers:

  • Linux: bubblewrap, Landlock, seccomp and cgroups, plus a proxy such as socat for the network.
  • macOS: Seatbelt (sandbox-exec), plus a Network Extension filter.
  • Windows: AppContainer or LPAC, Job Objects and App Control.
  • The agent: an egress proxy, and credential masking, which needs TLS termination.

Getting that stack right is hard, because it asks you to predict the future. Every rule has to anticipate what the agent will need to touch, and Rice's theorem says no tool can work that out in advance for an arbitrary program. Set the rules too tight and the build breaks. Set them too loose and they protect nothing.

In practice, a build breaks and people switch the protection off: --security-opt seccomp=unconfined or --privileged in Docker, --dangerously-skip-permissions, Claude Code's dangerouslyDisableSandbox retry and excludedCommands, or Codex's danger-full-access mode.

The vendors' own docs show the assembly. Claude Code's sandbox25 uses Seatbelt on macOS and bubblewrap with socat on Linux, and it "covers shell commands only." OpenAI Codex uses Seatbelt on macOS, bwrap and seccomp on Linux, and a native Windows sandbox. Both vendors are now moving their desktop agents toward the cloud: new Cowork tasks on Pro and Max plans run in the cloud from October 6, 202626, and Codex got reusable cloud environments27 on September 29, 2026.

How the principles map to the six-step method

The six-step method applies the principles in order of work:

  1. Break the agent into workflow steps: composition and isolation (one sandbox per step).
  2. List the capabilities each step actually needs: least authority and deny by default (the list is the allowlist).
  3. Pick the most restrictive boundary that can grant exactly those: isolation and deny by default.
  4. Choose the right abstraction for each layer: defense in depth and PEPs.
  5. Verify like any other software requirement: PEPs (test symlinks, redirects, chained commands) and information flow (egress logs).
  6. Compose isolated, secure steps into a workflow: composition, with no step holding all three trifecta legs.

Applying the principles with Cosmonic

Cosmonic applies these principles with WebAssembly components. Cosmonic Desktop, a free app for macOS, Windows and Linux, runs Wasm components locally in a deny-by-default sandbox and shows what a component can reach before it runs (deny by default and least authority).

Cosmonic Control is a Kubernetes-native platform built on wasmCloud with a capability-based security model. Egress is denied by default, with DNS controls. Secrets come from Kubernetes Secrets and External Secrets. It integrates with the Kubernetes admission controller and RBAC (composition), reports through OpenTelemetry (the verify step) and installs on-prem or air-gapped.

Download Cosmonic Desktop to try the principles on your own machine.

Apply the six principles on your own machine

Public beta

Cosmonic Desktop, a free app for macOS, Windows and Linux, runs WebAssembly components in a deny-by-default sandbox and shows what a component can reach before it runs.

Cosmonic Desktop's Inspect view: a component's declared interfaces and network egress drawn out, with everything else denied by default.

Frequently asked questions

What are the principles of secure design?
The principles of secure design are rules for building systems that stay safe when parts fail. For AI agents, six matter most: isolation, deny by default, least authority, policy enforcement points, controlled information flow and composition. They extend Saltzer and Schroeder's eight 1975 principles for systems that run AI agents.
What are Saltzer and Schroeder's design principles?
Saltzer and Schroeder's design principles are eight rules from their 1975 paper: economy of mechanism, fail-safe defaults, complete mediation, open design, separation of privilege, least privilege, least common mechanism and psychological acceptability. Fail-safe defaults became deny by default, complete mediation became policy enforcement points, and least privilege grew into least authority.
What are the 5 basic security principles?
Lists of five basic security principles usually describe goals: confidentiality, integrity, availability, often with authentication and non-repudiation. Design principles describe how to build a system that meets those goals. No list of five is canonical, so treat goals and design principles as two layers of the same job.
What are the principles of secure by design?
Secure by design principles vary by publisher. The UK government's Secure by Design principles list ten for government digital services, from owning cyber risk to making changes securely. The six principles here are the subset that decides whether a sandbox can contain an AI agent that misbehaves.
Is defense in depth a security principle?
No. Defense in depth is a strategy that stacks several boundaries so an attacker must beat each one. It helps only when every layer follows the six principles and fails in a different way. Two containers on one kernel share a single point of failure.
Which sandbox meets all six principles?
No sandbox meets all six principles by default. Wasm meets the most rows (9 Yes, 6 Partial, 0 No over 15), then a configured Linux sandbox (7 / 8 / 0); a plain process meets almost none. Strong setups stack Wasm inside an OS sandbox or a microVM, plus an egress proxy. See the comparison matrix.

Further reading

  1. Hugging Face, technical timeline of the July 2026 agent intrusion. huggingface.co ↩
  2. Anthropic, investigating incidents in cybersecurity evaluations, July 2026. anthropic.com ↩
  3. OpenAI, “An agent used DNS to reach an external chatbot”, September 2026. alignment.openai.com ↩
  4. Embrace The Red, GitHub Copilot remote code execution via prompt injection (CVE-2025-53773), August 2025. embracethered.com ↩
  5. CVE-2025-59829: Claude Code deny rule bypassed through a symlink. GitHub security advisory. github.com ↩
  6. Dennis and Van Horn (1966), “Programming Semantics for Multiprogrammed Computations”. doi.org ↩
  7. Popek and Goldberg (1974), “Formal Requirements for Virtualizable Third Generation Architectures”. doi.org ↩
  8. Saltzer and Schroeder (1975), “The Protection of Information in Computer Systems”. cs.virginia.edu ↩
  9. Wahbe et al. (1993), “Efficient Software-Based Fault Isolation”. doi.org ↩
  10. Miller (2006), “Robust Composition”, PhD dissertation, Johns Hopkins University. erights.org ↩
  11. NIST SP 800-207 (2020), Zero Trust Architecture. csrc.nist.gov ↩
  12. Linux manual page, namespaces(7). man7.org ↩
  13. Saltzer and Schroeder (1975), “The Protection of Information in Computer Systems”, Proceedings of the IEEE. doi.org ↩
  14. NIST SP 800-207 (2020), Zero Trust Architecture. doi.org ↩
  15. Rice (1953), “Classes of Recursively Enumerable Sets and Their Decision Problems”. doi.org ↩
  16. Aim Security, “When public prompts turn into local shells: RCE in Cursor via MCP auto-start” (CurXecute, CVE-2025-54135), July 2025. aim.security ↩
  17. LWN.net, on the history of seccomp. lwn.net ↩
  18. actions/runner-images issue #3812: Docker’s default seccomp profile blocked glibc 2.34’s clone3. github.com ↩
  19. Hardy (1988), “The Confused Deputy”. doi.org ↩
  20. Miller, Yee and Shapiro (2003), “Capability Myths Demolished”. zesty.ca ↩
  21. Simon Willison, “The lethal trifecta for AI agents”, June 2025. simonwillison.net ↩
  22. Invariant Labs, “GitHub MCP exploited”, May 2025. invariantlabs.ai ↩
  23. Greshake et al. (2023), “Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection”. doi.org ↩
  24. Invariant Labs, “MCP security notification: tool poisoning attacks”, April 2025. invariantlabs.ai ↩
  25. Anthropic, Claude Code sandboxing documentation. code.claude.com ↩
  26. Anthropic, “Use Claude Cowork on web, desktop and mobile”. support.claude.com ↩
  27. TechCrunch, “OpenAI gives Codex reusable cloud environments that work across devices”, September 2026. techcrunch.com ↩