Secure composition for AI agents:
safe parts, safely connected
Build agents out of parts that don't trust each other, so the whole stays as safe as its parts.
Containers, microVMs and OS sandboxes share a major flaw: they don't describe their inputs or their outputs. A container image says what is inside it, not what it expects to receive or what it will emit. At most it lists some ports, environment variables and folders1.
Complexity doesn't go away. It is only relocated, so the burden falls on the sender and the receiver: every pair of parts invents its own way to talk. In an AI agent, those improvised joints are where attacks such as MCP tool poisoning get in.
This page covers composition, one of The 6 Principles of Secure Agent Platforms. It is how you put the parts of your agent sandbox back together safely. Untyped I/O costs you five ways:
- Ad hoc formats: JSON over stdout, environment variables, mounted files, sockets, shared volumes.
- Duplicate code: parsing and validation on both sides of every joint.
- Version drift: one side renames a field, and the other breaks quietly or runs on bad data.
- Attack surface: wide, undeclared I/O is where attacks land. Tool poisoning is one symptom.
- No checks up front: nobody can check a composition before it runs.
Secure composition replaces those joints with declared, typed ones. In the series' building analogy, rooms connect only through guarded doors, so you can add a new room without re-inspecting the whole building.

How composition puts the pieces back together
Composition comes last because the other five principles all push you to split an agent into smaller parts. Each split makes the agent safer and adds a joint to secure.
- Isolation puts each part in its own box.
- Deny by default means each new box starts with nothing.
- Least authority sizes each grant to one job, so one broad token becomes several narrow ones.
- Policy enforcement points put a checkpoint on every path from a part to the outside world.
- Controlled information flow keeps the part that reads private data away from the part that can send data out.
Follow all five and one coding agent becomes many parts: the harness, a fetcher, a code runner, two MCP servers, a skill, and the apps it builds.
Join them through a shared folder and one shared token, and you have rebuilt the single big box you started with. Join them through narrow, typed, declared interfaces, and each part stays as small as you made it. You've put the pieces back together: securely, this time.
What is secure composition?
Secure composition is the practice of combining parts that don't trust each other so the whole stays as safe as its parts. Each part declares what it takes in (inputs), returns (outputs), and what it may reach. Nothing else crosses the boundary.
Take an agent that fetches a web page, summarizes it and posts the summary to Slack. Composed securely, the summarizer has no network access, the fetcher never sees your Slack token, and the poster can reach one host. A poisoned page can spoil a summary, but the summarizer has no way to send your files anywhere.
Three research ideas sit under this:
- Reynolds 2002, separation logic: parts that share no state can be checked one at a time.
- Miller 2006, Robust Composition2: you can build reliable systems from parts that don't trust each other, if each holds only the authority it was handed.
- Cardelli and Gordon 1998, mobile ambients: crossing a boundary between domains takes a capability.
The strategies, each covered below: typed interfaces, one sandbox per tool, verify before you compose, and re-approve on change.
Controls by layer
| Layer | Control | Where it applies |
|---|---|---|
| Build | Sigstore cosign signing, SBOMs, pin by hash | Any OS |
| OS | One bubblewrap sandbox per tool; XPC services; one AppContainer per tool; service mesh mTLS and admission policy | Linux; macOS; Windows; Kubernetes |
| Runtime | Compose Wasm components through WIT3 and WAC, sharing nothing | Wasm |
| Agent | Namespace tool names; re-consent when tool descriptions change | Any OS |
The Unix philosophy: small parts, one interface
The Unix philosophy is a way to build software from small programs that each do one job and connect through a shared interface. The 1978 foreword to the Unix issue of the Bell System Technical Journal put it this way: "Make each program do one thing well" and "Expect the output of every program to become the input to another, as yet unknown, program"4.
The pipe, written |, connects one program's output (stdout) to the next program's input (stdin):
cat access.log | grep " 404 " | sort | uniq -cFour programs, none written with the others in mind, count the "not found" errors in a web log.
Why it worked
The parts were small, so each was easy to test and replace. The interface was uniform: every program read and wrote a stream of bytes, so any program could plug into any other. Doug McIlroy had asked for this in a 1964 memo, wanting ways of "connecting programs like garden hose"5.
Where it falls short for agents
- The interface is untyped bytes. sort doesn't know it is sorting web requests. If the log format adds a column, grep still runs and quietly returns the wrong answer. Every program parses text in and formats text out, the same burden JSON over stdout puts on containers.
- Every program has ambient authority. Ambient authority is the power a program gets just by running as you: your home folder, your SSH keys, the whole network. grep only needs a stream of text, but it could read
~/.sshor open a connection, and nothing in the pipe would stop it.
That was fine for trusted tools that shipped with the OS, but not for an agent that adds a marketplace tool every week. Containers, microVMs and OS sandboxes narrow the second gap. The first gap stays.
WebAssembly components keep what worked and close both gaps. Think of them as Unix pipes with types and without ambient authority.

Typed interfaces, not untyped bytes: the WebAssembly Component Model
The WebAssembly Component Model builds software from components that share no memory and talk only through typed interfaces. A component "is restricted to interact only through the modules' imported and exported functions"6. Joining components this way is called shared-nothing linking: they agree on types, not on a shared memory.
The interfaces are written in WIT (WebAssembly Interface Types). In WIT, a world "describes a set of imports and exports" (WIT reference). Imports are everything a component can reach. Exports are everything it offers. Here are the summarizer and fetcher from the example above:
package example:digest@0.1.0;
interface summarize {
record summary {
title: string,
bullets: list<string>,
}
summarize: func(page: string) -> result<summary, string>;
}
interface fetch {
fetch: func(url: string) -> result<string, string>;
}
world summarizer {
export summarize;
}
world fetcher {
import wasi:http/client@0.3.0;
export fetch;
}Line by line: summarize takes a page as a string and returns a summary record, or an error. The summarizer world exports that function and imports nothing: no network, no files, no clock. The fetcher world imports outbound HTTP and nothing else. WAC, "a tool for composing WebAssembly components together," wires one component's exports into another's imports.
Back to the five costs from the top of this page:
- Formats: one WIT contract replaces ad hoc formats.
- Parsing: generated bindings deliver the summary record as a native type.
- Drift: packages carry versions, and a type mismatch shows up when you compose the parts.
- Attack surface: a component reaches only what it imports.
- Checks: you can read every component's world before anything runs.
The parts don't have to share a language. Per the docs, "A component implemented in Go can communicate directly and safely with a C or Rust component." Your fetcher can be Rust and your summarizer Python, joined by types, not by text each side must parse.
A poisoned summarizer can still write a misleading summary, but it can't read your files or reach the network: its world imports neither. Firefox applies the same idea with RLBox (Narayan et al. 2020).
Unforgeable handles make it safe to pass authority between composed parts: a component can use only what it was handed, and it can't guess a capability, build one from a name or string, or copy one it was never given7. The runtime enforces this, not the component's good behavior: a component can call only the imports the host linked, and the runtime checks every resource handle, so handing one to the summarizer gives it that one thing and nothing more.
The limits, plainly. A component given an import can still misuse it. Code must compile to Wasm, so arbitrary Linux binaries won't run and the harness often stays in a VM or container. WASI is still maturing, and isolation is enforced in software, so the compiler is in the trusted base.
Cosmonic Control runs MCP servers as sandboxed components with a capability-based security model. Cosmonic Desktop, a free app for macOS, Windows and Linux, lets coding agents build tools as components and shows what each can reach before it runs. See capabilities for the model behind it.

How authority combines when you compose: additive, transitive, inherited, composable
Typed calls have a side benefit: security teams can audit an agentic workflow from its declarations. When you join parts into one workflow, their authority combines in four ways, and each property decides whether the whole stays as safe as its parts.
- Additive. When you put parts together, their authority adds up: a workflow can do everything any of its steps can do. With declared imports you can total it up before anything runs, because the composed component's imports are the union of what its parts still need. With containers or scripts, you can't see the sum.
- Transitive. Authority flows along calls. If step A can call step B, and B can reach the database, then A can reach the database through B. This is how a confused deputy happens8. Keep each interface narrow: in Wasm, A reaches B only through B's typed exports, so what flows through B is only what B's interface allows.
- Inherited. In an ordinary OS process or container, a child process inherits the parent's whole environment by default: its user, its environment variables (often API keys), its open files and its network. Authority leaks down the tree without anyone deciding it. In the Wasm Component Model nothing is inherited. Each component gets only the imports it is explicitly linked with, and a capability passed to one component is not visible to its siblings.
- Composable. This is the goal: when two parts are each safe on their own, the combination is still safe, and you can check the whole by checking each part and the typed connections between them. Shared state and inherited authority break this. Typed, shared-nothing interfaces and unforgeable handles preserve it (Miller, 2006, "Robust Composition").
The earlier principles split the work into small parts with little authority each, and these four properties decide whether putting them back together keeps those guarantees.
Composition across the sandbox types
Give every tool, MCP server and skill its own sandbox, so a bad one can only damage itself: an MCP sandbox for every server, not one box around the agent. Agent sandboxes often stop at the shell: "The sandbox covers shell commands only. Claude's file tools, MCP servers, and hooks run outside it"9.
We rated the eight sandbox types in our comparison matrix against composition, with a plain process as a baseline. As elsewhere, we call these "sandbox types" after the conventional usage, though containers and VMs are not sandboxes on their own. Yes means it meets the principle by default; Partial means partly, or with the listed controls; No means not by default. OS sandboxes are rated as configured.
ProcessNo
- Pro
- Processes can talk over IPC
- Con
- Plugins and libraries share a process and the user's authority
- Controls to add
- Move to OS sandboxes, one per tool
Linux sandboxPartial
- Pro
- One sandbox per tool is cheap
- Con
- Interfaces between sandboxes are untyped (sockets, files)
- Controls to add
- One bubblewrap sandbox per tool
macOS sandboxPartial
- Pro
- XPC services split privileges
- Con
- Per-tool profiles are manual
- Controls to add
- XPC services; per-tool Seatbelt profiles (deprecated, no Apple docs)
Windows sandboxPartial
- Pro
- Separate AppContainers per tool; brokered COM
- Con
- Interfaces are untyped
- Controls to add
- One AppContainer per tool
ContainerPartial
- Pro
- One service per container; mesh mTLS
- Con
- Trust between services is by network address
- Controls to add
- Istio mTLS; signed images plus admission policy
gVisorPartial
- Pro
- One sandbox per tool server is practical
- Con
- Interfaces between sandboxes are untyped (network)
- Controls to add
- One runsc sandbox per tool server (gVisor)
VMPartial
- Pro
- Strong separation between trust levels
- Con
- Machine granularity; heavy
- Controls to add
- Separate VMs per trust level, as in Qubes OS
microVMPartial
- Pro
- Cheap per-tool VMs
- Con
- Interfaces are untyped (network, vsock)
- Controls to add
- One Firecracker microVM per tool server
WasmYes
- Pro
- Typed WIT interfaces, shared nothing
- Con
- A component given an import can still misuse it
- Controls to add
- Sign components; pin by digest; grant only the imports each needs
A plain process runs every tool with your whole account, so one bad tool is everyone's problem. OS sandboxes start wide open, and you add layers to close them: bubblewrap, Landlock and seccomp on Linux, Seatbelt on macOS, AppContainer on Windows. They, containers, gVisor and VMs can give each tool its own box, but the joints between boxes stay untyped. Wasm starts with typed, shared-nothing joints. For a side-by-side comparison, see containers vs microVMs vs Wasm.
MCP tool poisoning: a symptom of untyped, shared context
MCP tool poisoning is an attack where a malicious MCP server hides instructions in a tool's description, which the model reads and the user usually never sees10. It is a composition failure.
MCP does type part of each tool: an inputSchema, and optionally an outputSchema11. But the description is free text, and every server's descriptions land in one shared model context. The model becomes the joint between tools, and it reads everything as instructions it might follow. The spec says clients "MUST consider tool annotations to be untrusted unless they come from trusted servers."
Two variants share that gap:
- Tool shadowing: one server's description changes how the agent uses another server's tool, as with the WhatsApp MCP "sleeper" server.
- Rug pull: a tool changes after you approved it. In Cursor MCPoison (CVE-2025-54136), an approved config was swapped for malicious commands and still trusted by name.
With typed composition, a component's imports, not its description, decide what it can reach. For the full list of threats and fixes, see MCP security.
Running MCP servers as Wasm components
You can build an MCP server as a WebAssembly component instead of a native process or a container. The server's code runs in a Wasm runtime such as Wasmtime, and it can reach only the interfaces it imports: no file system, network or environment variables unless the host grants them. Two tools do this today:
- Wassette (Microsoft, 2025) is a runtime that loads components and exposes their typed WIT interfaces as MCP tools. Its components "can't access system resources without explicit access permissions."
- Cosmonic Desktop is a free app for macOS, Windows and Linux that runs components in a local deny-by-default sandbox. Its integrated MCP server lets coding agents such as Claude Code, Codex and Gemini CLI build tools as components, deploy them into that sandbox, and show what each one can reach before it runs.
What that means in practice:
- The server's reach is written down. Its imports are its full list of permissions, and you can read them before it runs. A server that formats dates has no network import, so it can't send anything anywhere.
- Servers share nothing. Each server is its own component with its own memory and its own grants, so one server can't read another's files or credentials.
- A rug pull has to ask. A new version that wants a new host or a new folder needs a new grant, which the host can refuse or flag for re-approval. Changes inside what a server already holds still get through, such as the postmark-mcp backdoor (a malicious MCP server, in the table below) adding a BCC to mail it was already allowed to send. That is why pinning and re-approval stay on the list.
Wasm does not fix the description problem on its own. A Wasm server's tool descriptions still land in the shared model context, so a poisoned description can still steer the model into misusing other tools. Wasm limits what a server can do; namespacing, showing full descriptions and re-approval handle what it says. In production, Cosmonic Control runs MCP servers as sandboxed components on Kubernetes.
The evidence: eight composition failures
Each case broke at a joint between parts. Newest first:
| Date | Case | What was shared or trusted by name | What happened |
|---|---|---|---|
| Feb 2026 | SANDWORM_MODE npm worm | Agent config files any package could write | Typo-squatted packages added a fake MCP server to Claude Desktop, Cursor and Windsurf configs that told the AI to read SSH keys |
| Feb 2026 | ClawHavoc, OpenClaw skills | A skills marketplace | Hundreds of fake skills tricked users into installing infostealers |
| Jan 2026 | Git MCP plus Filesystem MCP chain (CVE-2025-68143/4/5) | Two servers in one agent | Each Git MCP flaw was limited alone; combined with the Filesystem MCP server they reached code execution |
| Sep 2025 | postmark-mcp backdoor | A package trusted by name across versions | Normal for 15 versions, then one line BCC'd every email to the attacker |
| Jul 2025 | Cursor MCPoison (CVE-2025-54136) | An approved MCP config, trusted by name | The attacker swapped in malicious commands after approval |
| Jul 2025 | mcp-remote (CVE-2025-6514) | A popular proxy between client and remote server | An untrusted remote MCP server could run commands on the user's machine |
| Apr 2025 | WhatsApp MCP exfiltration | One agent context for two servers | A sleeper server made the agent forward chat history to an attacker |
| Apr 2025 | MCP tool poisoning, shadowing and rug pulls | Tool descriptions in the shared context | Hidden instructions, changes after approval, redirected trusted tools |
Across these vulnerabilities we see two root causes repeat. The first is shared state: tool descriptions share the model's context, plugins share a process, and servers share a user's credentials. The second is trust by name: a part approved once is trusted forever, as with MCPoison and postmark-mcp. The Git plus Filesystem chain shows harm adding up across parts, which is why defense in depth matters.
AI agent supply chain security: verify what you compose
Before you plug a tool, MCP server or skill into an agent, prove who built it, pin the exact version, and re-approve it when it changes. Ken Thompson's "Reflections on Trusting Trust" (1984) showed you can't fully trust code you didn't build. In the MCP supply chain, SANDWORM_MODE and postmark-mcp both arrived as ordinary packages.
- Sign and verify. Sigstore cosign ties a signature to an identity, keyless by default.
- Record provenance for each build step with in-toto12.
- Keep an SBOM, a list of every component in a build.
- Pin by digest, not tag13. Pinned images don't pick up patches, so schedule updates.
- Re-approve on any change to code, config or tool description.
- Enforce at deploy time: reject unsigned or unpinned workloads with Kubernetes ValidatingAdmissionPolicy (GA in v1.30).
How to compose an agent safely (and prevent MCP tool poisoning)
- Define each part's inputs and outputs as a typed interface, not JSON over stdout or a shared folder.
- Give every tool, MCP server and skill its own sandbox.
- Grant each part only the imports it needs. A part with no network import can't leak.
- Give each server its own scoped credential, never a shared all-repo token (least authority).
- Keep the part that reads private data apart from the part that can send data out (controlled information flow).
- Show users full tool descriptions, and namespace tool names so one server can't shadow another.
- Pin every part by digest, and re-approve on any change.
- Never auto-start new servers (deny by default).
- Log every tool call, so you can trace a bad result to the part that produced it.
Composition and defense in depth
Composition joins parts side by side. Defense in depth stacks boxes inside each other; the principles hub compares weak, dense and strong stacks.
To apply it, decompose the agent, sandbox each part, then compose them back, as in the six-step method.
Running tools and MCP servers in production? Cosmonic Control composes them as sandboxed Wasm components on Kubernetes, with egress denied by default and OpenTelemetry observability. To start on a laptop, download Cosmonic Desktop.
Related topics
Run every MCP server in its own sandbox
Public betaCosmonic Control runs MCP servers sandboxed, with a capability-based security model. Build and run a tool as a component locally with Cosmonic Desktop.

Frequently asked questions
What is secure composition?
What is the WebAssembly Component Model?
What is typed I/O, and why does it matter for AI agents?
Can tools written in different languages be composed safely?
How are Wasm components different from Unix pipes?
What is MCP tool poisoning?
Further reading
- Reynolds, "Separation Logic" (2002), and Cardelli and Gordon, "Mobile Ambients" (1998)
- Narayan et al., "Retrofitting Fine Grain Isolation in the Firefox Renderer" (RLBox, 2020)
- The Component Model docs and WAC
- OCI image specification, image configuration. github.com ↩
- Miller (2006), “Robust Composition”, PhD dissertation, Johns Hopkins University. erights.org ↩
- The Component Model book, WIT reference. component-model.bytecodealliance.org ↩
- McIlroy, Pinson and Tague (1978), “UNIX Time-Sharing System: Foreword”, Bell System Technical Journal. archive.org ↩
- McIlroy (1964), the memo proposing pipes. Bell Labs. nokia.com ↩
- Bytecode Alliance, WebAssembly Component Model: why the component model. component-model.bytecodealliance.org ↩
- Dennis and Van Horn (1966), “Programming Semantics for Multiprogrammed Computations”. doi.org ↩
- Hardy (1988), “The Confused Deputy”. doi.org ↩
- Anthropic, Claude Code sandboxing documentation. code.claude.com ↩
- Invariant Labs, “MCP security notification: tool poisoning attacks”, April 2025. invariantlabs.ai ↩
- Model Context Protocol specification (2025-06-18), Tools. modelcontextprotocol.io ↩
- Torres-Arias et al. (2019), “in-toto: Providing farm-to-table guarantees for bits and bytes”, USENIX Security. usenix.org ↩
- Docker CLI reference, docker image pull (pull by digest). docs.docker.com ↩