Download Cosmonic Desktop (Beta)

Secure composition for AI agents:
safe parts, safely connected

Build agents out of parts that don't trust each other, so the whole stays as safe as its parts.

Containers, microVMs and OS sandboxes share a major flaw: they don't describe their inputs or their outputs. A container image says what is inside it, not what it expects to receive or what it will emit. At most it lists some ports, environment variables and folders1.

Complexity doesn't go away. It is only relocated, so the burden falls on the sender and the receiver: every pair of parts invents its own way to talk. In an AI agent, those improvised joints are where attacks such as MCP tool poisoning get in.

This page covers composition, one of The 6 Principles of Secure Agent Platforms. It is how you put the parts of your agent sandbox back together safely. Untyped I/O costs you five ways:

  • Ad hoc formats: JSON over stdout, environment variables, mounted files, sockets, shared volumes.
  • Duplicate code: parsing and validation on both sides of every joint.
  • Version drift: one side renames a field, and the other breaks quietly or runs on bad data.
  • Attack surface: wide, undeclared I/O is where attacks land. Tool poisoning is one symptom.
  • No checks up front: nobody can check a composition before it runs.

Secure composition replaces those joints with declared, typed ones. In the series' building analogy, rooms connect only through guarded doors, so you can add a new room without re-inspecting the whole building.

Secure composition for AI agents: isolated tools joined by typed interfaces instead of one shared context

How composition puts the pieces back together

Composition comes last because the other five principles all push you to split an agent into smaller parts. Each split makes the agent safer and adds a joint to secure.

Follow all five and one coding agent becomes many parts: the harness, a fetcher, a code runner, two MCP servers, a skill, and the apps it builds.

Join them through a shared folder and one shared token, and you have rebuilt the single big box you started with. Join them through narrow, typed, declared interfaces, and each part stays as small as you made it. You've put the pieces back together: securely, this time.

What is secure composition?

Secure composition is the practice of combining parts that don't trust each other so the whole stays as safe as its parts. Each part declares what it takes in (inputs), returns (outputs), and what it may reach. Nothing else crosses the boundary.

Take an agent that fetches a web page, summarizes it and posts the summary to Slack. Composed securely, the summarizer has no network access, the fetcher never sees your Slack token, and the poster can reach one host. A poisoned page can spoil a summary, but the summarizer has no way to send your files anywhere.

Three research ideas sit under this:

The strategies, each covered below: typed interfaces, one sandbox per tool, verify before you compose, and re-approve on change.

Controls by layer

LayerControlWhere it applies
BuildSigstore cosign signing, SBOMs, pin by hashAny OS
OSOne bubblewrap sandbox per tool; XPC services; one AppContainer per tool; service mesh mTLS and admission policyLinux; macOS; Windows; Kubernetes
RuntimeCompose Wasm components through WIT3 and WAC, sharing nothingWasm
AgentNamespace tool names; re-consent when tool descriptions changeAny OS

The Unix philosophy: small parts, one interface

The Unix philosophy is a way to build software from small programs that each do one job and connect through a shared interface. The 1978 foreword to the Unix issue of the Bell System Technical Journal put it this way: "Make each program do one thing well" and "Expect the output of every program to become the input to another, as yet unknown, program"4.

The pipe, written |, connects one program's output (stdout) to the next program's input (stdin):

shell
cat access.log | grep " 404 " | sort | uniq -c

Four programs, none written with the others in mind, count the "not found" errors in a web log.

Why it worked

The parts were small, so each was easy to test and replace. The interface was uniform: every program read and wrote a stream of bytes, so any program could plug into any other. Doug McIlroy had asked for this in a 1964 memo, wanting ways of "connecting programs like garden hose"5.

Where it falls short for agents

  1. The interface is untyped bytes. sort doesn't know it is sorting web requests. If the log format adds a column, grep still runs and quietly returns the wrong answer. Every program parses text in and formats text out, the same burden JSON over stdout puts on containers.
  2. Every program has ambient authority. Ambient authority is the power a program gets just by running as you: your home folder, your SSH keys, the whole network. grep only needs a stream of text, but it could read ~/.ssh or open a connection, and nothing in the pipe would stop it.

That was fine for trusted tools that shipped with the OS, but not for an agent that adds a marketplace tool every week. Containers, microVMs and OS sandboxes narrow the second gap. The first gap stays.

WebAssembly components keep what worked and close both gaps. Think of them as Unix pipes with types and without ambient authority.

Unix pipes pass untyped bytes between programs that each hold the user's whole account; Wasm components pass typed values through declared imports.

Typed interfaces, not untyped bytes: the WebAssembly Component Model

The WebAssembly Component Model builds software from components that share no memory and talk only through typed interfaces. A component "is restricted to interact only through the modules' imported and exported functions"6. Joining components this way is called shared-nothing linking: they agree on types, not on a shared memory.

The interfaces are written in WIT (WebAssembly Interface Types). In WIT, a world "describes a set of imports and exports" (WIT reference). Imports are everything a component can reach. Exports are everything it offers. Here are the summarizer and fetcher from the example above:

wit
package example:digest@0.1.0;

interface summarize {
  record summary {
    title: string,
    bullets: list<string>,
  }
  summarize: func(page: string) -> result<summary, string>;
}

interface fetch {
  fetch: func(url: string) -> result<string, string>;
}

world summarizer {
  export summarize;
}

world fetcher {
  import wasi:http/client@0.3.0;
  export fetch;
}

Line by line: summarize takes a page as a string and returns a summary record, or an error. The summarizer world exports that function and imports nothing: no network, no files, no clock. The fetcher world imports outbound HTTP and nothing else. WAC, "a tool for composing WebAssembly components together," wires one component's exports into another's imports.

Back to the five costs from the top of this page:

  • Formats: one WIT contract replaces ad hoc formats.
  • Parsing: generated bindings deliver the summary record as a native type.
  • Drift: packages carry versions, and a type mismatch shows up when you compose the parts.
  • Attack surface: a component reaches only what it imports.
  • Checks: you can read every component's world before anything runs.

The parts don't have to share a language. Per the docs, "A component implemented in Go can communicate directly and safely with a C or Rust component." Your fetcher can be Rust and your summarizer Python, joined by types, not by text each side must parse.

A poisoned summarizer can still write a misleading summary, but it can't read your files or reach the network: its world imports neither. Firefox applies the same idea with RLBox (Narayan et al. 2020).

Unforgeable handles make it safe to pass authority between composed parts: a component can use only what it was handed, and it can't guess a capability, build one from a name or string, or copy one it was never given7. The runtime enforces this, not the component's good behavior: a component can call only the imports the host linked, and the runtime checks every resource handle, so handing one to the summarizer gives it that one thing and nothing more.

The limits, plainly. A component given an import can still misuse it. Code must compile to Wasm, so arbitrary Linux binaries won't run and the harness often stays in a VM or container. WASI is still maturing, and isolation is enforced in software, so the compiler is in the trusted base.

Cosmonic Control runs MCP servers as sandboxed components with a capability-based security model. Cosmonic Desktop, a free app for macOS, Windows and Linux, lets coding agents build tools as components and shows what each can reach before it runs. See capabilities for the model behind it.

Wasm Component Model composition: each component gets only the imports it needs, so a poisoned tool can't reach the network

How authority combines when you compose: additive, transitive, inherited, composable

Typed calls have a side benefit: security teams can audit an agentic workflow from its declarations. When you join parts into one workflow, their authority combines in four ways, and each property decides whether the whole stays as safe as its parts.

  • Additive. When you put parts together, their authority adds up: a workflow can do everything any of its steps can do. With declared imports you can total it up before anything runs, because the composed component's imports are the union of what its parts still need. With containers or scripts, you can't see the sum.
  • Transitive. Authority flows along calls. If step A can call step B, and B can reach the database, then A can reach the database through B. This is how a confused deputy happens8. Keep each interface narrow: in Wasm, A reaches B only through B's typed exports, so what flows through B is only what B's interface allows.
  • Inherited. In an ordinary OS process or container, a child process inherits the parent's whole environment by default: its user, its environment variables (often API keys), its open files and its network. Authority leaks down the tree without anyone deciding it. In the Wasm Component Model nothing is inherited. Each component gets only the imports it is explicitly linked with, and a capability passed to one component is not visible to its siblings.
  • Composable. This is the goal: when two parts are each safe on their own, the combination is still safe, and you can check the whole by checking each part and the typed connections between them. Shared state and inherited authority break this. Typed, shared-nothing interfaces and unforgeable handles preserve it (Miller, 2006, "Robust Composition").

The earlier principles split the work into small parts with little authority each, and these four properties decide whether putting them back together keeps those guarantees.

Composition across the sandbox types

Give every tool, MCP server and skill its own sandbox, so a bad one can only damage itself: an MCP sandbox for every server, not one box around the agent. Agent sandboxes often stop at the shell: "The sandbox covers shell commands only. Claude's file tools, MCP servers, and hooks run outside it"9.

We rated the eight sandbox types in our comparison matrix against composition, with a plain process as a baseline. As elsewhere, we call these "sandbox types" after the conventional usage, though containers and VMs are not sandboxes on their own. Yes means it meets the principle by default; Partial means partly, or with the listed controls; No means not by default. OS sandboxes are rated as configured.

ProcessNo
Pro
Processes can talk over IPC
Con
Plugins and libraries share a process and the user's authority
Controls to add
Move to OS sandboxes, one per tool
Linux sandboxPartial
Pro
One sandbox per tool is cheap
Con
Interfaces between sandboxes are untyped (sockets, files)
Controls to add
One bubblewrap sandbox per tool
macOS sandboxPartial
Pro
XPC services split privileges
Con
Per-tool profiles are manual
Controls to add
XPC services; per-tool Seatbelt profiles (deprecated, no Apple docs)
Windows sandboxPartial
Pro
Separate AppContainers per tool; brokered COM
Con
Interfaces are untyped
Controls to add
One AppContainer per tool
ContainerPartial
Pro
One service per container; mesh mTLS
Con
Trust between services is by network address
Controls to add
Istio mTLS; signed images plus admission policy
gVisorPartial
Pro
One sandbox per tool server is practical
Con
Interfaces between sandboxes are untyped (network)
Controls to add
One runsc sandbox per tool server (gVisor)
VMPartial
Pro
Strong separation between trust levels
Con
Machine granularity; heavy
Controls to add
Separate VMs per trust level, as in Qubes OS
microVMPartial
Pro
Cheap per-tool VMs
Con
Interfaces are untyped (network, vsock)
Controls to add
One Firecracker microVM per tool server
WasmYes
Pro
Typed WIT interfaces, shared nothing
Con
A component given an import can still misuse it
Controls to add
Sign components; pin by digest; grant only the imports each needs

A plain process runs every tool with your whole account, so one bad tool is everyone's problem. OS sandboxes start wide open, and you add layers to close them: bubblewrap, Landlock and seccomp on Linux, Seatbelt on macOS, AppContainer on Windows. They, containers, gVisor and VMs can give each tool its own box, but the joints between boxes stay untyped. Wasm starts with typed, shared-nothing joints. For a side-by-side comparison, see containers vs microVMs vs Wasm.

MCP tool poisoning: a symptom of untyped, shared context

MCP tool poisoning is an attack where a malicious MCP server hides instructions in a tool's description, which the model reads and the user usually never sees10. It is a composition failure.

MCP does type part of each tool: an inputSchema, and optionally an outputSchema11. But the description is free text, and every server's descriptions land in one shared model context. The model becomes the joint between tools, and it reads everything as instructions it might follow. The spec says clients "MUST consider tool annotations to be untrusted unless they come from trusted servers."

Two variants share that gap:

  • Tool shadowing: one server's description changes how the agent uses another server's tool, as with the WhatsApp MCP "sleeper" server.
  • Rug pull: a tool changes after you approved it. In Cursor MCPoison (CVE-2025-54136), an approved config was swapped for malicious commands and still trusted by name.

With typed composition, a component's imports, not its description, decide what it can reach. For the full list of threats and fixes, see MCP security.

Running MCP servers as Wasm components

You can build an MCP server as a WebAssembly component instead of a native process or a container. The server's code runs in a Wasm runtime such as Wasmtime, and it can reach only the interfaces it imports: no file system, network or environment variables unless the host grants them. Two tools do this today:

  • Wassette (Microsoft, 2025) is a runtime that loads components and exposes their typed WIT interfaces as MCP tools. Its components "can't access system resources without explicit access permissions."
  • Cosmonic Desktop is a free app for macOS, Windows and Linux that runs components in a local deny-by-default sandbox. Its integrated MCP server lets coding agents such as Claude Code, Codex and Gemini CLI build tools as components, deploy them into that sandbox, and show what each one can reach before it runs.

What that means in practice:

  1. The server's reach is written down. Its imports are its full list of permissions, and you can read them before it runs. A server that formats dates has no network import, so it can't send anything anywhere.
  2. Servers share nothing. Each server is its own component with its own memory and its own grants, so one server can't read another's files or credentials.
  3. A rug pull has to ask. A new version that wants a new host or a new folder needs a new grant, which the host can refuse or flag for re-approval. Changes inside what a server already holds still get through, such as the postmark-mcp backdoor (a malicious MCP server, in the table below) adding a BCC to mail it was already allowed to send. That is why pinning and re-approval stay on the list.

Wasm does not fix the description problem on its own. A Wasm server's tool descriptions still land in the shared model context, so a poisoned description can still steer the model into misusing other tools. Wasm limits what a server can do; namespacing, showing full descriptions and re-approval handle what it says. In production, Cosmonic Control runs MCP servers as sandboxed components on Kubernetes.

The evidence: eight composition failures

Each case broke at a joint between parts. Newest first:

DateCaseWhat was shared or trusted by nameWhat happened
Feb 2026SANDWORM_MODE npm wormAgent config files any package could writeTypo-squatted packages added a fake MCP server to Claude Desktop, Cursor and Windsurf configs that told the AI to read SSH keys
Feb 2026ClawHavoc, OpenClaw skillsA skills marketplaceHundreds of fake skills tricked users into installing infostealers
Jan 2026Git MCP plus Filesystem MCP chain (CVE-2025-68143/4/5)Two servers in one agentEach Git MCP flaw was limited alone; combined with the Filesystem MCP server they reached code execution
Sep 2025postmark-mcp backdoorA package trusted by name across versionsNormal for 15 versions, then one line BCC'd every email to the attacker
Jul 2025Cursor MCPoison (CVE-2025-54136)An approved MCP config, trusted by nameThe attacker swapped in malicious commands after approval
Jul 2025mcp-remote (CVE-2025-6514)A popular proxy between client and remote serverAn untrusted remote MCP server could run commands on the user's machine
Apr 2025WhatsApp MCP exfiltrationOne agent context for two serversA sleeper server made the agent forward chat history to an attacker
Apr 2025MCP tool poisoning, shadowing and rug pullsTool descriptions in the shared contextHidden instructions, changes after approval, redirected trusted tools

Across these vulnerabilities we see two root causes repeat. The first is shared state: tool descriptions share the model's context, plugins share a process, and servers share a user's credentials. The second is trust by name: a part approved once is trusted forever, as with MCPoison and postmark-mcp. The Git plus Filesystem chain shows harm adding up across parts, which is why defense in depth matters.

AI agent supply chain security: verify what you compose

Before you plug a tool, MCP server or skill into an agent, prove who built it, pin the exact version, and re-approve it when it changes. Ken Thompson's "Reflections on Trusting Trust" (1984) showed you can't fully trust code you didn't build. In the MCP supply chain, SANDWORM_MODE and postmark-mcp both arrived as ordinary packages.

  1. Sign and verify. Sigstore cosign ties a signature to an identity, keyless by default.
  2. Record provenance for each build step with in-toto12.
  3. Keep an SBOM, a list of every component in a build.
  4. Pin by digest, not tag13. Pinned images don't pick up patches, so schedule updates.
  5. Re-approve on any change to code, config or tool description.
  6. Enforce at deploy time: reject unsigned or unpinned workloads with Kubernetes ValidatingAdmissionPolicy (GA in v1.30).

How to compose an agent safely (and prevent MCP tool poisoning)

  1. Define each part's inputs and outputs as a typed interface, not JSON over stdout or a shared folder.
  2. Give every tool, MCP server and skill its own sandbox.
  3. Grant each part only the imports it needs. A part with no network import can't leak.
  4. Give each server its own scoped credential, never a shared all-repo token (least authority).
  5. Keep the part that reads private data apart from the part that can send data out (controlled information flow).
  6. Show users full tool descriptions, and namespace tool names so one server can't shadow another.
  7. Pin every part by digest, and re-approve on any change.
  8. Never auto-start new servers (deny by default).
  9. Log every tool call, so you can trace a bad result to the part that produced it.

Composition and defense in depth

Composition joins parts side by side. Defense in depth stacks boxes inside each other; the principles hub compares weak, dense and strong stacks.

To apply it, decompose the agent, sandbox each part, then compose them back, as in the six-step method.

Running tools and MCP servers in production? Cosmonic Control composes them as sandboxed Wasm components on Kubernetes, with egress denied by default and OpenTelemetry observability. To start on a laptop, download Cosmonic Desktop.

Run every MCP server in its own sandbox

Public beta

Cosmonic Control runs MCP servers sandboxed, with a capability-based security model. Build and run a tool as a component locally with Cosmonic Desktop.

Cosmonic Desktop's Inspect view: a component's declared interfaces and network egress drawn out, with everything else denied by default.

Frequently asked questions

What is secure composition?
Secure composition is the practice of building a system from isolated parts that connect only through declared, typed interfaces. Each part can be checked alone, and a bad part can't spoil the others. For AI agents, it means one sandbox per tool and typed connections instead of shared folders or one shared model context.
What is the WebAssembly Component Model?
The WebAssembly Component Model is a standard for building programs from WebAssembly components that share no memory and connect only through typed interfaces written in WIT. A component's imports list everything it can reach, so you can see its full set of permissions before it runs.
What is typed I/O, and why does it matter for AI agents?
Typed I/O means every input and output between parts has a declared shape, such as a record with a title and a list of bullets. Mismatches show up when you compose the parts, not in production. It also narrows the attack surface: a part can't receive or reach anything it didn't declare.
Can tools written in different languages be composed safely?
Yes, if they share a typed contract. With the WebAssembly Component Model, a Rust component can call one written in Go or Python through a WIT interface. Bindings are generated for each language, so neither side hand-parses text, and the two never share memory.
How are Wasm components different from Unix pipes?
Both build software from small parts joined in a chain. Unix pipes pass untyped bytes, and every program runs with the user's full authority. Wasm components pass typed values, and each reaches only what it imports. Think of them as Unix pipes with types and without ambient authority.
What is MCP tool poisoning?
MCP tool poisoning is an attack where a malicious MCP server hides instructions in a tool's description, which the model follows and the user rarely sees. It works because every server's descriptions share one model context. Defend with one sandbox per server, pinning, and re-approval whenever a tool changes.

Further reading

  1. OCI image specification, image configuration. github.com ↩
  2. Miller (2006), “Robust Composition”, PhD dissertation, Johns Hopkins University. erights.org ↩
  3. The Component Model book, WIT reference. component-model.bytecodealliance.org ↩
  4. McIlroy, Pinson and Tague (1978), “UNIX Time-Sharing System: Foreword”, Bell System Technical Journal. archive.org ↩
  5. McIlroy (1964), the memo proposing pipes. Bell Labs. nokia.com ↩
  6. Bytecode Alliance, WebAssembly Component Model: why the component model. component-model.bytecodealliance.org ↩
  7. Dennis and Van Horn (1966), “Programming Semantics for Multiprogrammed Computations”. doi.org ↩
  8. Hardy (1988), “The Confused Deputy”. doi.org ↩
  9. Anthropic, Claude Code sandboxing documentation. code.claude.com ↩
  10. Invariant Labs, “MCP security notification: tool poisoning attacks”, April 2025. invariantlabs.ai ↩
  11. Model Context Protocol specification (2025-06-18), Tools. modelcontextprotocol.io ↩
  12. Torres-Arias et al. (2019), “in-toto: Providing farm-to-table guarantees for bits and bytes”, USENIX Security. usenix.org ↩
  13. Docker CLI reference, docker image pull (pull by digest). docs.docker.com ↩