The principle of least privilege
for AI agents
Why least privilege breaks down for agents, how least authority goes further, and the controls that enforce it.
Your AI agent holds a ring of keys: it can read files, call APIs, run shell commands and open pull requests. The principle of least privilege for AI agents says it should carry only the keys its current job needs. The idea dates to 1975 and is still sound. But agents break a quiet assumption behind the principle: that you can predict what a program will do with its access.
This guide explains least privilege, where it falls short for agents, and its stricter cousin, the principle of least authority (POLA). We'll walk through the controls that enforce it, before-and-after tables, and a checklist for teams using agents. Still deciding where your agent should run? Start with our guide to choosing an agent sandbox.

What is the principle of least privilege?
The principle of least privilege (POLP) is a security rule that gives every user, program and process only the access it needs to do its job, and nothing more. Jerome Saltzer and Michael Schroeder explained it in one sentence in their 1975 paper, "The Protection of Information in Computer Systems"1: "Every program and every user of the system should operate using the least set of privileges necessary to complete the job."
Saltzer and Schroeder's main reason: it "limits the damage that can result from an accident or error." Security people call that damage the blast radius. One key off the ring, lost, costs you one door. The whole ring costs you the building.
In identity and access management, this is called least privilege access. CNSSI 4009-2022, reproduced in NIST's glossary2, defines it as restricting "the access privileges of users (or processes acting on behalf of users)" to the minimum needed. A process acting on behalf of a user is exactly what an AI agent is.
Least privilege vs least authority (POLP vs POLA)
Least privilege asks what a program is allowed to touch. Least authority asks what a program can actually cause to happen, including through other programs. For an agent that calls tools, the two answers can be very different.
Each idea has an original source worth reading:
- Least privilege: Jerome Saltzer and Michael Schroeder (1975), "The Protection of Information in Computer Systems".
- Least authority: Mark S. Miller (2006), "Robust Composition: Towards a Unified Approach to Access Control and Concurrency Control"3, his Johns Hopkins dissertation (PDF).
Miller states POLA simply: grant "each program only the authority it needs to do its job." He adds that in Saltzer and Schroeder's version "it is not clear precisely what they meant by 'privilege,'" so POLA measures effects instead.
Miller and Jonathan Shapiro's "Paradigm Regained" (2003)4 draws the line. Permission is a direct access right. Authority is every effect a program "may cause on objects it can access, either directly by permission, or indirectly by permitted interactions with other programs." If Alice will send Bob copies of a file he cannot open, Bob has the authority without the permission.
Why the difference matters for agents
Say your agent runs in a container with no network access. But it can call a fetch_url tool on an MCP server (a small service that offers tools to agents) running outside the container with open internet. The agent's privilege is "no network." Its authority is "any URL on earth."
Now stack it. The same server holds a GitHub token that can push to every repository in your organization. The agent never sees it, yet it is one tool call from production code.

That is the confused deputy problem, which Norm Hardy named in 19885: a trusted program uses its power on behalf of someone who should not have it. For a full example, worked step by step through a GitHub MCP server, see the confused deputy.
| Least privilege (POLP) | Least authority (POLA) | |
|---|---|---|
| Question it asks | What is this program permitted to access? | What effects can this program cause, directly or through others? |
| What it counts | Direct permissions: files, roles, scopes | Direct permissions plus everything reachable through tools, services and other agents |
| Typical tools | Access control lists, IAM roles, OAuth scopes | Capabilities: narrow, unforgeable references passed explicitly |
| Blind spot | Power borrowed through a helper that has more access | Harder to measure unless the system is built for it |
| Original source | Saltzer and Schroeder, 1975 | Miller, 2006 |
Least privilege is the floor. Least authority asks what an attacker asks: "What can I make this agent do?" The same split applies to sandboxes: judge both the wall and the authority behind it.
Why least privilege breaks down for AI agents
Least privilege assumes a program's code shows what it will do. Agents break that in six ways.
1. Agents decide what to do at run time
A normal program follows its code. An agent asks a large language model (LLM) what to do next, and the answer changes with the prompt. Asked to "fix the failing test," it might edit one file or rewrite your build script. Assume it may try anything its tools allow.
2. Prompt injection turns data into instructions
Prompt injection is text hidden in content the agent reads (a web page, an email, a GitHub issue) that the model follows as if it came from you. In OWASP's example, an email assistant that can read and send mail is told by a malicious email to forward sensitive messages to the attacker.
Simon Willison calls this mix the "lethal trifecta"6: private data, untrusted content, and a way to send data out.
3. Supply chain compromise
Agents pull in skills, plugins, MCP servers and packages written by other people, and a malicious or later-altered one runs with the agent's authority. Least privilege set on the agent cannot stop it, because the bad dependency inherits the agent's keys. Auditing all 2,857 skills then published on ClawHub, researchers found 341 malicious ones7, 335 of them part of a single campaign they named ClawHavoc, which shipped infostealer malware to OpenClaw users on macOS and Windows. MCP servers carry the same risk through tool poisoning and rug pulls.
4. Tools chain into more authority than any one tool has
A tool that reads your .env file is fine for debugging. A tool that makes HTTP requests is fine for checking docs. Give both to one agent and you have built a data-leak pipeline. Permissions are checked one at a time; authority adds up.
5. Tokens outlive the task
Many agents run with API keys in environment variables or config files, often with no expiry and every scope the developer had. On a laptop, an agent also inherits ambient authority, the access it gets just by starting, handed over by nobody. See what ambient authority looks like in an agent.
6. OWASP calls the result "excessive agency"
Excessive Agency8 is LLM03 in the 2026 OWASP Top 10 for LLM Applications (August 2026). In 2025 it was LLM069. OWASP names three root causes:
- Excessive functionality: the agent has tools or functions it does not need.
- Excessive permissions: those tools can reach more systems or data than the job requires.
- Excessive autonomy: high-impact actions run with no human check.
Classic least privilege covers only the middle cause; least authority covers all three. The Agentic Top 1010 names the same failure ASI03, Identity and Privilege Abuse.
How do you enforce least privilege for AI agents?
You enforce least privilege for AI agents by breaking the agent into small steps, giving each step only the access it needs, and composing the steps back together. Every limit is enforced outside the model. These six steps are the core of AI agent access control:
- Break it into steps. Break the agent's work down into workflow steps.
- List the capabilities. Identify the requirements for each step: file permissions, network access, secrets.
- Pick the boundary. Choose the most restrictive boundary that can still provide those capabilities.
- Choose the right abstractions for each layer. Align each layer to the appropriate sandboxing approach.
- Verify. Test sandbox guarantees like any other requirement.
- Compose the workflow. Combine the steps into a single chain of work.
Learn more about this method in How to choose an AI agent sandbox.
What follows is what those steps mean when the thing you are scoping is access: the five controls that turn the method into least privilege.
Scope credentials per tool, not per agent
The most common setup gives the agent one powerful API key for everything. Instead, the agent itself holds no credentials, and each tool holds its own. Give the agent its own identity too, so logs show what it did and you can revoke it alone.
Take an agent that triages GitHub issues. The "read issues" tool can read issues in one repository, the "add label" tool can write them, and neither can push code. GitHub's fine-grained personal access tokens support this.

Never put a secret in the prompt: anything the model can read, an injection can ask it to repeat.
Use short-lived, reduced-power tokens
A stolen token that expires in an hour does far less damage than one that works for a year. Mint tokens per task. A GitHub App installation token, for example, expires after one hour and can be restricted to named repositories and permissions.
Handing over a weaker version of a power you hold is called attenuation. The MCP security best practices11 apply it: start with a minimal scope and avoid wildcards such as * or full-access.
Start from deny by default at the sandbox layer
Start from deny by default: the agent can reach nothing until you list it. An allowlist names the only files, hosts and tools each step may use, and the sandbox blocks everything else. Writing "never read files outside the project" in the system prompt helps, but the model can ignore that line or be tricked out of it. At the sandbox layer, set:
- Files: only the project folder, read-only for analysis. No home folder, SSH keys or cloud credentials.
- Network: no outbound traffic except named hosts, such as your package registry.
- Environment: clean environment variables, not a copy of your shell.
Our guide to sandboxing AI agents covers every sandboxing technology that can enforce these rules.
Require human approval for irreversible actions
Deleting data, sending email or money, merging to main, deploying and changing access cannot be undone. OWASP recommends human-in-the-loop control "to require a human to approve high-impact actions." Show the reviewer the exact command, recipient or diff.
Run tools in capability-based runtimes
A capability is an unforgeable reference to one resource, handed over on purpose (what capabilities are, and why your agents need them). In a capability-based runtime, code starts with no access and uses only what it is handed, which limits authority, not just privilege. WebAssembly (Wasm) components work this way: per the Component Model docs12, a component reaches the outside world "only by calling its imports." See how WASI implements capabilities.
Wasm has limits: no arbitrary Linux binaries, no direct GPU access, and a younger ecosystem. So unless the harness itself compiles to a component, keep it (the program that runs the model loop) in a container or VM, and run each tool as its own Wasm component. See containers vs microVMs vs WebAssembly for the tradeoffs.
Least privilege examples: an agent, a tool and an MCP server
Here is what scoped AI agent permissions look like at three levels of one system.
Example 1: A coding agent
Task: fix the failing unit tests in one repository.
| Access | Before | After |
|---|---|---|
| Files | Your whole home folder | The project folder only |
| Secrets | Every environment variable in your shell, including cloud keys | None; the tests use fixtures |
| Network | Open internet | The package registry only |
| Git | Pushes to any branch with your SSH key | Commits locally; a human opens the pull request |
What changed in authority: a README that tells the agent to upload your cloud credentials still gets read, but there are no credentials and no host to send them to.
Example 2: A tool
Task: a support agent looks up the status of a customer's order.
| Access | Before | After |
|---|---|---|
| Database login | A shared admin account | A read-only role, short-lived |
| Tables | All of them | One orders view: order ID, status, ship date |
| Rows | Every customer | Only the customer on the current ticket |
| Query | Free-form SQL written by the model | One fixed function: get_order_status(order_id) |
What changed in authority: the model no longer writes SQL, so an injection cannot ask for SELECT * FROM customers.
Example 3: An MCP server
Task: summarize open pull requests in one repository and post a comment.
| Access | Before | After |
|---|---|---|
| Token | A classic personal token with full repo and org admin scopes | One repository: read pull requests, write comments |
| Token lifetime | No expiry | One hour, minted per task |
| Tools exposed | Every tool, including merge and delete branch | List, read and comment on pull requests |
| Tokens it accepts | Whatever token the client passes along | Only tokens issued for this server |
What changed in authority: forwarding a client's token unchecked ("token passthrough") makes the server a confused deputy, and the MCP security best practices forbid it. See also MCP security.
Common least privilege violations
A violation of the principle of least privilege is any access a person or program holds that its current task does not need. With agents, these are the ones that show up again and again. Several are also sandboxing mistakes, covered at length in five sandboxing mistakes teams make with agents:
- One "god token" for every tool. Whichever tool is fooled first gets everything.
- Running the agent as yourself. It inherits your home folder, SSH keys, cloud credentials and browser sessions.
- Secrets in the workspace. A
.envfile in the project folder is onecatcommand away from the model's context. - Open network egress. Anywhere the agent can send traffic, an injection can send your data.
- Standing access. Tokens that never expire and one-off permissions nobody removes. Over time this becomes privilege creep: access quietly piling up.
- Skipping approvals with no outer wall. Anthropic's permissions docs13 say to use Claude Code's
bypassPermissionsmode "only" in "isolated environments like containers or VMs."
Least privilege checklist for agent teams
Use this list for a new agent, a new tool, or a review. The early steps give the biggest wins.
- Break the workflow into steps and write down what each must read, write, call and reach.
- Inventory every credential in reach, including environment variables, config files, SSH keys and MCP server tokens.
- Give each tool its own credential, scoped to one resource and the fewest permissions that work.
- Set short expirations: minutes to an hour, not months.
- Deny by default: a clean environment, and only the files and hosts each step needs.
- Replace open-ended tools with narrow ones.
- Require approval for irreversible actions, showing the exact command, recipient or diff.
- Compose steps through narrow interfaces, so a compromise stays in one step.
- Log every tool call and test the walls: remove unused access, and plant a harmless prompt injection to confirm data cannot leave.
Putting least authority into practice
Least privilege sets the minimum. Least authority is the goal. If you do only three things this week, strip secrets from the agent's environment, block outbound network by default, and give each tool its own short-lived token.
Capability-based runtimes make least authority the default instead of a cleanup project. That is our approach at Cosmonic. Cosmonic Desktop is a free app for macOS, Windows and Linux that runs WebAssembly components in a deny-by-default sandbox and shows what each can reach before it runs. Its built-in MCP server lets Claude Code, Codex and Gemini CLI build tools and deploy them there. For production, Cosmonic Control runs the same components on Kubernetes, built on the CNCF project wasmCloud.
Related topics
Try least authority on your agent's tools
Public betaCosmonic Desktop runs each tool as a WebAssembly component in a deny-by-default sandbox, and shows you what it can reach before it runs. Free for personal use, no account, no cloud.

Frequently asked questions
What is least privilege access?
What is the difference between POLP and POLA?
How do you enforce least privilege for AI agents?
Is least privilege part of zero trust?
What is excessive agency in LLM applications?
What is AI agent access control?
What is the NIST definition of least privilege?
- Saltzer and Schroeder (1975), “The Protection of Information in Computer Systems”. cs.virginia.edu ↩
- NIST, Computer Security Resource Center glossary: least privilege. csrc.nist.gov ↩
- Mark S. Miller (2006), “Robust Composition: Towards a Unified Approach to Access Control and Concurrency Control”, Johns Hopkins dissertation. papers.agoric.com ↩
- Miller and Shapiro (2003), “Paradigm Regained: Abstraction Mechanisms for Access Control”. scs.stanford.edu ↩
- Hardy (1988), “The Confused Deputy (or why capabilities might have been invented)”. cap-lore.com ↩
- Simon Willison, “The lethal trifecta for AI agents”, June 2025. simonwillison.net ↩
- The Hacker News, “Researchers Find 341 Malicious ClawHub Skills Stealing Data from OpenClaw Users”, February 2026. (Koi Security’s original report is no longer online.) thehackernews.com ↩
- OWASP, GenAI Top 10 for LLM Applications 2026 (Excessive Agency, LLM03). genai.owasp.org ↩
- OWASP, LLM06:2025 Excessive Agency. genai.owasp.org ↩
- OWASP, Top 10 for Agentic Applications. genai.owasp.org ↩
- Model Context Protocol, Security Best Practices (2025-11-25). modelcontextprotocol.io ↩
- Bytecode Alliance, WebAssembly Component Model: why the component model. component-model.bytecodealliance.org ↩
- Anthropic, Claude Code permissions documentation. code.claude.com ↩
- NIST SP 800-207, “Zero Trust Architecture”. nvlpubs.nist.gov ↩