What it is
An agent sandbox is a walled-off space in which an AI agent does its work. Inside the walls, the agent can run code, edit files and browse. Outside them, it cannot reach anything it was not explicitly given.
The word comes from ordinary software security, where a sandbox runs untrusted programs in isolation. AI agents need the same idea because they take actions on their own, and a model that misreads an instruction or is tricked by a malicious web page can do real damage if it has the run of your machine.
How it works
A sandbox enforces limits outside the model, in the operating system or a container, so the agent cannot talk its way past them. Typical limits include:
- Files: the agent sees only a chosen folder, not your whole disk.
- Network: it may reach an approved list of sites, or none at all.
- Tools and permissions: it can call only the tools it was granted, with narrow access.
- Resources: caps on time, memory and spending stop runaway jobs.
- Logging: every action is recorded so you can review what happened.
The key point is that these rules are enforced by the environment, not requested politely in a prompt. A prompt can be ignored or overridden by a clever attacker. A sandbox boundary cannot be argued with.
Why it matters to you
Agents are useful because they can act, and risky for the same reason. A sandbox lets you give an agent real capability while keeping the worst case small. If you run coding agents, automation or browser agents, the sandbox is often the difference between a contained mistake and a deleted folder or a leaked credential.
Limits
A sandbox only protects what it isolates. If you hand the agent a password, a broad API key or write access to your production database, the agent can use them. Start with the least access that gets the job done, and widen it only when you must. Sandboxes are one layer of AI guardrails, and are worth testing with red teaming.
An example
You ask a coding agent to fix a bug. In a sandbox, it works in a copy of one project folder, can install packages only from an approved source, has no access to your email or cloud accounts, and shows you the changes before anything is merged. If a hidden instruction in a dependency tells it to upload your files, there is nowhere for them to go.
Standards such as the Model Context Protocol connect agents to outside tools, which makes clear permission limits and sandboxing even more important.