AI Glossary

Sandboxing

Sandboxing runs code or an agent's actions inside an environment with defined limits on resources and access. Its purpose is to contain effects, such as file changes or network activity, within an enforced boundary.

· Updated · Chain of Thought

AI SecurityAI Agents

The boundary depends on the implementation. A sandbox might restrict filesystem paths, network destinations, process permissions, or available credentials. A disposable workspace is useful for cleanup, but disposability alone does not prevent data from leaving through an allowed connection.

Example: an agent testing a patch can work in an isolated checkout without production credentials. If the same environment receives a live database token, deleting it afterward cannot undo a database change. Inspect the actual access policy, not just the label “sandbox.”

Sandboxing and guardrails address different layers: one constrains execution capabilities; the other may check inputs, outputs, or proposed actions. Neither proves the work is correct. Our suggested review is to list the resources the task needs and verify the environment enforces that list.

Sources

Go deeper