AI Security
Sandboxing AI Agents and Generated Code
A practical architecture for containing AI-generated code, shell commands, browser actions, and artifacts.
An AI agent may propose shell commands, code, browser actions, or file changes, but a proposal is not a trusted instruction. The application must decide whether to execute it, under which identity, and inside what boundary. Sandboxing contains the consequences when a model is mistaken, manipulated, or simply overconfident.
Why generated execution needs isolation
Model output is data produced by a probabilistic component. It can be syntactically valid while pointing at the wrong file, network, package, or account. Treating a completion as if an operator typed it grants the model all privileges inherited by the host process. A safer design inserts a policy gateway between proposal and execution. The gateway validates the requested operation, establishes a task identity, and provisions a restricted environment. Approval and authorization remain outside the model.
Execution creates risks that ordinary chat does not: filesystem modification, credential discovery, process creation, network access, persistence, lateral movement, environment-variable leakage, and denial of service. Generated code can also install arbitrary dependencies or write artifacts that later become trusted. The objective is not perfect prediction of model behavior; it is reducing blast radius when behavior is wrong.
Define the sandbox boundary
A useful boundary specifies filesystem visibility, network destinations, process and syscall policy, credentials, kernel or host access, resource quotas, and lifetime. Start with no host home-directory mount, no ambient cloud credentials, no privileged device access, and no ability to change the sandbox policy itself. Add one capability at a time for a documented task. The policy should be machine-enforced and versioned.
Filesystem policy commonly uses a read-only base image and a temporary writable workspace. Mount only the input files required for the task. Restrict path traversal, symlinks, archive extraction, and artifact export. Treat every file produced by the workload as untrusted until a result gateway scans it and checks its type, size, provenance, and destination policy.
Network policy should default to deny where feasible. If a task needs package downloads or a public API, allowlist hosts and protocols through a proxy. Observe DNS, redirects, and connection attempts. Do not assume an approved hostname remains safe after an unreviewed redirect. Block metadata services and internal control-plane addresses unless the task has a compelling, separately approved need.
Containers, VMs, and runtime sandboxes
Containers provide packaging, namespaces, and resource controls, but they share a host kernel and are not automatically a complete security boundary. A container escape, vulnerable runtime, excessive capability, or mounted socket can defeat the intended isolation. Harden images, drop privileges, remove unnecessary capabilities, and keep the host and runtime patched.
A virtual machine or microVM adds a stronger boundary at higher startup and operational cost. Language-level sandboxes can be efficient for restricted expressions but are unsafe when they expose native escapes or unmaintained interpreters. Choose based on threat model, workload, and acceptable latency; defense in depth matters more than a technology label. Firecracker and gVisor documentation are useful primary examples, not universal prescriptions.
Ephemeral execution
Provision a fresh environment for each task or short batch, then destroy it. Ephemerality removes persistence from prior jobs, limits cross-task contamination, and makes rollback simpler. Keep only explicitly exported results in a separate gateway. Do not reuse writable layers, caches, browser profiles, or credentials across tenants without a documented isolation guarantee.
Use execution deadlines, process limits, CPU and memory quotas, storage caps, and output-size limits. These controls contain infinite loops, fork bombs, oversized archives, and accidental expensive workloads. A supervisor outside the sandbox should be able to terminate it even if the agent process is stuck.
Credentials and dependencies
The sandbox should not inherit broad host secrets. Use the credential-broker pattern from Secure Credential Handling for AI Agents: the agent requests a capability, the broker checks policy, and a short-lived scoped credential is injected only into the adapter that needs it. Never put master keys in prompts, environment variables visible to arbitrary child processes, or copied configuration files.
Generated dependency installation is a supply-chain event. Pin versions and trusted registries, use lockfiles, scan packages, and install only inside the disposable environment. Do not let a package install modify the production image or CI host. AI Supply Chain Security and Model Security provide related provenance controls.
Human approval and policy
Require approval for destructive file operations, external messages, production changes, privileged commands, unusual destinations, or actions with large blast radius. Human Approval for AI Agents explains how to bind approval to exact parameters. Low-risk formatting or test execution can run under pre-approved limits. A reviewer should see the command, target, inputs, requested network, expected effect, and expiry.
Secure execution architecture
Use this reference path:
Agent → Action Proposal → Policy Gateway → Sandbox Provisioner → Restricted Execution Environment → Output Inspection → Artifact/Result Gateway → Agent
The proposal carries a typed operation and correlation ID. The policy gateway checks identity, tenant, resource, and risk. The provisioner selects an image, mounts, network policy, quotas, and lifetime. The result gateway validates outputs before they return to the agent or another interpreter. Each hop emits evidence and can fail closed.
Sandbox policy matrix
| Capability | Default | Allowed condition | Control | Evidence |
|---|---|---|---|---|
| Host filesystem | Deny | Explicit task input | Read-only mounts, path filter | Mount manifest |
| Network egress | Deny | Approved endpoint | Proxy, DNS log, redirect check | Connection record |
| Package install | Deny | Reproducible build task | Lockfile, registry allowlist | Package manifest |
| Credentials | None | Broker-approved adapter | Short-lived scoped token | Grant ID |
| Process creation | Limited | Known runtime | seccomp/job object, count limit | Process events |
| CPU/memory/storage | Quota | Task budget | cgroups or VM limits | Usage metrics |
| Artifact export | Quarantine | Type and policy pass | Malware/content scan | Export digest |
Logging and response
Record code and input hashes, command or tool name, image digest, policy version, identity, network requests, resource usage, exit status, and exported artifact digests. Redact secrets and personal data. Keep logs outside the sandbox so a compromised process cannot rewrite its history. Alert on denied capability requests, unexpected destinations, privilege attempts, repeated failures, or quota exhaustion.
When a sandbox is suspected compromised, stop the workload, revoke broker grants, preserve logs and disk evidence where safe, invalidate exported artifacts, and identify downstream consumers. Rebuild from a trusted image rather than repairing an unknown mutable environment. Test the kill switch, tenant isolation, and rollback path regularly.
Practical checklist
- Treat every model-generated command as an untrusted proposal.
- Enforce policy outside the model before provisioning execution.
- Use fresh, least-privileged environments with explicit mounts and egress.
- Remove host credentials, sockets, metadata access, and unnecessary devices.
- Apply CPU, memory, process, storage, output, and time limits.
- Pin and scan dependencies in disposable environments.
- Require step-up approval for destructive or externally visible effects.
- Inspect exported files before they enter trusted systems.
- Log hashes, decisions, requests, results, and artifact lineage without secrets.
- Revoke, destroy, and rebuild when compromise is suspected.
Sources
- NIST SP 800-190, Application Container Security Guide
- NIST SP 800-53, System and Services Acquisition and boundary controls
- OWASP Top 10 for LLM Applications and Agentic AI guidance
- MITRE ATLAS
- Firecracker and gVisor official security documentation