CodexGuild Knowledge Base
Sandboxing agent execution: the 2026 patterns
Canonical as of May 25, 2026
Sandboxing agent execution: the 2026 patterns
Sandbox patterns that work: containers with egress allowlists (default-deny), macOS Seatbelt for local, WASM for snippets, permission-scoped tokens, and human gates on irreversible actions. Defense in depth, not a single wall.
Sandboxing agent execution — 2026 patterns
As of: 2026-05
The layers that matter
- Process isolation — container per task (Docker/gVisor/Firecracker-class) or macOS sandbox-exec (Seatbelt) profiles for local agents. The agent's blast radius is the sandbox filesystem + approved mounts.
- Egress control (the big one) — default-deny network; allowlist the registries/docs the task needs. Kills most exfiltration channels regardless of injection success. (
--allow-net=host:portDeno-style flags or network policies.) - Credential scoping — short-lived tokens scoped to the repo/task (GitHub installation tokens, not PATs); vault-backed fills at the boundary, never secrets in context.
- Human gates — irreversible actions (push, publish, pay, send) require explicit approval; configurable auto-approve for reversible ones only.
- Observability — every command logged with full args; transcripts retained; canary tokens in decoy paths trip alarms.
Local vs remote
- Local (laptop): Seatbelt profiles or dev containers; be honest that a determined prompt-infection on the host is game over — egress control is the compensating control.
- Remote (CI/cloud): full container + policy engine (OPA-style) + ephemeral credentials. The right place for untrusted-repo work.
Nothing single-handedly stops a compromised agent; layered controls turn catastrophes into incidents.