Knowledge base
CodexGuild Knowledge Base

Sandboxing agent execution: the 2026 patterns

as of May 25, 2026 · canonical · codexguild.com/kb/kb-agent-sandboxing-2026 · exported 2026-10-11
Canonical as of May 25, 2026

Sandboxing agent execution: the 2026 patterns

Sandbox patterns that work: containers with egress allowlists (default-deny), macOS Seatbelt for local, WASM for snippets, permission-scoped tokens, and human gates on irreversible actions. Defense in depth, not a single wall.

Sandboxing agent execution — 2026 patterns

As of: 2026-05

The layers that matter

  1. Process isolation — container per task (Docker/gVisor/Firecracker-class) or macOS sandbox-exec (Seatbelt) profiles for local agents. The agent's blast radius is the sandbox filesystem + approved mounts.
  2. Egress control (the big one) — default-deny network; allowlist the registries/docs the task needs. Kills most exfiltration channels regardless of injection success. (--allow-net=host:port Deno-style flags or network policies.)
  3. Credential scoping — short-lived tokens scoped to the repo/task (GitHub installation tokens, not PATs); vault-backed fills at the boundary, never secrets in context.
  4. Human gates — irreversible actions (push, publish, pay, send) require explicit approval; configurable auto-approve for reversible ones only.
  5. Observability — every command logged with full args; transcripts retained; canary tokens in decoy paths trip alarms.

Local vs remote

  • Local (laptop): Seatbelt profiles or dev containers; be honest that a determined prompt-infection on the host is game over — egress control is the compensating control.
  • Remote (CI/cloud): full container + policy engine (OPA-style) + ephemeral credentials. The right place for untrusted-repo work.

Nothing single-handedly stops a compromised agent; layered controls turn catastrophes into incidents.