docker-agent-deploy
Use this skill when exposing a Docker Agent as a server (MCP, HTTP API, A2A, ACP, or OpenAI-compatible chat), distributing an agent via an OCI registry with `docker agent share`, or measuring agent quality with `docker agent eval`. Even if the user just says they want to "turn my agent into an MCP s
- 0
- Installs
- —
- Rating
- —
- Success rate
- 7
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 0bad486bf243e175… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Docker Agent: Serving, Sharing, and Evaluating
Overview
This skill owns the integration surface of Docker Agent: making an agent
reachable by other software (docker agent serve), distributing it through
an OCI registry the way container images are distributed (docker agent share), and proving it still behaves after a change (docker agent eval).
It assumes the agent config already exists — see docker-agent-config for
authoring it, and docker-agent-run for interactive/local invocation.
When to use this skill
Activate this skill when:
- The user wants an agent reachable over MCP, an OpenAI-compatible chat endpoint, a plain HTTP API, or A2A/ACP.
- The user wants to publish an agent to Docker Hub (or any OCI registry) or pull one someone else published.
- The user wants automated evaluations (regression tests) for an agent, or wants to gate CI on eval results.
Do not use this skill when
Do not use this skill when:
- The task is authoring the agent.yaml itself (models, toolsets, sub_agents) — use
docker-agent-config. - The task is running the agent interactively on a developer's machine, choosing
--safety/--sandbox, or aliases — usedocker-agent-run.
Core guidance
Serving an agent
-
Five server modes, each with its own default loopback listen address — never expose any of them beyond loopback without authentication:
Mode Default listen Auth flag Has --safety?serve mcp127.0.0.1:8081--auth-token(only with--http)Yes (only with --http)serve api127.0.0.1:8080--auth-tokenNo serve chat127.0.0.1:8083--api-key/--api-key-envYes serve a2a127.0.0.1:8082--auth-tokenYes serve acp(stdio only) n/a No docker agent serve mcp ./agent.yaml --http --listen 127.0.0.1:9090 --auth-token "$TOKEN" -
serve mcpdefaults to stdio transport (for local clients like Claude Desktop); pass--httponly when you need a network-reachable MCP endpoint, and set--auth-tokenwhenever you do. -
Binding any server flag to a non-loopback address without an auth token/key is refused;
--insecure-no-authexists to force it and must be treated as a deliberate, documented exception, never a default. -
serve mcp(with--http),serve chat, andserve a2aexpose--safety(strict/balanced/restricted/autonomous); Docker's docs state it defaults torestrictedfor these modes when unset.serve apiandserve acpexpose no--safetyflag at all. Never raise--safetytoautonomouson a network-reachable listener; if a served agent must approve more, preferbalancedand keep auth enabled. -
serve apiaccepts a directory instead of a single file: every.yaml/.yml/.hclin it is exposed under/api/agents. Use--session-workingdir-rootto confine session working directories when the server is reachable by more than one user.
Sharing agents via OCI registries
- Push and pull agent configs the same way you push and pull images — same
registry, same
docker loginauth:docker agent share push ./agent.yaml docker.io/username/my-agent:latest docker agent share pull docker.io/username/my-agent:latest instruction_filecontents are inlined into the pushed artifact automatically, so a published agent stays self-contained — you do not need to bundle the referenced files separately.- Pin
sub_agentsthat reference the pushed artifact to a digest (name@sha256:...) once published, to avoid a per-run registry lookup and to guarantee the exact config a consumer gets. - Use
--forceonshare pullonly when you intend to overwrite a local copy that already exists; without it, an existing local config is left untouched.
Evaluating agents
- Evals live in an
evals/directory next to the agent config by default; each eval is one JSON session file capturing a user message, the recorded tool calls, and anevalsobject with the scoring criteria. - Create eval sessions from real conversations rather than hand-writing
JSON: run the agent interactively, then use the
/evalslash command in the TUI to save the session, and edit inrelevance/size/assertionscriteria afterward. - Four scoring dimensions: Tool Calls (F1 against the recorded sequence),
Relevance (LLM-judge,
--judge-model, defaultanthropic/claude-opus-5), Size (S/M/L/XL response-length bucket), and Assertions (deterministic checks; see the complete assertion-type list inreferences/eval-format.md). Prefer assertions overrelevancewhen a check can be exact: they need no judge model and are deterministic, not approximation-prone. - Evaluations run inside containers for isolation; a Docker-compatible
runtime is required. Dedicated provider API keys
(
ANTHROPIC_API_KEY/OPENAI_API_KEY) are forwarded automatically.GITHUB_TOKEN/GH_TOKENare not forwarded automatically (they're broad host credentials, not model keys) — pass them explicitly with-e GITHUB_TOKENwhen an agent's provider needs one (e.g.github-copilot). - Gate CI on regressions, not on absolute scores, with
--baseline:
A previously-passing eval that now fails always gates regardless of tolerance; cost changes are reported but never gate. A baseline or run with zero evaluations (e.g. andocker agent eval ./agent.yaml --baseline results/2026-08-01-run.json --regression-tolerance 0.05--onlypattern matching nothing) is rejected rather than reported as passing. - Use
--keep-containersplus your runtime'sexecto inspect a failed eval's container; the eval's.dbsession file holds the full conversation for offline debugging.
Verify
- After changing a served agent's config, re-run its evals with the same
explicit
--safetyvalue used in the deployment before restarting the listener — this catches an approval-policy regression before it reaches traffic. If a rollout must be rolled back, restore the prior config and safety flag; never restore an unauthenticated listener as a rollback shortcut.
Related skills
- For writing or changing the underlying
agent.yaml, usedocker-agent-config. - For local/interactive runs, safety-mode choice, and sandboxing, use
docker-agent-run.
References
references/eval-format.md— full eval session JSON schema and CLI flag table.references/sources.md— provenance of every rule in this skill.
Assets
assets/eval-session-example.json— a minimal eval session file to copy and adapt.
Checks
checks/verification.md— Verification runbook for serving, sharing, and evaluating an agent.
Files
7- SKILL.md
f52287752b7.6 KB - agents/openai.yaml
5cf06cf791402 B - assets/eval-session-example.json
acefed88061.3 KB - checks/verification.md
4800d620732.3 KB - references/eval-format.md
347be32d472.6 KB - references/sources.md
d7c7809e8d1.6 KB - skill.yaml
449ccdd1d3923 B
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from docker/skills8
Use this skill when creating or editing an agent.yaml (or .yml/.hcl) configuration file for Docker Agent (cagent), including defining agents, models/providers, built-in or MCP toolsets, multi-agent teams with sub_agents. Even if the user just says they want to "build an AI agent with Docker", "make
Use this skill when running a Docker Agent with `docker agent run`, choosing a safety/approval mode, using the `--sandbox` isolation flag, setting up aliases, or troubleshooting a run (missing credentials, worktrees). Even if the user just says they want to "run my agent", "make my agent auto-approv
Use this skill when writing, reviewing, or optimizing Dockerfiles, even if the user just says their image is too large, their build is slow, or they need to harden a container for production. Covers multi-stage builds, layer caching, .dockerignore, non-root users, and image size optimization.
Use this skill when creating, modifying, or debugging Docker Compose configurations, even if the user just says they need to wire services together, add a database to their stack, or set up a local development environment with multiple containers. Covers service definitions, health checks, dependenc
Use this skill before running, or recommending, any Docker command that deletes, wipes, resets, or otherwise irreversibly changes state — even if the user just says to "clean up", "clear the cache", "start fresh", "wipe everything", "nuke it", "reset", "force remove", or "tear down" Docker resources
Use this skill when setting up, initializing, or Dockerizing a project, even if the user doesn't explicitly mention Docker but describes a need for containerized local development, adding a database or cache dependency, or running services without host-level installs. Covers Dockerfile, compose.yaml
Use this skill when authoring, planning, or running a declarative `sbxenv.yaml` file for Docker Sandboxes (`sbx env create/run/plan/exec/rm`), even if the user just says they want to "check in a sandbox config", "make onboarding reproducible for a sandbox", "run a setup script before the agent start
Use this skill when authoring, validating, packaging, signing, or composing a Docker Sandboxes kit `spec.yaml` (`sbx kit add/inspect/pack/pull/push/sign/validate/verify`), even if the user just says they want to "add a tool to a sandbox agent", "build a reusable sandbox extension", "publish a kit to
Related devops skillsscan passed
Run repeated rollouts ("Prime Gauss" style recursive prompting) while keeping an append-only decision ledger of trials, marks, coherence checks, and promotion gates, so recursive confidence never auto-approves live trading, deploy, or destructive actions. Use when the user asks for repeated rollouts
Configure deployment settings for /land-and-deploy.
Build or maintain Cloudflare Sandbox apps on @cloudflare/sandbox@next (SDK 1.0 preview). Use sandbox-migrate-to-next when porting a stable app.
Deploy tRPC on WinterCG-compliant edge runtimes with fetchRequestHandler() from @trpc/server/adapters/fetch. Supports Cloudflare Workers, Deno Deploy, Vercel Edge Runtime, Astro, Remix, SolidStart. FetchCreateContextFnOptions provides req (Request) and resHeaders (Headers) for context creation. The
Automates CI/CD pipeline setup. Use when setting up or modifying build and deployment pipelines. Use when you need to automate quality gates, configure test runners in CI, or establish deployment strategies.
Deploys and manages full-stack web applications (Next.js, Angular) with Server-Side Rendering (SSR) using Firebase App Hosting. Use when deploying Next.js/Angular apps, configuring apphosting.yaml or firebase.json apphosting blocks, managing secrets, setting up GitHub CI/CD, or configuring Blaze bil