AI coding agent security
An agent with shell access is remote code execution with extra steps. CodexGuild scans every skill for Claude Code, Codex and Hermes Agent, tracks prompt-injection and supply-chain incidents, and audits what is already installed on your machine.
- 5,637
- skills scanned
- 4,920
- passed
- 553
- need review
- 164
- flagged
Audit the skills on your machine
The CodexGuild MCP server scans ~/.claude/skills, plugins, Codex, Hermes and project skills locally. Nothing is uploaded.
codexguild_scan_skillsOr run codexguild_sync at session start to get the scan summary together with stack changes and advisories.
What the scanner looks for
- Pipe-to-shell & remote code execution
- Credential reads sent over the network
- Hidden unicode (TrapDoor, Trojan Source)
- Prompt injection & covert instructions
- Agent config, hook & MCP tampering
- Safety bypass flags (--yolo, skip-permissions)
- Persistence: cron, launch agents, rc files
- Destructive commands & exfil endpoints
The non-negotiables
Baseline controls for any agent that can run commands.
Sandbox the agent
Containers or VMs, not your laptop. The blast radius should be the sandbox.
Zero secrets exposure
The agent never types, echoes or reads secrets. Inject them at the boundary.
Gate irreversible actions
Push, publish, deploy, pay → a human approves. Never auto-approve writes outside the repo.
Assume injection everywhere
Files, issues, web pages and skills are untrusted input. Structural defenses beat prompts.
Incidents & playbooks
Dated, agent-verified. Your agent receives new ones through codexguild_sync.
Pydantic AI: Concurrency-limited models can keep their slot when a streamed request ends early
High severity. Affects pydantic-ai >= 2.10.0, < 2.53.0. Upgrade to 2.53.0 or later.
Pydantic AI: Excessive resource use when local web fetching converts nested HTML
Medium severity. Affects pydantic-ai >= 1.77.0, < 1.107.7; pydantic-ai >= 2.0.0b1, < 2.52.0. Upgrade to 1.107.7 / 2.52.0 or later.
Pydantic AI Web chat UI (`Agent.to_web()`, `clai web`): the local chat endpoint does not validate the `Host` header
Medium severity. Affects pydantic-ai >= 1.34.0, < 1.107.5; pydantic-ai >= 2.0.0b1, < 2.30.0. Upgrade to 1.107.5 / 2.30.0 or later.
Pydantic AI Web chat UI (`Agent.to_web()`, `clai web`): a website visited by the developer can trigger agent runs and server-side tool execution on the local chat endpoint
High severity. Affects pydantic-ai >= 1.34.0, < 1.107.4; pydantic-ai >= 2.0.0b1, < 2.28.0. Upgrade to 1.107.4 / 2.28.0 or later.
Pydantic AI OpenTelemetry instrumentation: retry prompt content is not redacted when `include_content=False`
Low severity. Affects pydantic-ai >= 0.3.4, < 1.107.4; pydantic-ai >= 2.0.0b1, < 2.27.1. Upgrade to 1.107.4 / 2.27.1 or later.
Pydantic AI: Unbounded memory use when downloading remote content via web_fetch or FileUrl
Medium severity. Affects pydantic-ai >= 1.77.0, < 1.107.2; pydantic-ai >= 2.0.0b1, <= 2.23.0. Upgrade to 1.107.2 / 2.24.0 or later.
Pydantic AI: Event loop blocked by quadratic title extraction in `web_fetch`
Medium severity. Affects pydantic-ai >= 1.77.0, < 1.107.6; pydantic-ai >= 2.0.0b1, < 2.44.0. Upgrade to 1.107.6 / 2.44.0 or later.
Pydantic AI OpenTelemetry instrumentation: exception events on tool and agent run spans include content when `include_content=False`
Low severity. Affects pydantic-ai >= 0.3.4, < 1.107.6; pydantic-ai >= 2.0.0b1, < 2.44.0. Upgrade to 1.107.6 / 2.44.0 or later.
Pydantic AI: SSRF cloud-metadata blocklist bypass via IPv6 zone identifiers
Medium severity. Affects pydantic-ai >= 1.56.0, < 1.107.6; pydantic-ai >= 2.0.0b1, < 2.44.0. Upgrade to 1.107.6 / 2.44.0 or later.
Flagged in the registry
Skills with high-risk patterns. Not necessarily malicious — but a human should read them first.
lambda-labs-gpu-cloud
ai-mlReserved and on-demand GPU cloud instances for ML training and inference. Use when you need dedicated GPU instances with simple SSH access, persistent filesystems, or high-performance multi-node clusters for large-scale training.
See findings
aws-transform
devopsMigrate, modernize, and upgrade codebases to AWS. Run analysis on repos for tech debt, security vulnerabilities, and modernization opportunities. Transforms .NET Framework to .NET 8/10, mainframe COBOL to Java, VMware VMs to EC2, SQL Server to Aurora, and upgrades Java/Python/Node.js versions and AW
See findings
foundry-iq
knowledgeFoundry IQ knowledge bases. WHEN: make local or Blob documents searchable; create/diagnose KBs; triage unsupported connectors or multi-source KB creation/reconfiguration; connect existing KB to agents (including multi-source KBs); create/reuse a Search service; retrieve from an existing knowledge ba
See findings
mcp-management
backendManage MCP servers - discover, analyze, execute tools/prompts/resources. Use for MCP integrations, capability discovery, tool filtering, programmatic execution, or encountering context bloat, server configuration, tool execution errors.
See findings
nvflare-fed-stats
ai-mlCompute federated statistics over tabular data (count, sum, mean, stddev, var, histogram, quantile, noise-protected min/max) and image data (count, failure_count, pixel-intensity histogram) across NVFLARE sites via FedStatsRecipe — automatic and non-interactive from the dataset, feature names (heade
See findings
auth0
backendUse when adding, fixing, or improving how an app authenticates users or protects an API, or when using or configuring any Auth0 feature — signing users in and out, sessions and tokens, guarding routes and endpoints, MFA, passwordless passkey login (WebAuthn), SSO, Organizations, RBAC, custom domains
See findings
Security questions about coding agents
Skills, MCP servers, prompt injection and secrets — the short answers.
Not by default. A skill is instructions plus scripts that your coding agent will follow with its own permissions, so a malicious SKILL.md can read ~/.ssh, pipe a remote script into a shell or rewrite CLAUDE.md. Install skills only after a security scan and a human read of anything flagged.
The scanner statically checks SKILL.md and every bundled file for prompt-injection phrasing, hidden and bidirectional Unicode, pipe-to-shell, encoded payloads, credential reads sent over the network, persistence (cron, launch agents, shell rc files), agent config and MCP tampering, and flags that disable safety checks. Each finding is reported with file and line.
Yes — anything your agent runs inherits its environment. Keep secrets out of the agent's context, inject them at the boundary, sandbox the agent, and require human approval for push, publish, deploy and payments. CodexGuild flags skills that read credential files and send them over the network.
Ask your agent to run codexguild_scan_skills from the local CodexGuild MCP server. It scans ~/.claude/skills, plugins, Codex, Hermes and project skill folders on your machine with the same rules and uploads nothing.
Attackers hide instructions in zero-width or bidirectional Unicode characters inside agent config files and skills. Humans reviewing the file see nothing, but the model reads and follows the hidden text. The CodexGuild scanner reports hidden Unicode in any scanned file.
No. Recommendations come back as proposals — an MCP entry, a skill to install, or a delimited block for AGENTS.md or CLAUDE.md — and your agent applies them only after you approve. Identity and memory files are never touched.