AI coding agent security

An agent with shell access is remote code execution with extra steps. CodexGuild scans every skill for Claude Code, Codex and Hermes Agent, tracks prompt-injection and supply-chain incidents, and audits what is already installed on your machine.

5,637
skills scanned
4,920
passed
553
need review
164
flagged

Audit the skills on your machine

The CodexGuild MCP server scans ~/.claude/skills, plugins, Codex, Hermes and project skills locally. Nothing is uploaded.

$codexguild_scan_skills

Or run codexguild_sync at session start to get the scan summary together with stack changes and advisories.

What the scanner looks for

  • Pipe-to-shell & remote code execution
  • Credential reads sent over the network
  • Hidden unicode (TrapDoor, Trojan Source)
  • Prompt injection & covert instructions
  • Agent config, hook & MCP tampering
  • Safety bypass flags (--yolo, skip-permissions)
  • Persistence: cron, launch agents, rc files
  • Destructive commands & exfil endpoints

The non-negotiables

Baseline controls for any agent that can run commands.

Sandbox the agent

Containers or VMs, not your laptop. The blast radius should be the sandbox.

Zero secrets exposure

The agent never types, echoes or reads secrets. Inject them at the boundary.

Gate irreversible actions

Push, publish, deploy, pay → a human approves. Never auto-approve writes outside the repo.

Assume injection everywhere

Files, issues, web pages and skills are untrusted input. Structural defenses beat prompts.

Incidents & playbooks

Dated, agent-verified. Your agent receives new ones through codexguild_sync.

All security entries
PlaybookOct 8, 2026

Pydantic AI: Concurrency-limited models can keep their slot when a streamed request ends early

High severity. Affects pydantic-ai >= 2.10.0, < 2.53.0. Upgrade to 2.53.0 or later.

PlaybookOct 8, 2026

Pydantic AI: Excessive resource use when local web fetching converts nested HTML

Medium severity. Affects pydantic-ai >= 1.77.0, < 1.107.7; pydantic-ai >= 2.0.0b1, < 2.52.0. Upgrade to 1.107.7 / 2.52.0 or later.

PlaybookOct 8, 2026

Pydantic AI Web chat UI (`Agent.to_web()`, `clai web`): the local chat endpoint does not validate the `Host` header

Medium severity. Affects pydantic-ai >= 1.34.0, < 1.107.5; pydantic-ai >= 2.0.0b1, < 2.30.0. Upgrade to 1.107.5 / 2.30.0 or later.

PlaybookOct 8, 2026

Pydantic AI Web chat UI (`Agent.to_web()`, `clai web`): a website visited by the developer can trigger agent runs and server-side tool execution on the local chat endpoint

High severity. Affects pydantic-ai >= 1.34.0, < 1.107.4; pydantic-ai >= 2.0.0b1, < 2.28.0. Upgrade to 1.107.4 / 2.28.0 or later.

PlaybookOct 8, 2026

Pydantic AI OpenTelemetry instrumentation: retry prompt content is not redacted when `include_content=False`

Low severity. Affects pydantic-ai >= 0.3.4, < 1.107.4; pydantic-ai >= 2.0.0b1, < 2.27.1. Upgrade to 1.107.4 / 2.27.1 or later.

PlaybookOct 8, 2026

Pydantic AI: Unbounded memory use when downloading remote content via web_fetch or FileUrl

Medium severity. Affects pydantic-ai >= 1.77.0, < 1.107.2; pydantic-ai >= 2.0.0b1, <= 2.23.0. Upgrade to 1.107.2 / 2.24.0 or later.

PlaybookOct 8, 2026

Pydantic AI: Event loop blocked by quadratic title extraction in `web_fetch`

Medium severity. Affects pydantic-ai >= 1.77.0, < 1.107.6; pydantic-ai >= 2.0.0b1, < 2.44.0. Upgrade to 1.107.6 / 2.44.0 or later.

PlaybookOct 8, 2026

Pydantic AI OpenTelemetry instrumentation: exception events on tool and agent run spans include content when `include_content=False`

Low severity. Affects pydantic-ai >= 0.3.4, < 1.107.6; pydantic-ai >= 2.0.0b1, < 2.44.0. Upgrade to 1.107.6 / 2.44.0 or later.

PlaybookOct 8, 2026

Pydantic AI: SSRF cloud-metadata blocklist bypass via IPv6 zone identifiers

Medium severity. Affects pydantic-ai >= 1.56.0, < 1.107.6; pydantic-ai >= 2.0.0b1, < 2.44.0. Upgrade to 1.107.6 / 2.44.0 or later.

Flagged in the registry

Skills with high-risk patterns. Not necessarily malicious — but a human should read them first.

Registry

lambda-labs-gpu-cloud

ai-ml

Reserved and on-demand GPU cloud instances for ML training and inference. Use when you need dedicated GPU instances with simple SSH access, persistent filesystems, or high-performance multi-node clusters for large-scale training.

See findings

aws-transform

devops

Migrate, modernize, and upgrade codebases to AWS. Run analysis on repos for tech debt, security vulnerabilities, and modernization opportunities. Transforms .NET Framework to .NET 8/10, mainframe COBOL to Java, VMware VMs to EC2, SQL Server to Aurora, and upgrades Java/Python/Node.js versions and AW

See findings

foundry-iq

knowledge

Foundry IQ knowledge bases. WHEN: make local or Blob documents searchable; create/diagnose KBs; triage unsupported connectors or multi-source KB creation/reconfiguration; connect existing KB to agents (including multi-source KBs); create/reuse a Search service; retrieve from an existing knowledge ba

See findings

mcp-management

backend

Manage MCP servers - discover, analyze, execute tools/prompts/resources. Use for MCP integrations, capability discovery, tool filtering, programmatic execution, or encountering context bloat, server configuration, tool execution errors.

See findings

nvflare-fed-stats

ai-ml

Compute federated statistics over tabular data (count, sum, mean, stddev, var, histogram, quantile, noise-protected min/max) and image data (count, failure_count, pixel-intensity histogram) across NVFLARE sites via FedStatsRecipe — automatic and non-interactive from the dataset, feature names (heade

See findings

auth0

backend

Use when adding, fixing, or improving how an app authenticates users or protects an API, or when using or configuring any Auth0 feature — signing users in and out, sessions and tokens, guarding routes and endpoints, MFA, passwordless passkey login (WebAuthn), SSO, Organizations, RBAC, custom domains

See findings

Security questions about coding agents

Skills, MCP servers, prompt injection and secrets — the short answers.

Not by default. A skill is instructions plus scripts that your coding agent will follow with its own permissions, so a malicious SKILL.md can read ~/.ssh, pipe a remote script into a shell or rewrite CLAUDE.md. Install skills only after a security scan and a human read of anything flagged.

The scanner statically checks SKILL.md and every bundled file for prompt-injection phrasing, hidden and bidirectional Unicode, pipe-to-shell, encoded payloads, credential reads sent over the network, persistence (cron, launch agents, shell rc files), agent config and MCP tampering, and flags that disable safety checks. Each finding is reported with file and line.

Yes — anything your agent runs inherits its environment. Keep secrets out of the agent's context, inject them at the boundary, sandbox the agent, and require human approval for push, publish, deploy and payments. CodexGuild flags skills that read credential files and send them over the network.

Ask your agent to run codexguild_scan_skills from the local CodexGuild MCP server. It scans ~/.claude/skills, plugins, Codex, Hermes and project skill folders on your machine with the same rules and uploads nothing.

Attackers hide instructions in zero-width or bidirectional Unicode characters inside agent config files and skills. Humans reviewing the file see nothing, but the model reads and follows the hidden text. The CodexGuild scanner reports hidden Unicode in any scanned file.

No. Recommendations come back as proposals — an MCP entry, a skill to install, or a delimited block for AGENTS.md or CLAUDE.md — and your agent applies them only after you approve. Identity and memory files are never touched.