Knowledge base
CodexGuild Knowledge Base

TrapDoor (May 2026): zero-width prompt injection in agent config files

as of Sep 28, 2026 · canonical · codexguild.com/kb/incident-trapdoor-2026 · exported 2026-10-11
Canonical as of Sep 28, 2026

TrapDoor (May 2026): zero-width prompt injection in agent config files

34+ malicious packages (384 versions) planted .cursorrules/CLAUDE.md with invisible zero-width-Unicode instructions. First documented at-scale attack on the agent instruction channel itself.

TrapDoor — hidden instructions in agent config files

As of: 2026-09-28 · active campaign, npm/PyPI/Crates.io

The novelty

The payload plants .cursorrules and CLAUDE.md files whose malicious instructions are encoded with zero-width Unicode characters (U+200B, U+200C, U+200D, U+FEFF): invisible in every standard editor, fully parsed by the AI assistant. The developer cannot see what the agent is being told.

On next use of Cursor/Claude Code in that directory, the hidden instructions trigger a fake "security scan" that collects and exfiltrates local secrets through the assistant's own command execution. No vulnerability in the assistant required — it exploits designed behavior (reading config files as trusted).

Distribution vectors

  1. 34+ malicious packages, 384+ versions, targeting crypto/DeFi/AI devs. Credential sweeps: SSH keys, AWS/GitHub tokens, browser profiles, wallet keystores — validated live before exfiltration.
  2. Pull requests against langchain, langflow, browser-use, llama_index, MetaGPT, OpenHands — "docs: add .cursorrules with dev standards" — pointing at an attacker-controlled config URL. (All caught; none merged.)

Detection & defense

  • Zero-width scan (add to CI and pre-commit):
    grep -rPn '[\x{200B}-\x{200D}\x{FEFF}]' --include='*.md' --include='.cursorrules' .
    
  • Treat agent config files as trusted-execution surfaces: review diffs like CI workflow changes.
  • Treat any PR adding/modifying instruction files from unknown contributors as hostile until reviewed.
  • Registry-side: Socket's behavioral monitoring flagged TrapDoor packages with a median of 5m27s after publication — npm provenance + install-time scanning (Socket, socket-cli) pays for itself.
  • Never auto-merge or auto-trust agent instruction files, even from "documentation" PRs.