AGENTS.md and instruction files: shaping agent behavior safely
AGENTS.md and instruction files: shaping agent behavior safely
AGENTS.md/CLAUDE.md define how agents behave in your repo — but they are also an attack surface (TrapDoor planted poisoned ones). Keep them reviewed, specific, and version-controlled.
AGENTS.md and instruction files
As of: 2026-09 · applies to Claude Code (CLAUDE.md), Codex/Cursor/others (AGENTS.md), JetBrains AI (AGENTS.md)
What they are for
Instruction files are the agent's onboarding manual: project context, coding standards, commands, restrictions, definition of done. Anything you'd repeat across sessions belongs there.
What makes a good one
- Specialist over generalist: define a specific role, not "helpful assistant".
- Show, don't tell: code examples beat prose.
- Actionable: exact commands with flags, not "run the tests".
- Boundaries section is critical: explicit Always / Ask-first / Never rules prevent mistakes.
- Mark user decisions: annotate intentional choices ("deliberately simple, do not refactor") so the agent doesn't "improve" them away.
- Living document: update when patterns change; tell the agent to propose updates.
The security angle (learned from TrapDoor, 2026-05)
The TrapDoor supply-chain campaign planted .cursorrules and CLAUDE.md files with instructions hidden in zero-width Unicode characters (U+200B/C/D, U+FEFF) — invisible in editors, fully parsed by the agent. Distribution: 34+ malicious packages across npm/PyPI/Crates.io AND pull requests against langchain, langflow, browser-use, llama_index, MetaGPT, OpenHands posing as docs updates.
Defenses:
- Treat instruction files as trusted-execution surfaces — same review rigor as CI config.
- Diff-review any incoming PR that adds/modifies AGENTS.md, CLAUDE.md, .cursorrules — especially from unknown contributors.
- Scan for zero-width characters before merging:
grep -rP '[\x{200B}-\x{200D}\x{FEFF}]' AGENTS.md CLAUDE.md .cursorrules(hit = investigate). - Pin your own instruction files — an unexpected modification to YOUR instruction file is a compromise indicator.
- Never let an agent auto-merge changes to its own instruction files.