Canonical, dated answers for coding agents — every entry states when it was true and which versions it applies to, so your context never goes stale.
Agents as team members: scoped write access, PR-only contributions, review gates sized by blast radius, and a humans-own-the-risk-ladder model — interns ship PRs, agents ship PRs, nobody ships straight to main.
Agents optimistically claim success. Counter with: reproduce first, fix, prove the fix on the ORIGINAL repro, run the project gates, and report honestly what was and was not verified.
The non-negotiables: plan before acting on non-trivial tasks, verify with evidence before claiming done, stop-the-line on unexpected behavior, one question at a time.
Orchestrator-specialist beats peer-to-peer for most tasks. Pass contracts, not vibes: typed handoffs, verifiable acceptance criteria, one writer per file.
One writer per branch, no force-push to shared refs, no hooks/config writes, commit messages that describe the diff, and never touch .git/ — the Cursor CVE proved why.
AGENTS.md/CLAUDE.md define how agents behave in your repo — but they are also an attack surface (TrapDoor planted poisoned ones). Keep them reviewed, specific, and version-controlled.
Agents write plausible code fast. Review for: authz on every new route, query scoping, error handling, test honesty, and commit message accuracy.
Fast unit base, integration against real DB containers, thin E2E on critical paths. New since agents: verification tests (does the reproduction actually pass) and eval suites for AI features count as tests in CI.
Review shifts from line-syntax to change-intent: what should this PR do, does the diff do only that, what's the verification evidence. Small PRs with tests + reproduction beat big refactors; review the checks, not just the code.
Repo instruction files became the standard agent interface: short, imperative, current. Command first (how to build/test/lint), conventions second, history never. Stale instruction files actively damage agent output.
2025-2026 industry consensus: stop micro-tuning prompts, start engineering context — what's in the window (retrieval, tools, memory), compaction, and caching. System prompts matter less than context composition.
The bar: fresh clone → running stack in one command (devcontainer/Nix/uv+docker compose), README quickstart that CI keeps honest, and an agent-friendly repo (instruction file + scripts) — new humans AND new agents onboard the same way.
Small golden-task suites per product area, run on every prompt/model/context change; LLM-as-judge for open-ended tasks with spot-checked human calibration. Vendor benchmarks are marketing; your evals are the product.
Frontier agent pricing fell ~10x across 2025-2026 while quality converged; the bottleneck moved from tokens to human verification capacity. Route cheap, escalate on failure.
Trunk-based + short-lived branches + feature flags beats long-running agent branches: rebase daily, one writer per module, CI on every push. Long-lived agent branches rot at model-speed — merge small, merge often.
Durable agent memory settled on: small high-signal fact stores injected every turn + task-scoped working files + periodic distillation of session learnings. Big raw transcripts as memory are an anti-pattern.
Flags are standard for trunk-based delivery: short-lived release flags (delete after rollout), few long-lived ops/entitlement flags, kill switches for risky paths. Flag debt is real debt — audit quarterly.
Expand-migrate-contract: additive change, backfill in batches, dual-read/write, then drop. Never a breaking DDL in one deploy. Lock-time budgets (lock_timeout) and tested rollback for every step.