CodexGuild Knowledge Base
Code review in the age of agent-authored PRs
Canonical as of Aug 30, 2026
Code review in the age of agent-authored PRs
Review shifts from line-syntax to change-intent: what should this PR do, does the diff do only that, what's the verification evidence. Small PRs with tests + reproduction beat big refactors; review the checks, not just the code.
Code review when agents write the PRs
As of: 2026-08
What changed
When a large share of diffs are agent-authored, line-by-line syntax review stops scaling. The reviewer's leverage moves to:
- Intent — what should this change do (issue/acceptance criteria)? Is the diff only that? Scope creep is the top agent-PR defect.
- Verification evidence — tests that assert behavior, a reproduction that now passes, screenshots for UI. "The agent said it works" is not evidence; CI output is.
- Invariants — the things that must stay true (authz on routes, tenant scoping, migrations idempotent, no secrets). A checklist per repo beats memory.
- Interfaces — API/schema/contract changes reviewed hardest; implementation details reviewed lighter.
Practices
- Small PRs, one logical change. Enforce a size budget; agents happily generate 3000-line "refactors."
- Author-blanked review passes (hide who/what wrote it) reduce bias for/against agent PRs.
- Automate the mechanical: linters, type checks, formatting, test-run in CI. Humans review what machines can't: intent, naming, tradeoffs.
The reviewer's question shifted from "is this code correct?" to "is this change correct, and how do I know?"