brownfield-retrospective
Run the ECC engineering process retroactively against an existing project slice. Audit current state, reconstruct past decisions, produce a four-axis gap register (thinking, coding practice, missing features, technology selection), then execute fixes forward through the plan-TDD-review pipeline. Use
- 0
- Installs
- —
- Rating
- —
- Success rate
- 1
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 d6e0c8701050bf14… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Brownfield Retrospective
Run ECC's engineering process retroactively against a project that was built without it. Pick any slice of an existing codebase, audit what exists, reconstruct how it got there, extract a four-axis gap register, then execute the fixes forward through the standard ECC pipeline.
The workflow is slice-first: no whole-repo onboarding is required before it is useful, and every phase is re-enterable from a durable artifact on disk. Start in the middle, work backwards, then move forward.
When to Activate
- Adopting ECC in a repository that already exists and is already shaped
- User asks to "audit my existing project", "retro this codebase", or "what did we miss?"
- Taking over, inheriting, or returning to a codebase after a long gap
- User asks "why is it like this?" followed by "should it stay like this?"
- A module, feature, or service needs a rethink before the next feature lands on top of it
- Deciding whether a legacy area should be fixed forward, rewritten, or left documented as-is
When Not to Use
- Greenfield work with no existing code — use
plannerandtdd-workflowdirectly - Onboarding to spec-driven development without the redesign loop — the
spec-mineragent does that on its own - A launch go/no-go production-readiness pass — use
production-audit - Security-only sweeps — use
security-review - Team process retrospectives (sprint ceremonies) — this audits a codebase, not people
The Working Set
All artifacts are written inside the audited project, under docs/retro/<slice>/:
| File | Phase | Purpose |
|---|---|---|
gap-register.md | 0-6 | The state machine — perimeter, Current phase: marker, and rows. The only file required to resume mid-stream. |
baseline-audit.md | 1 | Reviewer findings, test/CI snapshot, TODO inventory. |
reconstructed-spec.md | 2 | What the code actually does today (Requirements + Invariants). |
target-design.md | 4 | Target shape for the slice and the fix sequence. |
Backfilled decisions go to docs/adr/ per the architecture-decision-records skill. If the project uses OpenSpec, reconstructed-spec.md may instead be openspec/specs/<capability>/spec.md produced by spec-miner.
How It Works
Phase 0 — Pick the Entry Point
Start from what the user cares about today. Do not map the whole repo first.
Any of these is a valid slice:
| User signal | Slice |
|---|---|
| Names a directory, module, or service | That path plus its direct callers |
| Names a feature or user flow | The files that implement it, traced from the entry point |
| Names a bug, incident, or pain point | The code that produced it plus one level of call chain |
| Names a dependency or tech choice | Everything that touches it |
| Names nothing | The 2-3 hottest paths by churn — propose them and let the user pick |
Record the perimeter at the top of gap-register.md before anything else: paths in scope, entry points, callers, blast radius. Budget the slice at roughly 15 files or 2,000 lines; if the user's pick is bigger, cut it into named sub-slices and register each as its own row of work. A slice too big to audit in one session is a slice that never finishes. Keep a Current phase: marker beside the perimeter and advance it each time a phase completes.
Phase 1 — Present-State Audit
Establish what exists now, with evidence. Do not fix anything yet.
- Recon the slice: manifest and framework fingerprint, entry points, test layout, CI checks (borrow the recon steps from
codebase-onboarding, scoped to the perimeter). - Run the reviewers that match the stack, in parallel where possible. Every finding must carry a
file:linepointer:- The language reviewer for the stack (
typescript-reviewer,python-reviewer,go-reviewer, ...) security-reviewersilent-failure-hunterfor swallowed errorstype-design-analyzerwhere invariants are encoded in typesperformance-optimizeronly if the slice is on a hot path
- The language reviewer for the stack (
- Snapshot the current state: test count and coverage, lint/CI status,
TODO/FIXME/HACKinventory inside the perimeter.
Write everything to baseline-audit.md. Findings that get fixed during this phase corrupt the baseline — hold them for the register.
Phase 2 — Backward Reconstruction
Work backwards from the present state to the decisions that produced it.
- Mine the behavioral spec: run
spec-mineragainst the slice. Its Requirements (WHEN → THEN) and Invariants becomereconstructed-spec.md— the baseline truth for Phase 5 tests. - Git archaeology on the perimeter:
git log --oneline -- <slice>for directory history; use--followonly with a single file path — age, authors, cadence- Largest and most reverted commits — where the design changed under pressure
- Churn hotspots — files changed far more often than neighbors
- Dead ends — branches or scaffolding that never shipped
- Decision inventory: list the significant decisions frozen into the code (stack and framework choices, data model, boundaries, error strategy, state management). Check each against
docs/adr/, README, and comments. Every decision with no written record is itself a pre-existing thinking gap — record it in Phase 3, do not silently backfill it here. - Where behavior is opaque, use
code-explorerto trace execution paths before guessing at intent.
Phase 3 — Gap Analysis (Four Axes)
Convert the audit and reconstruction into register rows. One axis per class of gap:
| Axis | Question | Sources |
|---|---|---|
| Thinking | What was decided but never recorded? Which invariants live only in code? | Decision inventory, reconstructed spec |
| Coding practice | Where does the code diverge from how ECC would have written it? | Phase 1 findings, the rules/ pack for the language, spec items with no test anchor |
| Missing features | What would a day-one build have included that this slice never got? | Boundary input validation, observability, rate limiting, idempotency, migrations plus rollback, a11y, i18n, user docs, CI gates |
| Technology selection | Is the chosen tech fighting the codebase? Outdated, duplicated, risky, or missing? | Dependency inventory, git archaeology, recon |
Register format (the table is the skill's durable output):
| ID | Axis | Gap | Evidence | Forward fix | Severity | Effort | Status |
|----|------|-----|----------|-------------|----------|--------|--------|
| G-001 | practice | Errors swallowed in retry loop | src/api/retry.ts:42 | Propagate + map at boundary | P1 | S | open |
- Severity:
P0security exposure or data loss,P1correctness,P2practice and observability,P3selection and cleanup. - Effort:
Sunder a session,Ma few sessions,Lneeds its own plan. - Status:
open,in-progress,done,wont-fix. - Every row needs an evidence pointer. A gap you cannot point at is an opinion.
Phase 4 — Redesign / Rethink
Turn the register into a sequenced plan. Follow the project's existing task and approval authority. Keep the register as audit evidence linked to that system; do not create competing canonical task state. Confirm scope and obtain project-required approval before a rewrite, technology replacement, migration, or other material change.
- Run
architectorcode-architectover theP0andP1rows to producetarget-design.md: target shape for the slice, the fix sequence, and an explicit "what does not change" list to protect working behavior. - Backfill retro-ADRs for every undocumented decision from Phase 2. Mark them as reconstructed (
status: proposed (reconstructed), date of the retro, note the original decision predates the record). Separate observed history from inferred intent, and have the project owner confirm a decision before marking it accepted. Never present a backfilled ADR as if it were written at decision time. - Triage every register row into one of:
fix-forward(repair in place),rewrite(replace the slice),leave-documented(wont-fix). Awont-fixmust record its reason in the register — a wont-fix with a reason closes the thinking gap; a silent one reopens it. - Sequence the backlog:
P0first, then interleave quick wins with larger repairs to keep momentum.
Phase 5 — Forward Execution
Work the register through the standard ECC pipeline, one row per session as a baseline pace:
planner (scope the row)
→ tdd-guide (failing test first — for a legacy gap the first test
usually encodes the missing invariant from reconstructed-spec.md)
→ implement
→ code-reviewer
→ security-reviewer (if the row touches input, auth, or data)
→ doc-updater + ADR (if the decision changed)
→ conventional commit
→ mark the register row done
- Update
gap-register.mdstatuses as rows close; the register is the durable state, not the conversation. - On session boundaries, hand off through the project's memory mechanism (
unified-memory, context carriers) so the next session resumes from the register alone.
Phase 6 — Closure and Re-Audit
When every row is done or wont-fix:
- Re-run the Phase 1 audit over the same perimeter.
- Diff against
baseline-audit.mdand verify withverification-loop— the gaps claimed closed must stay closed. - Record the delta (findings then, findings now, rows closed, rows deferred) at the bottom of
gap-register.md. - Anything that resurfaced becomes a new row, not a failure — the register is a loop, not a one-shot.
Resuming Mid-Stream
Any session can start in the middle. Read docs/retro/<slice>/gap-register.md first:
- Register exists with a row
in-progress→ Phase 5, finish that row. - Register exists with at least one row, all rows closed → Phase 6.
- Register exists, rows open, none in progress → resume at the
Current phase:marker (Phase 3 register awaiting triage, or Phase 4 backlog awaiting its first row). - No register, but
baseline-audit.mdexists → Phase 2. - No artifacts at all → Phase 0 with the user.
Anti-Patterns
- Auditing the whole repo before any slice — the workflow is slice-first by design.
- Fixing code during Phases 1-3 — audit phases produce the register; edits mid-audit corrupt the baseline.
- Register rows without evidence pointers.
wont-fixwithout a recorded reason.- Presenting backfilled ADRs as contemporaneous records.
- Skipping the Phase 6 re-audit — closure is verified, not declared.
Related Skills
codebase-onboarding— recon baseline for Phase 1architecture-decision-records— retro-ADR backfill in Phase 4tdd-workflow— Phase 5 execution disciplineverification-loop— Phase 6 closureproduction-audit— launch-readiness lens (complementary, later)security-review— security-only requests
Files
1- SKILL.md
805439bc7911.2 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from affaan-m/everything-claude-code8
Design, implement, and audit accessible UI to WCAG 2.2 Level AA across Web, iOS, and Android — semantic ARIA roles and labels, accessibility traits and hints, focus management, contrast, target size, and screen-reader support. Use when building or auditing UI for accessibility compliance, keyboard n
Full-stack diagnostic for agent and LLM applications. Audits the 12-layer agent stack for wrapper regression, memory pollution, tool discipline failures, hidden repair loops, and rendering corruption. Produces severity-ranked findings with code-first fixes. Essential for developers building agent ap
Head-to-head comparison of coding agents (Claude Code, Aider, Codex, etc.) on custom tasks with pass rate, cost, time, and consistency metrics. Use when choosing between coding agents, or when a change to an agent setup needs measured pass rate, cost, and time rather than an impression.
Design and optimize AI agent action spaces, tool definitions, and observation formatting for higher completion rates. Use when defining or revising an agent's tool set, action space, or observation format.
Structured self-debugging workflow for AI agent failures using capture, diagnosis, contained recovery, and introspection reports. Use when an agent run fails and you need a reproducible diagnosis instead of a retry.
Add x402 payment execution to AI agents with per-task budgets, spending controls, and non-custodial wallets. Supports Base through agentwallet-sdk, X Layer through OKX Payments / OKX Agent Payments Protocol, and Solana plus multi-network EVM through the upstream x402 packages with facilitator-based
Verify a local agent API, temporary gateway tunnel, and remote sandbox callback with a tool-free task, then restore the original app connection.
Security hardening guidance for AI agent frameworks that process untrusted content, invoke tools, write workspace files, manage runtime identifiers, or handle credentials. Use when building or reviewing an agent runtime, autonomous worker, tool gateway, memory service, or multi-tenant agent deployme
Related methodology skillsscan passed
Use when asked to review a plan, challenge a proposal, run a CEO review, poke holes in an approach, think bigger about scope, or decide whether to expand or reduce the plan.
This skill should be used when the user wants to "create a skill", "add a skill to plugin", "write a new skill", "improve skill description", "organize skill content", or needs guidance on skill structure, progressive disclosure, or skill development best practices for Claude Code plugins.
Reviews a branch, pull request or uncommitted change for the interface problems it introduced or regressed, across accessibility, layout, writing, typography, color and UI.
Guides Stripe integration decisions across development and test environment planning (separate sandboxes vs the shared test mode sandbox), API selection (Checkout Sessions vs PaymentIntents), Connect platform setup (Accounts v2, controller properties), billing/subscriptions, tax and registrations (S
Cloudflare Workers best practices for production applications. Use when writing, reviewing, or configuring Workers.
Generate and evaluate grounded ideas. Use when the user wants ideas, improvements, or surprising directions before choosing one to develop. Not for refining an idea they already have (ce-brainstorm) or judging one already on the table (ce-pov).