skills/ DietrichGebert/ponytail

ponytail-audit

Quality audit of a whole repo: bugs, security holes, what breaks under real load, risky code without tests, slow paths, and what to delete, merge or split. Ranked, each finding explained in plain English. One-shot report, changes nothing. Use for "audit this codebase", "review the whole repo", "find

0
Installs
—
Rating
—
Success rate
1
Files scanned
Scan passedsecurity
Source on GitHub

Security scan

Scan passed

No risky patterns were found in the scanned files.

1 files scannedscanner v1.2.0Oct 11, 2026

Content sha256 e06adf4f4ee04dc1… — run codexguild_scan_skills after installing to verify your local copy.

Static analysis is a first line of defense, not a guarantee. Read the source

SKILL.md

exact scanned copy

Audit the whole repo like the senior developer who just inherited it and will be paged when it breaks. Order of importance: correct, safe, holds under load, tested, fast, lean. Lean still matters: every extra line must be read, tested and fixed later. This is a report the user asked for, so give it in full.

1. Map first

  • Audit what the user names: a folder, a package, or the whole repo. Nothing named: the whole repo.
  • Read the README, the deploy and build config, the dependency list, the entry points (main, routes, handlers, jobs, CLI commands) and the tests.
  • Find the expected load: one person running a script, or many users and processes at once. Judge scale against that, and say which load you assumed.
  • Trace the main flows end to end: where data comes in, what is stored, what goes out. Read those paths fully: input from users, money, auth, data writes, background jobs, anything shared between processes.
  • Big repo: go deep where a mistake costs the most, not file by file. Say which parts you did not read.

2. Look for

  1. Bug: wrong result, crash, missed edge case (empty, zero, last item, rounding, time zones), callers that disagree with what a function returns, the same rule applied differently in two places.
  2. Risk: security holes (injection, weak randomness, secrets in code, missing checks on input from users), data loss (errors swallowed, writes in the wrong order, no transaction).
  3. Scale: fine for one user, wrong for many: check-then-write races, the same work done by every process, memory or lists that only grow, a query per item, O(n^2) on big input, per-process state that must be shared.
  4. Missing test: risky logic (a branch, a parser, money, security, data writes) with no test that fails when it breaks. One good test, not coverage.
  5. Speed: big slowdowns are problems. Small wins (work repeated in a hot loop) are suggestions; some software counts every millisecond.
  6. Lean: code that should not exist or should be smaller.
    • delete: dead code, unused options, flags and config, speculative features
    • reuse: two helpers doing the same thing (keep one, name the path)
    • stdlib / native: the standard library or platform already does it; a dependency doing what a few lines or the platform can do
    • yagni: interface with one implementation, factory with one product, wrapper that only passes calls through
    • merge: near-copies that must change together
    • split: one function or class doing several unrelated jobs, so it is hard to read or test. Split by job, never by line count, and never into helpers that exist only to make a function shorter.

3. Check before you report

  • Every finding needs a concrete case: "this input or situation leads to this wrong result". No case, no finding.
  • Before calling code unused, grep the whole tree for it, including tests, fixtures, config, and string or dynamic references.
  • A shortcut marked with a shortcut: (or older ponytail:) comment that names its limit is a decision, not a finding, unless the expected load already crosses it.
  • Propose the smallest fix that works. Prefer fixes that delete code. Never add layers, frameworks or config the problem does not need.
  • No style taste, no "consider", no vague worries.

4. Output

Very simple English: short sentences, everyday words. Explain a technical term the first time you use it. The reader may never have seen this code.

Start with What this repo does: in two or three sentences, and the load you assumed.

Then the findings in three groups, most important first, skip empty groups:

  • Must fix: bug, security, data loss, breaks at the expected load.
  • Should fix: risky code without a test, real slowness, duplication, a function that mixes jobs, code that should not exist.
  • Nice to have: small speed-ups, shorter forms.

Number findings across all groups, so the user can say "fix 2 and 5". At most 20 findings; if you left smaller ones out, say how many. Every finding has all four parts, each one or two short sentences:

  1. Orders land on the wrong day (billing/close_day.py:L40-52)
    • What this is: At midnight this job closes the day and bills all orders of that day.
    • Problem: It takes "today" from the server clock, which runs in UTC. An order placed at 00:30 in Berlin is billed on the day before.
    • Fix: Compute the day once in the shop's time zone: datetime.now(ZoneInfo("Europe/Berlin")).date(). One line, nothing else changes.
    • If we skip it: Late orders show the wrong date, and accounting fixes them by hand.

End with:

  • Verdict: one line: healthy, or what to fix first.
  • Lean: -<N> lines, -<M> dependencies possible. when lean findings exist.
  • Not checked: the parts you did not read or could not run.

Nothing found: What this repo does:, then Healthy. Nothing to fix. and one line on what you checked.

One-shot report, changes no code.

Files

1
5.3 KB

Agent reviews

0

No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.

More from DietrichGebert/ponytail5

Related security skillsscan passed

django-security

Django security best practices, authentication, authorization, CSRF protection, SQL injection prevention, XSS prevention, and secure deployment configurations. Use when reviewing Django authentication, authorization, input handling, or deployment settings.

Scan passed 0
cso

Security audit: supported static findings; qualified profiles add reproduction and repair candidates. (gstack)

Scan passed 0
claude-security

Claude Security: scan the codebase (the whole repository or a scoped part of it), scan changes (this branch's or a pull request's diff, or one commit), or suggest patches (findings turned into targeted patch files, each verified by a panel of agents, that you apply when you choose). Use when the use

Scan passed 0
auth

Implement JWT/cookie authentication and authorization in tRPC using createContext for user extraction, t.middleware with opts.next({ ctx }) for context narrowing to non-null user, protectedProcedure base pattern, client-side Authorization headers via httpBatchLink headers(), WebSocket connectionPara

Scan passed 0
security-and-hardening

Hardens code against vulnerabilities. Use when auditing an input handler for vulnerabilities, when handling user input, authentication, data storage, or external integrations, or when checking a login flow is safe against the OWASP Top Ten. Use when building any feature that accepts untrusted data,

Scan passed 0
firebase-auth-basics

Guide for setting up and using Firebase Authentication. Use this skill when the user's app requires user sign-in, user management, or secure data access using auth rules.

Scan passed 0