ponytail-gain
Show ponytail's measured savings (code, cost, speed) from the benchmark. One-shot display. Use for /ponytail-gain, "what does ponytail save", "ponytail impact".
- 0
- Installs
- —
- Rating
- —
- Success rate
- 1
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 c5ac2a5e5222a678… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Ponytail Gain
Display this scoreboard when invoked. One-shot: do NOT change mode, write flag files, or persist anything.
The figures are the published agentic benchmark of Ponytail 5: headless Claude
Code (Opus 5.5, default effort) on 39 tasks (feature tickets in a real FastAPI +
React repo, bug fixes, security and privacy cases, small apps), 5 runs each,
against the same agent without the skill. 18 of the tasks have hidden checks
for correctness and safety. Each figure is the geometric mean of the per-task
medians. They are measured, not computed from the current repo.
Source: benchmarks/results/2026-10-07-agentic.md and the README.
Scoreboard
Render plain ASCII bars. The bar length shows ponytail as a share of the no-skill baseline; the label carries the exact figure:
ponytail gain benchmark · 39 tasks × 5 runs · Opus 5.5
no-skill ████████████████████ 100%
Lines of code █████████··········· 47% ▼ 53%
Output tokens ███████████········· 55% ▼ 45%
Cost ███████████████····· 74% ▼ 26%
Time ████████████········ 59% ▼ 41%
Hidden checks passed 97% (no-skill 96%)
Tests where the logic needs one 98% (no-skill 68%)
This repo: /ponytail-debt (shortcuts you deferred)
/ponytail-audit (what's still cuttable)
Honesty boundary
These are benchmark averages, not this repo. NEVER print a per-repo savings
number ("you saved X lines/tokens here"): the unbuilt version was never
written, so there is no real baseline to subtract from in a live repo. The
only real per-repo figures come from /ponytail-debt (a counted ledger), and
this card points there instead of inventing one.
Boundaries
One-shot display. Edits nothing, changes no mode. "stop ponytail" or "normal mode": revert.
Files
1- SKILL.md
60e05eaf652.2 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from DietrichGebert/ponytail5
Lazy senior dev mode: the smallest change that fully solves the task, and a reply a busy human understands in one read. Use on any coding task (writing, fixing, refactoring, reviewing, choosing dependencies) and when the user says "ponytail", "be lazy", "simplest solution", "yagni", or complains abo
Quality audit of a whole repo: bugs, security holes, what breaks under real load, risky code without tests, slow paths, and what to delete, merge or split. Ranked, each finding explained in plain English. One-shot report, changes nothing. Use for "audit this codebase", "review the whole repo", "find
List every `shortcut:` comment (and older `ponytail:` ones) as a debt ledger. One-shot report, changes nothing. Use for "ponytail debt", "what did ponytail defer", "list the shortcuts", /ponytail-debt.
Quick reference for ponytail levels, skills and commands. One-shot display. Use for /ponytail-help, "ponytail help", "how do I use ponytail".
Quality review of a change: is the logic right, is it safe, does it hold under real load, is risky code tested, is it fast enough, and is every line needed. Reads the connected code, not only the diff. Each finding is explained in plain English. Use for "review this", "code review", "review the last