Labsco
Labsco/Repos/dietrichgebert/ponytail
REPO PACKAGE

dietrichgebert ponytail

COMMUNITY
dietrichgebert · publisher75,041 repository starsMIT · Freegithub.com/dietrichgebert/ponytail
6skills
3groups
6ready to use

Six skills ship together, and the database marks all six "ready" — no paid account, no extra service, nothing beyond installing the plugin. ponytail itself is the ruleset: before touching code it walks a seven-rung ladder (does this need to exist, is it already in the codebase, does the standard library do it, is there a native platform feature, an installed dependency, can it be one line, only then the minimum that works), and it explicitly keeps validation, error handling, security, and accessibility off that ladder. The other five are single-purpose companions built on the same idea: ponytail-review and ponytail-audit look for over-engineering in a diff or an entire repository, ponytail-debt collects deferred-work comments into one ledger, ponytail-gain prints the benchmark numbers as a scoreboard, and ponytail-help is a reference for the rest.

It is aimed at anyone who has watched an agent reach for a library and a wrapper component to do what a native <input type="date"> already does, and wants a rule that catches that before the diff lands. The readme documents installation for roughly a dozen separate agent CLIs and harnesses beyond Claude Code — Codex, GitHub Copilot CLI, OpenCode, Gemini CLI, Devin CLI, Pi, Hermes, and others — which is unusually broad distribution for a repository this size.

READ THE FULL ANALYSIS

The headline number was revised down once, in public. An earlier single-shot benchmark reported 80-94% less code; a filed issue (#126) pointed out that its no-skill baseline pads answers with prose and options, inflating the gap. Measured again against a real agentic baseline — twelve feature tickets on a FastAPI-and-React template, Haiku 4.5, four runs each — the mean fell to about 54% less code, with cost down roughly a fifth and time down about a quarter, and ponytail was the only one of three arms tested that improved every metric while staying fully safe.

That number is still the maintainer's own measurement. We have not reproduced the FastAPI/React benchmark ourselves; the percentages above come entirely from the repository's own writeup and benchmark scripts, not an independent run.

One setup detail for two of the harnesses. The Claude Code and Codex plugins run their always-on activation through small Node.js lifecycle hooks; if node is not on the shell's PATH, the skills still work but that always-on trigger silently stays quiet instead of erroring.

6Grouped into 3 sets by the authors.
6Work with nothing else to set up.
75,041Stars on the GitHub repository, at last check.