Palenge · BLOG

Pulsa Esc para cerrar · / para abrir

IA

[ponytail] Writes Less Code, and Its Benchmark Graded Itself

Watch two contractors quote the same kitchen. The first arrives with a clipboard of additions: a new island, a wall to move, fresh wiring for appliances you already own. The second walks the room, opens a cabinet, and suggests rotating it ninety degrees. The smaller quote finishes sooner, and the kitchen works. Ask an AI agent for a date picker and it turns into the first contractor: it installs a library, wraps it in a component, adds a stylesheet, then opens a debate about timezones.

Ponytail exists to make the agent the second contractor.

It is one prompt, skills/ponytail/SKILL.md, with a compact AGENTS.md for agents that read a rules file. It adds no tools and no weights. What it changes is the order of the agent’s decisions. Before a single line gets written, the agent stops at the first rung that holds: does this need to exist, is it already somewhere in the codebase, will the standard library do the job, does the platform, does an installed dependency, can it be one line. Only past all of those does it write the smallest thing that works.

Why give a prompt this much room? The launch made Ponytail the fastest-spreading agent skill so far, around 44,000 stars in nine days, and the follow-up is more instructive than the speed. The README reported about 54% less code across twelve tasks on a real FastAPI and React repository. JetBrains then ran the same skill through its own harness and measured closer to 15%. Three-quarters of the advertised saving dissolved under a second pair of eyes.

The skill still trimmed code. The gap matters for another reason. Every prompt, plugin and skill in this young market arrives carrying numbers its author graded, red pen in hand. Incentive and measurement sit in the same palm, and that is why the figure that survives an independent run is the only one worth planning around.

Where does it pay off? On code that already exists. The prompt earns its keep on the grind of maintenance, where the tenth endpoint looks like the ninth and the glue between services never ends. It earns less on a fresh prototype, where a little wasteful scaffolding is how the shape of a problem becomes visible, and on problems where a mature dependency genuinely beats a hand-rolled line. Minimalism pushed too hard turns into its own debt: the abstraction deleted today is the one rebuilt next month. The ladder mostly sidesteps that, because the agent pauses before the first line, when thinking is cheapest.

Under the hood there is nothing to audit. The skill is a markdown file that steers whichever model reads it, which is why the same words work across Claude Code, Codex, Copilot, Cursor, OpenCode and Gemini. Claude Code and Codex install it through their plugin marketplaces; everywhere else the file gets copied in. Codex wants two lifecycle hooks trusted before a new thread. A handful of commands adjust the intensity, including /ponytail ultra for codebases that have earned it, and a short startup line reports the current mode. Nothing to compile and nothing to lock in.

Fuentes

Deja una respuesta

Tu dirección de correo electrónico no será publicada. Los campos obligatorios están marcados con *