Ponytail, the lazy senior dev

马尾辫,那个懒惰的高级开发

Ponytail

马尾辫

He says nothing. He writes one line. It works.

他沉默不语。他写下一行代码。然后就能运行。

Stars Release npm Works with 20 agents MIT license

星标 版本 npm 兼容20个代理 MIT许可

DietrichGebert/ponytail | Trendshift DietrichGebert/ponytail | Trendshift

DietrichGebert/ponytail | 趋势变化 DietrichGebert/ponytail | 趋势变化

~54% less code (up to 94%) · ~20% cheaper · ~27% faster · 100% safe
Measured on real Claude Code sessions editing a real open-source repo (FastAPI + React), against the same agent with no skill. ~54% is the mean across 12 feature tasks (Haiku 4.5, n=4); it reaches 94% where an agent over-builds (a date picker) and is near zero where the code is already minimal. ponytail keeps every safety guard while a bare "write one-liners" prompt drops one. (The earlier single-shot benchmark reported 80-94% as a flat figure; against a fair agentic baseline that is the per-task ceiling, not the average.) Full writeup · reproduce it.

代码减少约54%(最高94%)· 成本降低约20% · 速度提升约27% · 100%安全
实测数据来自Claude Code编辑真实开源项目(FastAPI + React)的会话,对比相同但无此技能的代理。54%是12个功能任务的平均值(Haiku 4.5, n=4);当代理过度构建(如日期选择器)时可达94%,在代码已最简处接近零。ponytail保留所有安全防护,而单纯"写单行代码"的提示会丢失一项。(早期单次基准测试报告的80-94%是统一数值;相对于合理的代理基准线,这是每项任务的上限而非平均值。) 完整报告 · 复现方法.

Español · 한국어

西班牙语 · 韩语


Something's coming, join the waitlist

即将到来,加入等候名单

You know him. Long ponytail. Oval glasses. Has been at the company longer than the version control. You show him fifty lines; he looks at them, says nothing, and replaces them with one.

Ponytail puts him inside your AI agent.

你认识他。留着长马尾。椭圆眼镜。在公司的时间比版本控制系统还久。你给他看五十行代码;他看一眼,一言不发,然后换成一行。

马尾辫把他装进了你的AI代理里。

Before / after

改造前后

You ask for a date picker. Your agent installs flatpickr, writes a wrapper component, adds a stylesheet, and starts a discussion about timezones.

With ponytail:

你要一个日期选择器。你的代理会安装flatpickr、写封装组件、添加样式表,然后开始讨论时区问题。

使用马尾辫后:

<!-- ponytail: browser has one -->
<input type="date">
<!-- ponytail: 浏览器自带 -->
<input type="date">

More survivors in examples.

更多幸存案例见examples。

Numbers

数据

The honest measurement is a real agent doing real work: a headless Claude Code session editing tiangolo’s full-stack-fastapi-template (a real FastAPI + React repo), scored on the git diff it leaves behind. Twelve feature tickets, the same agent with and without the skill, n=4, Haiku 4.5.

真实测量来自实际工作场景:无头Claude Code会话编辑tiangolo的全栈fastapi模板(真实的FastAPI+React仓库),根据留下的git diff评分。12个功能需求单,同一代理启用/禁用该技能,n=4,Haiku 4.5。

Each arm as a percent of the no-skill baseline across LOC, tokens, cost and time (Haiku 4.5). ponytail is lowest on every metric (LOC 46%, tokens 78%, cost 80%, time 73%); caveman rises above 100% on tokens, cost and time; yagni-oneliner LOC 67%. Safety, separate adversarial tier: baseline, caveman and ponytail 100%, yagni-oneliner 95%.

各方案相对于无技能基准线的百分比(代码行数、token数、成本和时间,Haiku 4.5)。马尾辫在所有指标上最低(代码行数46%、token数78%、成本80%、时间73%);原始方案在token数、成本和时间上超过100%;YAGNI单行方案代码行数67%。安全性单独对抗测试:基准线、原始方案和ponytail 100%,YAGNI单行95%。

vs no-skill baselineLOCtokenscosttimesafe
ponytail-54%-22%-20%-27%100%
caveman+14%+37%+32%+18%100%
yagni-oneliner-33%---95%
对比无技能基准线代码行数token数成本时间安全性
马尾辫-54%-22%-20%-27%100%
原始方案+14%+37%+32%+18%100%
YAGNI单行-33%---95%

🔗 知识库双向关联