Ponytail is an open-source skill that instructs AI coding agents to follow YAGNI principles and avoid over-building, accumulating over 82,000 GitHub stars since June 2026. Its original benchmark claimed 80–94% code reduction, but a contributor challenge revealed the baseline was flawed — a chatty agent inflated the comparison. The maintainer responded by rebuilding the benchmark as a real agentic run on a FastAPI/React repo, revising the figure down to ~54% code reduction on average, along with ~20% lower cost and ~27% faster execution. The episode highlights a broader problem: AI agent skills and prompt frameworks proliferate with no evaluation standards, and Ponytail's public correction and new behavioral test framework may set a precedent for how skills should prove their claims.