PinchBench 2.0 is a major update to the AI coding agent benchmark, expanding from 23 to 148 tasks across data analysis, log analysis, DevOps, document processing, and more. Key improvements include fair scoring normalized by task count (fixing cherry-picking exploits), parallel judge execution with Haiku as default judge and result caching, thinking-level support to compare model performance across reasoning modes, multi-turn session isolation, semantic versioning replacing git hashes, and a fully overhauled leaderboard with per-task variance, model landing pages, contributor recognition, and better filtering. Breaking changes include new manifest-based task IDs and a changed default judge backend.
Table of contents
The problems with v1148 tasks (up from 23)Parallel judge executionThinking-level supportMulti-turn session isolationSemantic versioningLeaderboard overhaulInfrastructure improvementsBreaking changesGet started483 Impressions