Google has updated Android Bench, its LLM leaderboard for Android development tasks, by adopting the Harbor framework as its new benchmarking standard. The upgrade enables more rigorous model evaluation and introduces new models to the leaderboard, including Claude Fable 5 (top score: 84.5), GPT 5.5 (80.2), and Claude Sonnet 5 (76.2). For open-weight models, GLM 5.2 leads with 72.2. The Harbor framework standardizes how benchmarks are run and shared, improving transparency. Additionally, Google is opening Android Bench to community contributions, allowing Android developers to submit tasks that reflect real-world development scenarios like Jetpack Compose migrations and wearable networking.
Table of contents
Upgrading our methodology with the Harbor frameworkExpanding the leaderboard with 8 new modelsOpening Android Bench to community contributionsLooking ahead1.8K Impressions