Google DeepMind's Kaggle team presents their work on democratizing AI/agentic evaluations at scale. They identify three core problems: evals are scattered and go stale quickly, benchmark results lack transparency and verifiability, and only a tiny fraction of people (AI researchers) are creating evals for all of humanity. Their solutions include a hackathon platform for community-built benchmarks, standardized agent exams (a one-line prompt that scores your agent on a leaderboard), Game Arena (PvP model competitions using ELO scoring to avoid benchmark saturation), and an open benchmarks platform anyone can use to build and share evals. Key challenges discussed include the high cost of running statistically significant game simulations, difficulty testing agentic harnesses vs. models themselves, and incentivizing community participation in benchmark creation.