---
title: "Agentic Evaluations at Scale, For Everybody — Nicholas Kang & Michael Aaron, Google DeepMind"
url: https://daily.dev/posts/agentic-evaluations-at-scale-for-everybody-nicholas-kang-michael-aaron-google-deepmind-fzgujdu27
source_url: https://www.youtube.com/watch?v=Ubwb6NzegyA
type: video:youtube
source: "AI Engineer"
published: 2026-05-25T17:06:57.306Z
updated: 2026-05-25T17:33:46.428Z
tags: ["agentic-ai", "google-deepmind"]
reading_time: 20
upvotes: 0
comments: 0
language: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Agentic Evaluations at Scale, For Everybody — Nicholas Kang & Michael Aaron, Google DeepMind

**[AI Engineer](https://daily.dev/sources/aidotengineer)** · 20 min read · 0 upvotes · 0 comments

## Summary

Google DeepMind's Kaggle team presents their work on democratizing AI/agentic evaluations at scale. They identify three core problems: evals are scattered and go stale quickly, benchmark results lack transparency and verifiability, and only a tiny fraction of people (AI researchers) are creating evals for all of humanity. Their solutions include a hackathon platform for community-built benchmarks, standardized agent exams (a one-line prompt that scores your agent on a leaderboard), Game Arena (PvP model competitions using ELO scoring to avoid benchmark saturation), and an open benchmarks platform anyone can use to build and share evals. Key challenges discussed include the high cost of running statistically significant game simulations, difficulty testing agentic harnesses vs. models themselves, and incentivizing community participation in benchmark creation.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://www.youtube.com/watch?v=Ubwb6NzegyA>

## Similar posts on daily.dev

- [Center for Responsible, Decentralized Intelligence at Berkeley](https://daily.dev/posts/center-for-responsible-decentralized-intelligence-at-berkeley-lkuvgfu7q) · Hacker News · 0 upvotes · 0 comments
- [EvalHub: Because "looks good to me" isn't a benchmark](https://daily.dev/posts/evalhub-because-looks-good-to-me-isn-t-a-benchmark-osnz9ucrd) · Red Hat Developer · 0 upvotes · 0 comments
- [Evaluating AI agents: Real-world lessons from building agentic systems at Amazon](https://daily.dev/posts/evaluating-ai-agents-real-world-lessons-from-building-agentic-systems-at-amazon-lkttxw05m) · AWS · 2 upvotes · 0 comments

---

Tags: [#agentic-ai](https://daily.dev/tags/agentic-ai), [#google-deepmind](https://daily.dev/tags/google-deepmind)

[View this post on daily.dev](https://daily.dev/posts/agentic-evaluations-at-scale-for-everybody-nicholas-kang-michael-aaron-google-deepmind-fzgujdu27)
