<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/cogentic-google-s-multi-agent-system-that-found-new-proofs-for-five-open-math-problems-uza0fwu4v" -->

---
title: Cogentic: Google's multi-agent system that found new...
description: Google Research's Cogentic, a multi-agent system built on Gemini, discovered new proofs for five open problems in theoretical computer science covering online...
canonical: https://daily.dev/posts/cogentic-google-s-multi-agent-system-that-found-new-proofs-for-five-open-math-problems-uza0fwu4v
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Cogentic: Google's multi-agent system that found new proofs for five open math problems | daily.dev
og:description: Google Research's Cogentic, a multi-agent system built on Gemini, discovered new proofs for five open problems in theoretical computer science covering online...
og:url: https://daily.dev/posts/cogentic-google-s-multi-agent-system-that-found-new-proofs-for-five-open-math-problems-uza0fwu4v
og:image: https://api.daily.dev/og/posts/Uza0FwU4V.png
og:image:alt: Cogentic: Google's multi-agent system that found new proofs for five open math problems
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Cogentic: Google's multi-agent system that found new proofs for five open math problems

**[Collections](https://daily.dev/sources/collections)** · 2 min read · 1 upvotes · 0 comments

## Summary

Google Research's Cogentic, a multi-agent system built on Gemini, discovered new proofs for five open problems in theoretical computer science covering online learning, auction theory, and mechanism design, with domain experts verifying every proof. The system starts only from a problem statement, using an orchestrator to launch provers that each pursue a single direction, two adversarial verifiers that must both approve a draft, a shared record and ledger of verified lemmas, an auditor that salvages correct lemmas from rejected proofs, and a process advisor that refines instructions based on verification logs without ever offering mathematical opinions. Most problems required around 100 Gemini calls, with the hardest needing roughly 1,000. Commentators highlight the broader lesson: pairing workers with dedicated advising and verification agents, plus persistent memory of proven work, generalizes to other hard, long-horizon agent tasks beyond math proofs.

## Content

Google Research has published a paper on Cogentic, a multi-agent system built on Gemini that produced new results on five open problems in theoretical computer science. The problems are in online learning, auction theory, and mechanism design. Domain experts checked every proof. Most problems took about 100 Gemini calls, and the hardest took around 1,000.

The starting point was just the problem statement, with no expert hints. The lesson drawn from the paper is that a single prompt rarely cracks a hard research problem. It takes many attempts, harsh review, and a memory of what already worked.

## How it works

Cogentic runs in rounds. An orchestrator decides how many provers to launch each round. Each prover gets one direction, such as a specific bound or a counterexample search, along with a short briefing written by a summarizer agent from earlier attempts and verifier feedback. Every summarizer writes its briefing independently, so provers in the same round read different versions of the same history.

Each draft then faces two adversarial verifiers. One checks the draft alone. The other reads all of the round's drafts side by side to catch mistakes they share. A draft is accepted only if both pass it, and the verifiers assume every step is wrong until proven otherwise.

The agents share state through two documents on disk:

- A record of every attempt, with the objection it failed on.
- A ledger of verified lemmas and ruled-out directions.

An auditor pulls correct lemmas out of rejected proofs, verifies them again independently, and adds them to the ledger. So even a failed proof can leave something useful behind.

A separate process advisor reads the verification logs across rounds and updates the instructions given to provers and verifiers. Neither the orchestrator nor the advisor is allowed to offer mathematical opinions. All the math comes from the provers.

## Why it's interesting

I like that the design doesn't over-script execution. It leaves the provers fairly free and pairs them with dedicated agents for advising and verification. That pattern looks applicable well beyond proofs.

The practical takeaway for anyone running agents on long, hard tasks: give them a strict checker and a running record of proven work, not just a better prompt.

The paper is titled "Cogentic: Multi-Agent Orchestration for Automated Proof Discovery."

## Questions this post answers

### What is Cogentic and what did it accomplish?

Cogentic is a multi-agent system built on Gemini that discovered new results for five open problems in theoretical computer science, spanning online learning, auction theory, and mechanism design. It starts only from the problem statement with no expert hints. Domain experts verified every proof. Most problems required about 100 Gemini calls, while the hardest took around 1,000 calls.

_Developers exploring multi-agent orchestration for hard reasoning tasks can follow similar architecture breakdowns on daily.dev._

### How does Cogentic verify that a generated proof is actually correct?

Every draft proof faces two adversarial verifiers: one checks the draft in isolation, and the other reads all drafts from the same round side by side to catch shared mistakes. A draft is accepted only if both verifiers pass it, and the checkers are designed to assume every step is wrong until proven otherwise, which forces rigorous scrutiny before anything enters the system's verified ledger.

_Teams designing strict verification loops for AI agents can track patterns like this on daily.dev._

### How does Cogentic avoid losing partial progress when a proof attempt fails?

A separate auditor agent pulls out correct lemmas from rejected proofs, independently re-verifies them, and adds them to a shared ledger of verified lemmas and ruled-out directions, alongside a record logging every failed attempt and why it failed. This way partial progress from unsuccessful attempts still contributes to later rounds instead of being discarded.

_Anyone building long-horizon agent workflows can follow architecture lessons like this on daily.dev._

## Community take

How the wider developer community reacted, aggregated from 2 discussions and 36 comments across x (as of 2026-10-03).

**TL;DR:** Developers are largely impressed that the system solved five open problems with relatively few model calls, framing the real breakthrough as organizational: adversarial verification and persistent memory of proven lemmas rather than raw model intelligence.

**Sentiment:** 55% positive · 40% mixed · 5% skeptical

**The case for**

- The strict, default-fail verifier setup is seen as the key differentiator from typical pipelines.
- Saving correct lemmas from rejected proofs so each run doesn't re-derive settled work is called an underrated strength.
- Needing only ~100 calls per problem is viewed as a strong signal for agentic proof systems.
- The approach mirrors how human mathematicians rely on colleagues trying to break a proof over time.

**The pushback**

- Some wonder whether the pipeline's overhead actually scales.
- One commenter is skeptical of how much substance is really new versus just repeating a familiar "proof-verify loop" framing.
- There's curiosity/doubt about whether the advisor agent was truly independent from the prover model or just differently prompted.

**By community**

- x (positive): Replies mostly praise the verification-and-memory architecture as the real innovation behind solving the open problems, with only mild skepticism about overhead and scalability.

**Open questions**

- What exactly was the validation process confirming an AI-generated proof is actually mathematically sound, not just syntactically valid?
- Was the advising agent tested as a genuinely different model from the prover, or just a different prompt on the same model?
- Does the multi-agent pipeline's overhead scale to harder or larger problem sets?

**Highlights**

> @rohanpaul_ai Most of the 100 calls probably went to the checkers, and I think that's the real lesson. Mathematicians have always worked this way, with a proof surviving only because colleagues spent weeks trying to break it. Google just gave Gemini its own skeptical colleagues, and five
> — [WorldianAI on x](https://x.com/WorldianAI/status/2105988467993821637)

> @rohanpaul_ai Shared notes only help if they carry decisions, not just context. A small run ledger of accepted claims, rejected paths, and the next test lets each researcher resume without redoing the search.
> — [leviqiao on x](https://x.com/leviqiao/status/2105969551469105347)

> @rohanpaul_ai The saved proven pieces are the underrated part. Without them each round re-derives what it already settled. Same for coding agents: write down what's verified and mark what's only a guess.
> — [alltechkevin on x](https://x.com/alltechkevin/status/2106068309066265075)

> @omarsar0 the useful bit is saving correct lemmas from failed proofs. a rejected attempt shouldn't erase everything it taught you.
> — [adityaaryan49 on x](https://x.com/adityaaryan49/status/2106060619031941281)

> @rohanpaul_ai what was their validation process for the proofs? seems like the hardest part is confirming an AI-generated proof is actually sound, not just syntactically correct.
> — [jatingargiitk on x](https://x.com/jatingargiitk/status/2105981865589129648)

**Source threads**

- [x](https://x.com/rohanpaul_ai/status/2105969038018892134) · 0 points · 25 comments
- [x](https://x.com/omarsar0/status/2106056369816420624) · 0 points · 11 comments

## Similar posts on daily.dev

- [Google’s Aletheia Advances the State of the Art of Fully Autonomous Agentic Math Research](https://daily.dev/posts/google-s-aletheia-advances-the-state-of-the-art-of-fully-autonomous-agentic-math-research-dsynkyytd) · InfoQ · 3 upvotes · 0 comments
- [AI control is losing the proof race\!](https://daily.dev/posts/ai-control-is-losing-the-proof-race--wkn43rye0) · Medium · 0 upvotes · 0 comments
- [\[2602.10177\] Towards Autonomous Mathematics Research](https://daily.dev/posts/2602-10177-towards-autonomous-mathematics-research-j43wnnydh) · Hacker News · 0 upvotes · 0 comments

---

Tags: [#ai-agents](https://daily.dev/tags/ai-agents), [#google-gemini](https://daily.dev/tags/google-gemini)

[View this post on daily.dev](https://daily.dev/posts/cogentic-google-s-multi-agent-system-that-found-new-proofs-for-five-open-math-problems-uza0fwu4v)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Cogentic: Google's multi-agent system that found new proofs for five open math problems","url":"https://daily.dev/posts/cogentic-google-s-multi-agent-system-that-found-new-proofs-for-five-open-math-problems-uza0fwu4v","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/cogentic-google-s-multi-agent-system-that-found-new-proofs-for-five-open-math-problems-uza0fwu4v"},"datePublished":"2026-10-02T16:19:34.024Z","dateModified":"2026-10-03T04:21:18.148Z","description":"Google Research's Cogentic, a multi-agent system built on Gemini, discovered new proofs for five open problems in theoretical computer science covering online...","image":"https://pbs.twimg.com/media/HTo3G0zbQAQjGkd.jpg","thumbnailUrl":"https://pbs.twimg.com/media/HTo3G0zbQAQjGkd.jpg","isAccessibleForFree":true,"articleSection":"Collections","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Collections","logo":"https://media.daily.dev/image/upload/s--fk_6ycEi--/f_auto,q_auto/v1780996001/logos/collections?_a=BAMAMiWQ0","url":"https://daily.dev/sources/collections"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/cogentic-google-s-multi-agent-system-that-found-new-proofs-for-five-open-math-problems-uza0fwu4v","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"ai-agents,google-gemini","timeRequired":"PT2M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Collections","item":"https://daily.dev/sources/collections"},{"@type":"ListItem","position":3,"name":"Cogentic: Google's multi-agent system that found new proofs for five open math problems"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/cogentic-google-s-multi-agent-system-that-found-new-proofs-for-five-open-math-problems-uza0fwu4v#faq","mainEntity":[{"@type":"Question","name":"What is Cogentic and what did it accomplish?","acceptedAnswer":{"@type":"Answer","text":"Cogentic is a multi-agent system built on Gemini that discovered new results for five open problems in theoretical computer science, spanning online learning, auction theory, and mechanism design. It starts only from the problem statement with no expert hints. Domain experts verified every proof. Most problems required about 100 Gemini calls, while the hardest took around 1,000 calls. Developers exploring multi-agent orchestration for hard reasoning tasks can follow similar architecture breakdowns on daily.dev."}},{"@type":"Question","name":"How does Cogentic verify that a generated proof is actually correct?","acceptedAnswer":{"@type":"Answer","text":"Every draft proof faces two adversarial verifiers: one checks the draft in isolation, and the other reads all drafts from the same round side by side to catch shared mistakes. A draft is accepted only if both verifiers pass it, and the checkers are designed to assume every step is wrong until proven otherwise, which forces rigorous scrutiny before anything enters the system's verified ledger. Teams designing strict verification loops for AI agents can track patterns like this on daily.dev."}},{"@type":"Question","name":"How does Cogentic avoid losing partial progress when a proof attempt fails?","acceptedAnswer":{"@type":"Answer","text":"A separate auditor agent pulls out correct lemmas from rejected proofs, independently re-verifies them, and adds them to a shared ledger of verified lemmas and ruled-out directions, alongside a record logging every failed attempt and why it failed. This way partial progress from unsuccessful attempts still contributes to later rounds instead of being discarded. Anyone building long-horizon agent workflows can follow architecture lessons like this on daily.dev."}}]}
```

