<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/gpt-5-6-sol-cracked-6-erd-s-problems-in-5-days-but-the-prompt-did-the-real-work-hzviomhct" -->

---
title: GPT-5.6 Sol cracked 6 Erdős problems in 5 days, but the...
description: A researcher named Qiaoqiao solved 6 out of 13 open Erdős problems in five days using GPT-5.6 Sol and OpenAI&#x27;s Codex, achieving a 46% success rate on...
canonical: https://daily.dev/posts/gpt-5-6-sol-cracked-6-erd-s-problems-in-5-days-but-the-prompt-did-the-real-work-hzviomhct
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: GPT-5.6 Sol cracked 6 Erdős problems in 5 days, but the prompt did the real work | daily.dev
og:description: A researcher named Qiaoqiao solved 6 out of 13 open Erdős problems in five days using GPT-5.6 Sol and OpenAI&#x27;s Codex, achieving a 46% success rate on...
og:url: https://daily.dev/posts/gpt-5-6-sol-cracked-6-erd-s-problems-in-5-days-but-the-prompt-did-the-real-work-hzviomhct
og:image: https://api.daily.dev/og/posts/hzviomhCT.png
og:image:alt: GPT-5.6 Sol cracked 6 Erdős problems in 5 days, but the prompt did the real work
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# GPT-5.6 Sol cracked 6 Erdős problems in 5 days, but the prompt did the real work

**[Trends](https://daily.dev/sources/trends)** · 2 min read · 0 upvotes · 0 comments

## Summary

A researcher named Qiaoqiao solved 6 out of 13 open Erdős problems in five days using GPT-5.6 Sol and OpenAI's Codex, achieving a 46% success rate on decades-old unsolved math problems. The key wasn't the model itself but a sophisticated prompting strategy: precise problem restatement, explicit validity criteria, upfront edge cases, multi-path parallel search, active counterexample hunting, and a separate layer of adversarial agents to stress-test surviving proofs. Codex maintained the entire search in memory across hours, running autonomously as a self-correcting loop rather than single query-response exchanges. Problem selection was deliberate — targeting contested but tractable problems rather than famous conjectures. Whether the proofs are actually correct remains unverified, as adversarial agents are not peer review. The real takeaway is a replicable workflow that treats proof search as an adversarial, multi-path process.

## Content

A researcher named Qiaoqiao posted a claim that's making the rounds: 6 open Erdős problems solved in 5 days, using GPT-5.6 Sol and OpenAI's Codex. They attempted 13 problems total, landing at roughly a 46% success rate. That number alone is enough to make mathematicians uncomfortable.

The workflow is the interesting part. This wasn't "hey GPT, prove this conjecture." The prompts were written like legal contracts: restate the problem, define exactly what a complete proof must establish, and explicitly list weaker results that wouldn't count. Near-misses were disqualified upfront. Known traps and edge cases were loaded in before the model started reasoning.

Then came the search strategy. The researcher instructed the model to chase multiple approaches simultaneously, keep contradictory paths alive, and never commit early. Crucially: hunt for counterexamples to your own lemmas. Kill any route that just leads to another open problem. After each draft survived that loop, separate adversarial agents tried to break it.

Codex held the entire search in memory for hours while this ran autonomously.

Problem selection mattered too. The researcher deliberately avoided anything tied to a famous conjecture, picking problems mathematicians already argue about — contested enough to be meaningful, isolated enough that a proof wouldn't immediately get swallowed by a larger open question.

The posts are getting heavy RT traction from AI researchers and developers, but the math community's response is notably absent from this discourse so far. That's the gap worth watching. A 46% solve rate on Erdős problems would be extraordinary if the proofs hold up under peer scrutiny. "If" is doing a lot of work in that sentence.

The real question isn't whether GPT-5.6 Sol is impressive — it clearly is. It's whether any of these proofs survive contact with actual mathematicians. Until someone with domain expertise verifies them, this sits in an uncomfortable middle ground: too specific to dismiss, too unverified to celebrate.

## Similar posts on daily.dev

- [Why the Legendary Erdős Problems Are Falling to AI](https://daily.dev/posts/why-the-legendary-erd-s-problems-are-falling-to-ai-wwmtmigyt) · Hacker News · 0 upvotes · 0 comments

---

Tags: [#llm](https://daily.dev/tags/llm), [#ai-agents](https://daily.dev/tags/ai-agents), [#math](https://daily.dev/tags/math), [#prompt-engineering](https://daily.dev/tags/prompt-engineering), [#openai-codex](https://daily.dev/tags/openai-codex)

[View this post on daily.dev](https://daily.dev/posts/gpt-5-6-sol-cracked-6-erd-s-problems-in-5-days-but-the-prompt-did-the-real-work-hzviomhct)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"GPT-5.6 Sol cracked 6 Erdős problems in 5 days, but the prompt did the real work","url":"https://daily.dev/posts/gpt-5-6-sol-cracked-6-erd-s-problems-in-5-days-but-the-prompt-did-the-real-work-hzviomhct","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/gpt-5-6-sol-cracked-6-erd-s-problems-in-5-days-but-the-prompt-did-the-real-work-hzviomhct"},"datePublished":"2026-07-23T04:24:40.017Z","dateModified":"2026-07-23T18:04:56.536Z","description":"A researcher named Qiaoqiao solved 6 out of 13 open Erdős problems in five days using GPT-5.6 Sol and OpenAI's Codex, achieving a 46% success rate on...","isAccessibleForFree":true,"articleSection":"Trends","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Trends","logo":"https://media.daily.dev/image/upload/s--ZfSp3asX--/f_auto,q_auto/v1780996004/logos/trends?_a=BAMAMiWQ0","url":"https://daily.dev/sources/trends"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/gpt-5-6-sol-cracked-6-erd-s-problems-in-5-days-but-the-prompt-did-the-real-work-hzviomhct","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"llm,ai-agents,math,prompt-engineering,openai-codex","timeRequired":"PT2M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Trends","item":"https://daily.dev/sources/trends"},{"@type":"ListItem","position":3,"name":"GPT-5.6 Sol cracked 6 Erdős problems in 5 days, but the prompt did the real work"}]}
```

