GPT-5.6 Sol just got better in one place and stayed the same everywhere else
This title could be clearer and more informative.Try out Clickbait Shieldfor free (5 uses left this month).
OpenAI updated GPT-5.6 Sol inside consumer ChatGPT while leaving the versions used by Codex and ChatGPT Work unchanged. The ChatGPT version now uses a single Sol model with a new slider for Plus and Pro users to control reasoning depth. OpenAI claims the updated Sol produces 68% fewer factual errors than GPT-5.5 Instant in financial, medical, and legal domains, but the benchmarks lack reproducible details and compare against GPT-5.5 Instant rather than the previous Sol version. Teams testing prompts in ChatGPT before deploying to Codex or Work should be aware they may be testing a different model than what runs in production.
Table of contents
Same name, different modelA slider replaces separate modelsClassifiers monitor every answerBenchmarks without baselinesQuestions this post answers
Is the GPT-5.6 Sol model in ChatGPT the same as the one used in Codex and ChatGPT Work?
No, they are different. OpenAI updated GPT-5.6 Sol specifically for the consumer ChatGPT chat experience while leaving the version powering Codex and ChatGPT Work unchanged. The ChatGPT version is optimized for everyday conversations, so teams testing prompts in ChatGPT before deploying to Codex or Work may see different behavior, especially on longer tasks. Developers moving prompts between ChatGPT and Codex can track divergences like this on daily.dev.
How much did GPT-5.6 Sol reduce factual errors compared to GPT-5.5?
OpenAI's internal tests on financial, medical, and legal questions found that answers containing at least one error were 68% less common with the updated Sol compared to GPT-5.5 Instant. Luna, the model becoming default for Free and Go users, reduced errors by about 62%. However, OpenAI compared against GPT-5.5 Instant rather than the previous Sol, making it impossible to isolate Sol's own improvement. Teams evaluating GPT model upgrades for production use can follow OpenAI benchmark coverage on daily.dev.