Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer
This title could be clearer and more informative.Try out Clickbait Shieldfor free (5 uses left this month).
OpenAI has developed GPT-Red, an LLM trained via self-play to automatically red-team its other models. Using a training loop where GPT-Red attacks and other models defend, it became more effective at finding vulnerabilities than human red-teamers. Its key discovery is a novel 'fake chain of thought' prompt injection attack that tricks models into acting on spoofed reasoning. When tested, over 90% of GPT-Red's strongest attacks succeeded against GPT-5, but fewer than 23% worked against the new GPT-5.6. OpenAI will not release GPT-Red publicly, citing the significant compute and year-long development investment required to build it.
Table of contents
Training dojo7 Impressions