OpenAI has developed GPT-Red, an LLM trained via self-play to automatically red-team its other models. Using a training loop where GPT-Red attacks and other models defend, it became more effective at finding vulnerabilities than human red-teamers. Its key discovery is a novel 'fake chain of thought' prompt injection attack that tricks models into acting on spoofed reasoning. When tested, over 90% of GPT-Red's strongest attacks succeeded against GPT-5, but fewer than 23% worked against the new GPT-5.6. OpenAI will not release GPT-Red publicly, citing the significant compute and year-long development investment required to build it.

6m read timeFrom technologyreview.com
Post cover image
Table of contents
Training dojo
7 Impressions