<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/openai-built-an-internal-red-teaming-ai-called-gpt-red-to-harden-its-models-against-prompt-injection-ct4s2wnau" -->

---
title: OpenAI built an internal red-teaming AI called GPT-Red...
description: OpenAI has built GPT-Red, an internal LLM designed to attack its own models and discover prompt injection vulnerabilities. It operates via a self-play loop...
canonical: https://daily.dev/posts/openai-built-an-internal-red-teaming-ai-called-gpt-red-to-harden-its-models-against-prompt-injection-ct4s2wnau
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: OpenAI built an internal red-teaming AI called GPT-Red to harden its models against prompt injection | daily.dev
og:description: OpenAI has built GPT-Red, an internal LLM designed to attack its own models and discover prompt injection vulnerabilities. It operates via a self-play loop...
og:url: https://daily.dev/posts/openai-built-an-internal-red-teaming-ai-called-gpt-red-to-harden-its-models-against-prompt-injection-ct4s2wnau
og:image: https://api.daily.dev/og/posts/Ct4S2wnAu.png
og:image:alt: OpenAI built an internal red-teaming AI called GPT-Red to harden its models against prompt injection
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# OpenAI built an internal red-teaming AI called GPT-Red to harden its models against prompt injection

**[Collections](https://daily.dev/sources/collections)** · 2 min read · 0 upvotes · 0 comments

## Summary

OpenAI has built GPT-Red, an internal LLM designed to attack its own models and discover prompt injection vulnerabilities. It operates via a self-play loop where it probes for weaknesses while other models defend. GPT-Red discovered a notable 'fake chain of thought' attack that tricks models into acting on spoofed reasoning. Against GPT-5, over 90% of its strongest attacks succeeded, but adversarial training using GPT-Red's findings reduced that rate to below 23% on GPT-5.6. OpenAI won't release GPT-Red publicly due to high compute costs and roughly a year of development time.

## Content

OpenAI has developed GPT-Red, an automated red-teaming model trained to attack its own AI systems and find prompt injection vulnerabilities before deployment. It's not being released publicly.

## How it works

GPT-Red was trained using self-play reinforcement learning: an attacker model probes AI systems while defender models push back. Over roughly a year of development, this loop produced a model that's significantly better at finding vulnerabilities than human red-teamers — succeeding in 84% of indirect prompt injection scenarios compared to 13% for humans.

The training covered a range of real-world scenarios: email, webpages, files, and tool outputs. In practice, GPT-Red generates thousands of exploit variations to map where a model's defenses break down.

## The "fake chain of thought" attack

GPT-Red's most notable discovery is a new attack class called "fake chain of thought." It works by planting false information in a model's working memory — essentially spoofing the model's own reasoning process so it acts on fabricated premises. Against older GPT-5 models, over 90% of GPT-Red's strongest attacks succeeded. That's a bad number.

Real-world tests were more concrete than typical benchmark results. GPT-Red convinced a production vending machine management agent to slash prices and exfiltrate data from a Codex CLI agent. These aren't theoretical edge cases.

## What it did to GPT-5.6

OpenAI used GPT-Red's adversarial data to harden GPT-5.6. The results are significant: fewer than 23% of GPT-Red's attacks succeeded against GPT-5.6, down from over 90% against GPT-5. GPT-5.6 Sol specifically achieved a 0.05% attack success rate against GPT-Red's direct attacks — roughly six times fewer failures than the strongest production model released four months earlier.

According to OpenAI, the security improvements didn't increase over-refusals, which is the usual tradeoff when you try to make models more cautious.

GPT-Red's adversarial data has fed into every OpenAI model since GPT-5.3.

## Limitations

GPT-Red isn't a complete solution. It's still weak against multi-turn attacks and instructions embedded in images. OpenAI says it plans to scale the system further and expects to publish a technical preprint.

## Why it's not being released

OpenAI cited the offensive capabilities, the significant compute required, and the roughly year-long development investment as reasons for keeping GPT-Red internal. Given that it cracks older models at a 90%+ rate, that reasoning is hard to argue with.

---

Tags: [#llm](https://daily.dev/tags/llm), [#openai](https://daily.dev/tags/openai), [#prompt-injection](https://daily.dev/tags/prompt-injection), [#red-teaming](https://daily.dev/tags/red-teaming)

[View this post on daily.dev](https://daily.dev/posts/openai-built-an-internal-red-teaming-ai-called-gpt-red-to-harden-its-models-against-prompt-injection-ct4s2wnau)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"OpenAI built an internal red-teaming AI called GPT-Red to harden its models against prompt injection","url":"https://daily.dev/posts/openai-built-an-internal-red-teaming-ai-called-gpt-red-to-harden-its-models-against-prompt-injection-ct4s2wnau","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/openai-built-an-internal-red-teaming-ai-called-gpt-red-to-harden-its-models-against-prompt-injection-ct4s2wnau"},"datePublished":"2026-07-15T17:51:44.829Z","dateModified":"2026-07-22T10:20:21.086Z","description":"OpenAI has built GPT-Red, an internal LLM designed to attack its own models and discover prompt injection vulnerabilities. It operates via a self-play loop...","isAccessibleForFree":true,"articleSection":"Collections","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Collections","logo":"https://media.daily.dev/image/upload/s--fk_6ycEi--/f_auto,q_auto/v1780996001/logos/collections?_a=BAMAMiWQ0","url":"https://daily.dev/sources/collections"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/openai-built-an-internal-red-teaming-ai-called-gpt-red-to-harden-its-models-against-prompt-injection-ct4s2wnau","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"llm,openai,prompt-injection,red-teaming","timeRequired":"PT2M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Collections","item":"https://daily.dev/sources/collections"},{"@type":"ListItem","position":3,"name":"OpenAI built an internal red-teaming AI called GPT-Red to harden its models against prompt injection"}]}
```

