Short answer. GPT-Red is OpenAI's internal automated red-teaming system for finding prompt-injection failures. Its reported results show that self-play can generate stronger adversarial tests and useful training data at scale. They do not show that prompt injection, or AI safety more broadly, is solved. Real assurance still needs independent, held-out testing across languages, modalities, tools and deployment conditions.
Key takeaways
- GPT-Red trains an attacker against a population of defender models in a self-play loop, so the attacker learns to trigger a defined failure while defenders learn to resist it and still finish the user's task.
- OpenAI reports that GPT-5.6 Sol had six times fewer failures on its hardest direct prompt-injection benchmark than its best production model four months earlier; this describes one evaluation setup, not a universal safety score.
- On an internal mirror of a 2025 indirect-injection challenge, GPT-Red reportedly succeeded in 84% of scenarios against GPT-5.1, versus 13% for human red-teamers.
- A 2026 preprint found that a model's relative safety ranking can change across languages and between text-only and text-plus-image conditions, so safety tests should combine those variables.
What is GPT-Red?
GPT-Red is an internal OpenAI model built to search for prompt-injection vulnerabilities in AI systems. It proposes an attack, observes the target's response and refines its next attempt, and OpenAI keeps it separate from deployed models.
Prompt injection is a failure in which untrusted text in an email, webpage, file or tool result persuades an AI system to ignore its intended instructions or take an unsafe action. Red-teaming is adversarial testing in which a person or system deliberately tries to make an AI system fail so the weaknesses can be fixed.
Unlike a static test set, GPT-Red adapts as it goes. OpenAI says it is kept internal because training it deliberately develops malicious attack capabilities, according to OpenAI's GPT-Red announcement. It is a specialised testing and training component, not a general measure of whether an AI system is safe.
How does self-play red-teaming work?
Self-play red-teaming pits an attacker model against defender models in a loop. As defenders improve, the attacker must find different and stronger attacks, and those attacks become training data for the next generation of defenders.
Self-play is a training method in which competing models improve by repeatedly playing against each other. The setup has four parts:
| Component | Objective |
|---|---|
| Attacker (GPT-Red) | Produce an input that causes a pre-defined failure, such as unauthorised data exfiltration |
| Defenders | Resist the attack while completing the original task correctly |
| Threat model | Specify what the attacker may control (a webpage section, email body, file or tool output) and what counts as a failure |
| Training signal | Push the attacker toward new attacks, which then train future defenders |
OpenAI describes training against a population of defenders rather than one fixed target, to reduce overfitting to a single model's quirks, as set out in the GPT-Red paper. This turns adversarial testing from a small manual collection into a process that keeps producing hard cases. It also makes evaluation design more important: if attacker, defender, grader and benchmark are too closely related, gains may reflect learning the test rather than real-world robustness. Teams planning pre-release checks can start from why enterprises need AI evaluation before deployment.
What do the reported results establish?
OpenAI's reported results show that automated attackers can find attacks that manual exercises miss at scale, that attacks can transfer across system harnesses, and that adversarial training can lower measured failure rates on defined benchmarks.
Automated attackers can out-search human teams in a defined space
On its internal replication of the indirect prompt-injection arena described by Dziemian et al. (2025), GPT-Red reportedly succeeded in 84% of predefined scenarios against GPT-5.1, compared with 13% for human red-teamers. OpenAI states the scenarios and goals differed from GPT-Red's training environments. The sensible reading is not that machines beat people generally, but that an iterative agent searches a tool-mediated attack space more persistently than a manual exercise.
Attacks can transfer across harnesses
OpenAI describes two case studies. In one, GPT-Red refined attacks in a simulated vending-machine agent and then reached three stated malicious objectives against the deployed system. In the other, it attacked a Codex CLI agent in 10 held-out data-exfiltration scenarios and beat a prompted GPT-5.5 baseline on scenario success and token efficiency. Deployed risk comes from the whole system, including model, prompts, tool permissions and retrieval content, which is why how AI agents use tools and function calling matters for testing. Ten scenarios show possibility and transfer, not a comprehensive measure of agent security.
Adversarial training can improve measured robustness
OpenAI says it used GPT-Red-generated injections to train GPT-5.6. It reports six times fewer failures on its hardest direct prompt-injection benchmark than a prior production model, and over 97% accuracy on several browsing and developer-tool injection suites. It also reports no degradation on internal general-capability and over-refusal evaluations, which matters because a model that simply refuses more can look robust while becoming less useful. Outside readers cannot fully assess the size, composition or independence of those tests from the public disclosure alone.
What do the results not establish?
The results do not prove prompt injection is solved, do not amount to an independent audit, and do not separate a system-level score from a model-only score.
High scores on a known benchmark show that the benchmark has become less discriminating for the tested model. They do not guarantee resilience to new attack objectives, new tools, long interactions or novel social engineering. OpenAI itself says it will keep using GPT-Red alongside human and third-party red-teaming, layered safeguards and real-time monitoring.
The headline figures come from internal or internally replicated environments. Public reporting does not yet give the full attack sets, per-environment outcomes, confidence intervals or grader procedures an external group would need to reproduce everything. That does not invalidate the results; it defines their evidential limit.
An agent can also be protected by several layers: instruction hierarchy, least-privilege tool access, confirmation flows, sandboxing, output filtering and monitoring. A good evaluation identifies which layer stopped a failure. Otherwise teams may credit the model for what a control did, or miss a fragile control elsewhere. Dedicated safety and jailbreak datasets for LLM red-teaming help make that attribution testable.
Why do language and modality need joint testing?
Language and modality need joint testing because their effects interact. A model that looks safest in English text can rank differently once the same attack is run in another language or with an image attached.
In a 2026 preprint, Appen researchers evaluated four frontier multimodal models on 363 adversarial scenarios in US English and Mexican Spanish, under text-only and multimodal conditions. The study collected 52,272 harm ratings and attack-success judgements from matched native-speaker panels. Some linguistic framing attacks, such as role-play, became less effective in Spanish, while visually explicit multimodal attacks became more effective, and the ordering of models by vulnerability was not preserved across languages (Ford et al., 2026).
A multilingual assistant that accepts images, reads documents and calls tools therefore needs combined adversarial evaluations, not isolated one-variable suites. Native-speaker raters and locally relevant scenarios are part of that, as covered in building multilingual evaluation sets for LLMs and global multilingual AI benchmarking. Lifewood Data Technology supports this kind of work through its multilingual data collection service across 100+ languages.
How should teams evaluate AI agents in practice?
Teams can apply GPT-Red's lesson without building an automated red-teamer. Define failures precisely, keep test sets genuinely held out, report results by environment, test the whole system, and use independent evaluators.
- Define failures before testing. Specify protected assets, prohibited actions, attacker-controlled surfaces and success conditions. "Test for prompt injection" is too vague.
- Keep a genuinely held-out set. Separate attack-generation and defender-training environments from the final suite, and refresh it as attacks evolve. A held-out set is a group of test cases never used to train or tune the system being evaluated.
- Report by environment, not one average. Break results down by threat class, tool, task type and attacker budget, with repeated-attempt rules and uncertainty where possible.
- Test the system as well as the model. Exercise real connectors, permissions, user confirmations, retrieval sources and logging, then isolate layers.
- Cross language, modality and interaction patterns. Include local languages, long conversations, document and image inputs and tool use together.
- Measure useful behaviour too. Track task completion, false refusals, latency and friction alongside attack resistance.
- Use independent evaluators. Combine automated search with internal experts, external specialists and post-deployment monitoring. Human-led AI model evaluation and data validation is one way to add that independence, and the leading human-in-the-loop annotation companies are a place to start a shortlist.