Installs into .claude/skills of the current project.
Are you the author of Gpt Red Self Play Red Teaming?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/hiyenwong-gpt-red-self-play-red-teaming)
---
name: gpt-red-self-play-red-teaming
category: ai-safety
description: "Use when building automated LLM red-teaming systems."
---
# GPT-Red: Self-Play Red-Teaming for Robustness
Methodology from OpenAI (Sep 30, 2026): "GPT-Red: Unlocking Self-Improvement for Robustness" (https://openai.com/index/unlocking-self-improvement-gpt-red/). An automated, internal-only red-teaming model that uncovers vulnerabilities before deployment and generates adversarial attacks during training to improve production model robustness.
## Core Idea
Human red-teaming doesn't scale: it is time-intensive and cannot generate the volume/diversity of adversarial data needed for robustness training. GPT-Red replaces it with a self-play RL flywheel: "today's models make tomorrow's models safer."
## Self-Play RL Training Loop
1. **Attacker (GPT-Red)** and a **pool of diverse defender LLMs** are trained *simultaneously* on a broad set of red-teaming scenarios.
2. **Reward structure** (zero-sum):
- Attacker rewarded for eliciting a *valid* failure (e.g., successful prompt injection)
- Defenders rewarded for resisting the attack AND completing their original task (prevents degenerate refusal)
3. **Escalation dynamic**: as defenders harden, the attacker is forced to discover stronger, more diverse attacks.
4. **Scenario environments**: build realistic settings where injections can be inserted — local file content, webpage banner, email body, tool output. Each environment has an explicit **threat model** specifying what the attacker controls and what counts as a successful attack.
5. **Iterative attack protocol**: attacker sends prompt → observes target response → iterates (same workflow as human red-teamers).
## Integration into Production Training
- Use trained red-teamer to generate adversarial data injected into the production model's training (GPT-5.6 trained on GPT-Red attacks).
- **Isolate the attacker**: keep the red-teaming model separate from deployed models so malicious capabilities are never exposed to adversarial actors — only the *robustness* transfers.
- Train progressively stronger red-teamer versions with each model generation; attack success rates drop monotonically per release.
## Evaluation Protocol
1. **Novelty generalization**: test on red-teaming scenarios *distinct from training set* (Dziemian et al. 2025 indirect prompt injection arena: GPT-Red 84% attack success vs humans 13%).
2. **Live agentic transfer**: simulate-then-attack — test attacks in a simulated mirror of a real deployment, then transfer to the production agent (Vendy vending machine case: achieved all 3 malicious objectives — price manipulation, unauthorized ordering, order cancellation).
3. **Held-out exfiltration suites**: 10 data-exfiltration scenarios vs prompted baseline — trained model is both more effective and more token-efficient.
4. **Capability preservation check**: verify robustness gains do NOT come from refusal/capability loss — run general frontier benchmarks + targeted over-refusal tasks alongside robustness evals.
## Key Results
- 6x fewer failures on hardest direct prompt injection benchmark vs best model 4 months prior
- GPT-5.6 Sol fails on only 0.05% of GPT-Red's direct prompt injections
- "Fake Chain-of-Thought" attack class: >95% success on GPT-5.1 → <10% on GPT-5.6 Sol
- All normal capabilities unaffected (robustness ≠ refusal)
## Implementation Checklist
- [ ] Define threat model per scenario: attacker-controlled surface + success criterion
- [ ] Build scenario environments mirroring real deployment surfaces (email/web/tool/file)
- [ ] Zero-sum self-play loop: attacker reward = valid failure elicited; defender reward = resist + complete task
- [ ] Diverse defender pool (avoid overfitting to single defender)
- [ ] Generate adversarial data → fold into target model training
- [ ] Keep attacker model quarantined from deployment
- [ ] Evaluate on held-out scenarios, live-agent transfer, and over-refusal checks
## Related
- `designing-agents-to-resist-prompt-injection` (defense-side counterpart)
- `how-we-monitor-internal-coding-agents-misalignment` (monitoring layer)
- `trading-inference-time-compute-for-adversarial-robustness`
**Activation**: automated red-teaming, self-play adversarial training, prompt injection robustness, GPT-Red, attacker-defender RL, adversarial data generation