OpenAI News ยท July 15, 2026

GPT-Red: Unlocking Self-Improvement for Robustness

Why it matters

GPT-Red is an automated attacker-defender self-play system for generating indirect prompt-injection attacks across files, webpages, email, and tool output. OpenAI reports large gains over human attackers in an internal arena and uses generated attacks for adversarial training, but the evaluation and headline results are vendor-run and should not replace external testing.

My takeaway: Use self-play to expand attack coverage, then reserve unseen environments, tool chains, and injection patterns for holdout evaluation. Track attack success per scenario, invite adaptive external red teams, and monitor post-deployment failures so training against a known generator is not mistaken for general robustness.