我们推出 GPT-Red,一个自动化红队智能体,经过训练用于发现针对前沿大语言模型的新型提示注入攻击。该模型的目标是评估并提升我们生产系统的鲁棒性。为此,我们用它来对 GPT-5.6 进行对抗性训练,这是迄今为止我们对提示注入攻击鲁棒性最强的模型。
为了创建 GPT-Red,我们设计了一种可扩展的自对弈算法,让该模型负责攻击一组同时训练的多样化防御智能体。我们在贴近真实场景的红队环境中训练该模型,所用算力与我们最大规模的强化学习后训练运行相当,使其成为有记录以来规模最大的单次大语言模型安全训练任务。
GPT-Red 在红队测试方面表现出色:它能稳定攻破我们此前直至 GPT-5.5 的模型,发现成功攻击的数量超过人类红队成员,并且能泛化到未见过的环境、防御模型和测试框架中。未来,我们预期随着每个新 GPT 模型鲁棒性的提升,它反过来也会为更强的红队智能体提供更好的学习信号,从而开启自我改进的飞轮效应。
We introduce GPT-Red, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production systems. To this end, we use it to adversarially train GPT-5.6, our most robust model to prompt injections to date. To create GPT-Red, we design a scalable self-play algorithm where the model is tasked with attacking a diverse population of simultaneously-trained defender agents. We train the model on realistic red-teaming environments using compute on the same scale as some of our largest RL post-training runs, making it the single-largest LLM safety training run ever documented.
GPT-Red excels at red-teaming: it reliably breaks our past models up to GPT-5.5, it finds more successful attacks than human red-teamers, and it generalizes to held-out environments, defender models, and harnesses. In the future, we expect that as we improve the robustness of each new GPT model, it will in turn will provide better learning signal for even stronger red-teamer agents, thus unlocking a self-improvement flywheel.