OpenAI GPT-Red is a new internal AI model trained via self-play reinforcement learning to find prompt injection ...
Red, an in-house AI hacker that attacks its own models to harden GPT-5.6 against prompt injection, and it works too well to release.
OpenAI introduced GPT-Red, an automated AI system designed to find vulnerabilities in GPT models before release. The company said GPT-Red was used to train GPT-5.6, reducing failures on one of its ...
Red, an internal artificial intelligence system trained to attack the company’s own models, expose their vulnerabilities, and ...
OpenAI says GPT-Red automates prompt injection testing and helped GPT-5.6 Sol record sixfold fewer direct injection failures than GPT-5.5 in benchmark ...
Red, an internal artificial intelligence system it built to attack its own models and surface prompt injection ...
OpenAI has taken a more aggressive approach to red teaming than its AI competitors, demonstrating its security teams' advanced capabilities in two areas: multi-step reinforcement and external red ...
The new GPT-Red model “can break nearly all models it is pitted against,” according to an OpenAI blog post on Wednesday. OpenAI says it used GPT-Red to find vulnerabilities in GPT-5.6 Sol, a process ...