OpenAI's GPT-Red: Teaching LLMs to Hack Themselves for Safety
OpenAI deployed an adversarial AI model called GPT-Red to discover new attack vectors against its systems, improving robustness of GPT-5.6 through automated red-teaming.
OpenAI deployed an adversarial AI model called GPT-Red to discover new attack vectors against its systems, improving robustness of GPT-5.6 through automated red-teaming.
OpenAI introduced GPT-Red, an automated red-teaming model that discovers vulnerabilities at scale and trains GPT-5.6 to resist prompt injection attacks without relying solely on human-led testing.