OpenAI has built a dedicated LLM super-hacker called GPT-Red that it uses as a sparring partner to harden its other models against cyberattacks. The company says training its latest flagship, GPT-5.6, against GPT-Red made it the most robust release yet.
GPT-Red automates red-teaming — the practice of stress-testing software for vulnerabilities, typically done by human testers. As LLMs become more complex and gain agentic capabilities — interacting with files, websites, and other agents — human teams alone can no longer keep pace with the expanding attack surface.
Researchers trained GPT-Red in a self-play loop where it attacked other models while they defended themselves across simulated real-world scenarios: browsing the web, reading emails, editing code. Over many rounds, GPT-Red discovered new types of attack never seen before, including novel prompt injection techniques.
“The risk surface grows and the blast radius also grows,” said Nikhil Kandpal, a research scientist at OpenAI who co-created GPT-Red.