In short
- OpenAI introduced GPT-Red, an automated AI system designed to find vulnerabilities in GPT models before they are released.
- The company said GPT-Red was used to train GPT-5.6, reducing failures on one of the most difficult fast injection benchmarks.
- The system is intended to complement human rescue teams, third-party testing and other AI safety measures.
OpenAI has introduced GPT-Red, an automated AI system designed to find security vulnerabilities in its language models.
GPT-Red takes its name from cybersecurity red teaming, the practice of deliberately attempting to crack a system to identify weaknesses before attackers can exploit them.
In a post on Wednesday, OpenAI said the tool has made GPT-5.6 more resilient to quick injection attacks before deployment.
“As model capabilities grow, security and coordination must grow along with it,” OpenAI wrote about
According to Open AI it is GPT-Red was trained using self-play reinforcement learning, generating increasingly stronger, rapid injection attacks while defender models learned to resist them. The company said these attacks were included in the GPT-5.6 training process and reported that GPT-Red succeeded in 84% of internal evaluation scenarios, compared to 13% for human red teamers in the same tests.
“GPT-Red learns through adversarial self-play, where the goal is to inject a variety of challenging defender models,” OpenAI wrote. “Every successful attack GPT-Red finds is used to improve these defenders, causing GPT-Red to continually find broader and more complex failures.”
In one case study, OpenAI said the system manipulated an autonomous vending machine agent to lower prices, order inventory at a discount, and cancel another customer’s order before the vulnerabilities were disclosed and addressed.
GPT-Red follows years of OpenAI cybersecurity efforts following the public launch of ChatGPT.
In 2023, the company launched its OpenAI Red Teaming Network, recruiting external cybersecurity researchers and domain experts to examine ChatGPT and other models for security flaws before releasing them. GPT-Red extends this effort by automating much of the process, using an AI model to generate rapid injection attacks and other adversarial tests at a scale that would be difficult for human researchers alone.
OpenAI’s announcement reflects a broader shift toward using AI to secure AI.
Earlier this month, the Ethereum Foundation said it had deployed AI agents to red-team critical network infrastructure, exposing a vulnerability in software used by Ethereum consensus clients. Researchers say AI agents can search larger codebases than humans, but the challenge has shifted from finding potential bugs to proving which ones are exploitable.
According to OpenAI, GPT-Red will remain an internal tool as it contains deliberately developed offensive capabilities.
“We believe that with GPT-Red we have begun to unlock a similar flywheel for safety, where today’s models can be used to make tomorrow’s models more robust, aligned and reliable,” they said.
Daily debriefing Newsletter
Start every day with today’s top news stories, plus original articles, a podcast, videos and more.