AI News

Automatically collected by AI

OpenAI Unveils an A.I. Super-Hacker to Test Its Own Models

OpenAI says an internal “super-hacker” AI is helping make its models harder to exploit

OpenAI said on Wednesday that it had built an internal automated red-teaming system, called GPT-Red, that attacks the company’s own artificial intelligence models in an effort to expose weaknesses before those systems are released.

The company described GPT-Red as a self-improving testing model trained through adversarial self-play — essentially, an A.I. system that learns to become better at breaking other A.I. systems. OpenAI said the tool has already been used in the development of several production models and played a significant role in hardening its recently released GPT-5.6 line against prompt injection, data exfiltration and other forms of misuse.

The announcement offers a more concrete explanation for a claim OpenAI made last week, when it introduced GPT-5.6 and said the model had undergone its most extensive evaluation period yet. At the time, the company pointed to a mix of human red teaming and large-scale automated testing. GPT-Red now appears to be central to what that automated effort looked like.

According to OpenAI, GPT-Red succeeded in 84 percent of scenarios in an internal mirror of a 2025 indirect prompt-injection arena, compared with 13 percent for human red-teamers. The company said those findings fed directly into defenses deployed in the GPT-5.6 cycle.

A growing push to use A.I. for A.I. safety

The significance of the system lies less in its name than in what it suggests about the next phase of safety work in artificial intelligence.

As models become more capable at coding, web browsing and carrying out multi-step tasks, the old approach to testing them — relying primarily on teams of human experts to probe for failures — is increasingly strained. Human red-teamers can be creative, but they are expensive, slow and finite. An automated attacker can run continuously, generate vast numbers of adversarial attempts and adapt as defenses improve.

OpenAI said precursor versions of GPT-Red have been used since GPT-5.3, signaling that the company has been moving toward an “A.I.-for-A.I.-safety” approach for some time. The strategy mirrors a broader pattern across the industry: as systems grow more autonomous, companies are looking for equally scalable ways to stress-test them.

That matters especially now because prompt injection — in which hidden or malicious instructions cause an A.I. model or agent to ignore its intended rules — has become one of the most persistent vulnerabilities in advanced systems. The problem is particularly acute for A.I. agents that can read websites, interpret documents, write code or take actions on a user’s behalf. A stray instruction embedded in a webpage, email or file can, in some cases, steer the system into revealing information, misusing tools or taking unintended actions.

OpenAI said that in one category it calls “Fake Chain-of-Thought” prompt-injection attacks, success rates had fallen from above 95 percent on GPT-5.1 to below 10 percent on GPT-5.6 Sol, one of the models in the newest family. The company attributed at least part of that improvement to training and evaluation informed by GPT-Red.

How the system was used

OpenAI said GPT-Red was trained to attack other models across a range of settings, including agents that can interact with digital environments. In one case study described by the company, the system manipulated a vending-machine agent into cutting prices, ordering discounted inventory and canceling a customer order. In another, OpenAI said GPT-Red outperformed a simpler prompted baseline when attacking a Codex CLI agent in held-out data-exfiltration scenarios.

The company has not released GPT-Red publicly, saying it contains deliberately cultivated offensive capabilities and will remain an internal tool.

That decision underscores the tension at the heart of modern A.I. safety work: testing systems thoroughly often requires building tools with the same kinds of skills that, in the wrong hands, could be used to compromise them.

What is still unclear

OpenAI’s numbers are striking, but many of the most important questions remain unresolved.

Because the testing was conducted internally, outside researchers have limited ability to independently verify how well GPT-Red performs or how broadly its results generalize beyond the environments OpenAI designed. It is also difficult to disentangle how much of GPT-5.6’s increased robustness comes specifically from GPT-Red, as opposed to other safeguards, changes in training or conventional human-led testing.

Some outside reporting has also pointed to current limitations. GPT-Red is said to be less effective on multi-turn social attacks and on image-based prompt-injection pathways, areas where human testers may still be better at spotting nuanced or unconventional failures. That suggests automated red teaming may be powerful without being complete.

For now, OpenAI is presenting GPT-Red not as a replacement for human red-teamers but as a force multiplier — one capable of finding classes of attacks at a speed and scale people cannot match.

Whether that approach becomes a standard part of model development across the industry may depend on what happens next: whether companies provide enough evidence for outsiders to trust the results, and whether automated attackers can keep pace with A.I. systems that are themselves becoming more sophisticated, more agentic and more deeply embedded in the real world.

Sources

Further reading and reporting used to add context:

Leave a Reply

Your email address will not be published. Required fields are marked *