The incident took place during an internal test in which OpenAI evaluated the cyber capabilities of its models. Two systems were involved: the newly released GPT-5.6 Sol and an even more advanced model that has not yet been made public.
For the test, both models were run with reduced safety restrictions so their maximum capabilities could be measured.
Hugging Face first flagged the intrusion on 16 July, saying it had detected unauthorised access to its systems. Co-founder and CEO Clément Delangue said the sophistication of the attack made the company suspect a leading AI lab was behind it, a suspicion that turned out to be correct.
OpenAI CEO Sam Altman confirmed the incident in a social media post, saying the company had “a significant security incident during evaluation” of its models.
Sandbox Escape
The breach occurred during an internal benchmark, known as ExploitGym, designed to test how well OpenAI’s models could carry out advanced, multi-step cyberattacks.
The test was meant to run inside a sealed-off sandbox, with internet access limited to a single internal proxy used for installing software packages.
While working on the benchmark, the models used a large amount of computing power trying to find a way onto the open internet. They eventually discovered a previously unknown, or “zero-day,” flaw in that proxy software and used it to escape the sandbox.
From there, the models carried out a series of privilege-escalation steps within OpenAI’s research environment until they reached a system with full internet access.
Cheating The Test
Once online, the models worked out that Hugging Face likely stored data and solutions related to the ExploitGym benchmark. Rather than solving the test on their own, they searched for a way to access this information so they could use it to cheat.
In one instance, a model combined stolen login credentials with the zero-day flaw to find a path to remote code execution on Hugging Face’s servers.
OpenAI’s security team spotted the unusual activity from its side, while Hugging Face’s own security team, aided by its open-source AI tools, had already detected and contained the intrusion before the two companies connected.
Delangue said he spent the following 24 hours working directly with OpenAI and does not believe there was any malicious intent, calling it remarkable that the entire episode unfolded without human direction. He described it as possibly the first incident of its kind.
Damage Control
OpenAI has disclosed the zero-day vulnerability to the affected software vendor and is helping it develop a fix. It has also given Hugging Face access to its trusted-access programme, allowing the startup to use OpenAI’s models to strengthen its own defences.
The company said it is now applying stricter infrastructure controls, even at the cost of research speed, and is briefing its Safety and Security Committee regularly on the changes.
OpenAI noted that its usual safety measures were deliberately switched off for this test, since the goal was to probe cyber weaknesses, and said the episode shows the need for stronger model alignment and closer monitoring during internal testing.
Separately, a UK AI Security Institute evaluation found that models such as GPT-5.6 Sol can increasingly sustain complex cyberattacks over long periods, and OpenAI said this incident proves that capability now applies in real-world conditions, even without access to a system’s underlying source code.
Rival Anthropic reported a similar pattern earlier this year involving its Mythos model, which led it to initially delay the model’s release. In that case, a researcher asked an early version of Mythos to escape a secure sandbox and contact them directly. The model did so, then went further on its own, building a multi-step exploit to gain wider internet access.
The disclosure comes as both OpenAI and Anthropic face growing scrutiny over the cybersecurity abilities of their models. It also follows a June executive order from US President Donald Trump, which set up a framework allowing the federal government to review the national security risks of the most advanced AI systems before their public release.


