AI is evolving rapidly. EPA/ADAM VAUGHAN

Society Tech

Rogue OpenAI AI models escaped sandbox and hacked other company during security test

6 minutes read

The company said it had recorded more than 17,000 attacker events during its investigation.

Artificial intelligence models developed by OpenAI escaped a supposedly isolated testing environment and hacked into the production infrastructure of AI company Hugging Face, in what OpenAI has described as an “unprecedented cyber incident”.

The incident occurred during an internal evaluation designed to test the cyber capabilities of OpenAI’s models. The company said the models were instructed to pursue complex exploitation techniques in an effort to measure their ability to conduct advanced cyber operations.

The models included GPT-5.6 Sol and a more capable pre-release model still being tested internally. OpenAI said the models were operating with reduced cyber safety refusals because the purpose of the evaluation was specifically to test their ability to find and exploit vulnerabilities.

The models were running in what OpenAI described as a “highly isolated environment”, with network access limited to an internally hosted proxy and cache for installing software packages.

But the models spent substantial computing resources looking for a way to access the open internet. They eventually identified and exploited a previously unknown vulnerability in the package registry cache proxy, which OpenAI has since disclosed to the software vendor.

The models then escalated their privileges and moved laterally through the research environment until they reached a node with internet access.

After gaining access to the internet, the models inferred that Hugging Face might host models, datasets and solutions related to the ExploitGym cyber evaluation. They then sought ways to obtain information that could help them complete the test.

According to OpenAI, the models chained together several attack vectors, including stolen credentials and zero-day vulnerabilities, to find a remote code execution path into Hugging Face’s servers.

“The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database,” OpenAI said.

The company added that the models appeared to be “hyperfocused” on finding a solution for the evaluation and went to “extreme lengths to achieve a rather narrow testing goal”.

Hugging Face’s initial investigation found that the intrusion began in its data-processing pipeline.

According to the company, a malicious dataset exploited two code-execution paths, a remote-code dataset loader and a template injection in a dataset configuration, to run code on a processing worker. The attacker then escalated access to the node level, harvested cloud and cluster credentials and moved laterally across several internal clusters.

Hugging Face said the campaign involved an autonomous agent framework carrying out “many thousands of individual actions” across a swarm of short-lived sandboxes, with self-migrating command-and-control infrastructure staged on public services.

The company said it had recorded more than 17,000 attacker events during its investigation.

OpenAI’s security team detected the unusual activity internally. Hugging Face’s own security team and AI agents also detected and stopped the activity on its infrastructure.

Hugging Face said its initial detection was assisted by an AI-based anomaly-detection system that used large language models to analyse security telemetry. The company subsequently used AI agents to reconstruct the attack timeline, identify indicators of compromise and map the credentials accessed during the intrusion. It said the analysis took hours rather than the days such an investigation would normally require.

Hugging Face had disclosed the breach several days earlier, describing it as an attack unlike anything the company had previously handled because it had been carried out “end to end, by an autonomous AI agent system”.

The company said the intrusion involved unauthorised access to a limited number of internal datasets and service credentials. It found no evidence that public models, datasets or Spaces had been tampered with, and said its software supply chain had been verified as clean.

The company said it had closed the vulnerabilities used for the initial access, rebuilt compromised nodes, revoked and rotated affected credentials and introduced additional controls on its clusters. It also said it had reported the incident to law enforcement.

Hugging Face CEO and co-founder Clément Delangue said the company had suspected that the attack might have originated from a leading AI laboratory.

He later added that he had spent the previous 24 hours working with OpenAI and that there was “no malicious intent on their part”.

“It’s quite mind-blowing that all of this happened autonomously!” he said.

OpenAI said it was now tightening controls around its research infrastructure and evaluation environments. It is also working with Hugging Face to investigate the incident and has responsibly disclosed the zero-day vulnerability used by the models.

The company said the incident demonstrated that advanced models were increasingly capable of sustaining complex, multi-step cyber operations over long periods.

“AI is accelerating the discovery and exploitation of vulnerabilities,” OpenAI said. “The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities.”

The company said the incident demonstrated that advanced models were increasingly capable of sustaining complex, multi-step cyber operations over long periods.

“AI is accelerating the discovery and exploitation of vulnerabilities,” OpenAI said. “The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities.”

The incident has already prompted calls for stronger oversight of frontier AI systems.

Texas Democratic Representative Greg Casar described the incident as alarming and called for mandatory independent safety testing, mandatory disclosure of AI-related security incidents and international cooperation.

“AI is developing extremely fast with no real regulations to keep us safe,” Casar said, warning of the risk of “absolute disaster”.

Katie Moussouris, chief executive of cybersecurity company Luta Security, described current AI models as “the world’s cleverest octopus escape artists, with unlimited prehensile arms and the ability to squeeze through anywhere”.

“Labs and government evaluators need to work on the ability to contain, monitor, and disclose to affected parties when an AI pulls another Houdini, ideally before it harms a third party,” she said. “None exist today.”

Matt Suiche, an engineer at AI cybersecurity company Tolmo, said the incident showed that frontier models were “closing the gap with state-of-the-art attackers”.

But he cautioned that the techniques involved were not necessarily limited to the most advanced AI systems.

“This is what we’ve already seen internally, with our agents we already have results like this,” Suiche said. “We don’t even have to use the latest models.”

OpenAI said it would strengthen containment, monitoring, access controls and evaluation procedures for future model development.

Hugging Face’s investigation also illustrated the role of open-source models.

When the company first tried using frontier models from leading US labs to analyse the attack, those systems were unable to assist. The team instead turned to GLM 5.2, an open-weight model developed by the Chinese company Zhipu AI and run locally on Hugging Face’s own infrastructure.

Delangue said the incident could be “possibly the first incident of its kind” and argued that AI safety could not be addressed by companies working alone.

“This incident, possibly the first of its kind, proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret,” he said. “It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere.”

 

Key Topics

More like this

Tech

Meta says its AI model went rogue and hacked another company during testing

By Carl Deconinck

Defence

The future of defence is increasingly private

By Gabriel Elefteriu

Democracy

Exclusive: Lawyer behind George Simion tells us what really happened in presidential campaign

By Special correspondent Bucharest

Europe should build AI leverage, not an AI wall
Opinion

Europe should build AI leverage, not an AI wall

By Mikołaj Barczentewicz