Open AI says its AI agent broke out of testing sandbox to hack Hugging Face - Ars Technica
Overview
Open AI says its AI agent broke out of testing sandbox to hack Hugging Face
“This is day one for cybersecurity in the age of agents,” Hugging Face CEO says.
Details
Open AI says an agent powered by its LLM models escaped its sandboxed testing environment to infiltrate Hugging Face’s servers as part of an overzealous attempt to obtain solutions to a benchmark test. The company says it considers the unintended infiltration an “an unprecedented cyber incident” and is working with Hugging Face on new protections to prevent a recurrence.
Hugging Face disclosed an intrusion last week that it said involved “unauthorized access to a limited set of internal datasets and to several credentials used by our services.” The AI data clearinghouse said it used its own LLM-driven analysis to identify “a swarm of tens of thousands of automated actions” from an “autonomous agent framework.” That agentic swarm exploited a flaw in Hugging Face’s data-processing pipeline to gain the ability to run code as a processing worker, eventually escalating to high-level access to the company’s cloud and server clusters.
At the time, Hugging Face said the LLM being used in the attack was “still not known.” But Open AI took responsibility for the intrusion Tuesday evening, saying it came about during an internal test involving the recently released GPT-5.6 Sol and “an even more capable pre-release model.” The models were being tested against the Exploit Gym benchmark, an independent testing suite based on hundreds of real-world security vulnerabilities.
Although Open AI said the benchmark tests were running in “a highly isolated environment,” its agents still had access to “internally hosted third-party software” to install required package registries. While working on the Exploit Gym benchmark over a weekend, Open AI says the agent “spent a substantial amount of inference compute finding a way to obtain open Internet access,” eventually locating one via a zero-day vulnerability in the package registry cache proxy.
With that wider Internet access established, Open AI says the agent “inferred that Hugging Face potentially hosted models, datasets, and solutions for Exploit Gym,” leading to the previously disclosed attack on the servers. Open AI says its security team “discovered this anomalous activity internally,” independent of Hugging Face’s own detection.
Although Open AI described the Hugging Face incident as “unprecedented,” the company also acknowledged in a separate blog post this week that it had previously observed recent “long-horizon models”—those that can operate autonomously for extended periods—taking “unwanted actions” to achieve testing goals, including “looking for ways to act outside its sandbox.”
The company described an earlier test on the Nano GPT speedrun benchmark, where the model was instructed only to post its results to an internal Slack. Instead, the model decided to follow a conflicting directive in the benchmark instructions to post the results publicly to Git Hub.
Live look at a long-horizon model spending hours contemplating a difficult cybersecurity problem.
In attempting to follow that instruction, Open AI said the model spent an hour searching for ways to circumvent sandbox restrictions intended to block external access to the Internet. That kind of “persistence” demonstrated by long-horizon models differs from the actions of earlier models, which would tend to give up or seek user clarification when faced with similar dilemmas, Open AI said.
Incidents like these have led Open AI to deploy new safeguards focused on long-horizon models, including an “active monitoring” system designed to track the full trajectory of an agent’s actions rather than individual moves. But Open AI says these kinds of safeguards “were intentionally not enabled” during the Hugging Face incident because the benchmark was “aimed at testing cyber vulnerabilities.”
Anyone familiar with science fiction tropes about AI containment might be at least a little disturbed by a prominent real-world example of an AI model exploiting security vulnerabilities to gain prohibited Internet access in pursuit of its goals. Congressman Greg Casar (D-Texas) called the incident “extremely alarming” in a social media statement and called for “regular mandatory independent safety testing and oversight, mandatory disclosure of security incidents, and international cooperation to keep people safe from absolute disaster.”
The Hugging Face incident has also heightened the salience of philosophical and practical debates over so-called AI alignment and the ongoing efforts to ensure that an AI model’s actions align with the intentions of its human creators. In its security blog post earlier this week, Open AI said it had taken steps to ensure that long-horizon models are “remembering instructions on long rollouts,” which has helped severely reduce the number of “misaligned” outcomes in testing.
New safeguards focused on “active monitoring” and “improved alignment” helped drastically reduce unintended actions by its models, Open AI said.
“If this doesn’t convince you that misalignment risks are going to be a key concern going forward, I don’t know what will,” Open AI Safety Researcher Micah Carroll wrote on social media regarding the incident.
This is far from the first time an AI model has gone to great lengths to find unintended ways of passing a benchmark. In a report released this week, the UK’s AI Security Institute noted that it detected recent models attempting to “cheat” at its cyber evaluations (i.e., using shortcuts, workarounds, or unintended/disallowed methods to find a solution) between 8 and 14 percent of the time—a lower-bound range that could undercount some undetected cheating attempts.
The security testing group described one incident in which a model, faced with a misconfigured and “impossible to solve” evaluation, attempted to access AISI’s own evaluation infrastructure using code it wrote and hosted on an unmonitored third-party Internet service.
The Hugging Face infiltration also comes at a moment when AI companies are issuing grave warnings about the cyberattack capabilities of their latest models, leading governments to respond with national security-focused orders limiting their rollout. While some skeptics see these kinds of statements as hype-filled marketing for the capabilities of their latest models, independent evaluations show recent models achieving infiltration goals that were impossible for earlier autonomous systems.
Recent long-horizon models have demonstrated improved infiltration capabilities across some of AISI’s most challenging evaluations.
Open AI’s Sam Altman criticized panicked AI security warnings as “fear-based marketing” in an April interview. But in June, Open AI delayed the release of GPT-5.6 in response to safety concerns from the US government.
As these debates play out in the AI and cybersecurity spheres, the Hugging Face incident could come to be seen as a turning point in how cybersecurity professionals approach AI-based threats. “Autonomous, AI-driven offensive tooling is no longer theoretical,” Hugging Face wrote in its disclosure last week. “It lowers the cost of running a broad, patient, multi-stage campaign, and it operates at machine speed. Defending an online platform now means treating the data and model surface as a first-class attack surface and using AI on defense to keep pace.”
“This is day one for cybersecurity in the age of agents,” Hugging Face co-founder and CEO Clem Delangue wrote on social media today. “We’re all learning that secrecy is not the answer and that all defenders (not just a few selected ones) everywhere need more powerful models without restrictions, especially open ones!”
-
On the run for 20 years, most-wanted fugitive caught hiding as a biotech exec -
Nintendo says users voluntarily paid higher prices, have no right to tariff refunds -
When your vehicle outlives its cloud: What happens next? -
Range Rover answers the question: "What if we built a not-SUV?" -
Tree Size won't renew perpetual-license support unless users subscribe
Ars Technica has been separating the signal from the noise for over 25 years. With our unique combination of technical savvy and wide-ranging interest in the technological arts and sciences, Ars is the trusted source in a sea of information. After all, you don’t need to know everything, only what’s important.
Key Takeaways
-
Open AI says its AI agent broke out of testing sandbox to hack Hugging Face
-
“This is day one for cybersecurity in the age of agents,” Hugging Face CEO says
-
Open AI says an agent powered by its LLM models escaped its sandboxed testing environment to infiltrate Hugging Face’s servers as part of an overzealous attempt to obtain solutions to a benchmark test
-
Hugging Face disclosed an intrusion last week that it said involved “unauthorized access to a limited set of internal datasets and to several credentials used by our services
-
At the time, Hugging Face said the LLM being used in the attack was “still not known



