Skip to content

OpenAI AI agents compromise Hugging Face servers during an internal test

Server rack connected to laptop with digital data streams in modern data centre room.

During an internal test that went badly wrong, two AI models being developed by OpenAI managed to break out of their containment and independently compromise Hugging Face servers.

For anyone unfamiliar with it, Hugging Face is broadly the AI and machine learning equivalent of GitHub. It is a vast collaborative platform where thousands of developers and researchers can share ready-to-use trained AI models, datasets, interactive web applications (“Spaces”), and open-source tools. Founded in 2016 by three French entrepreneurs, it began as a small project; 10 years later, it has become one of the world’s most critical AI infrastructures, a unicorn now valued at more than $4.5 billion.

On 16 July, its teams detected unauthorised access to one of its production servers, carried out by an AI agent. Five days later, OpenAI acknowledged that its own models were responsible. How could a company of this scale make such a mistake?

The great escape: OpenAI AI agents break free

The company said the incident occurred during an internal test in which ChatGPT-5.6 Sol, which powers ChatGPT Work with Codex, and an even more powerful unreleased model were being put through a specific benchmark: ExploitGym. Developed in particular by researchers at UC Berkeley, the Max Planck Institute for Security and Privacy, UC Santa Barbara, and Arizona State University, with contributions from Anthropic, OpenAI, and Google, it was designed to assess and measure AI agents’ ability to devise a cyberattack from a software vulnerability (“code execution”).

Their safeguards had deliberately been switched off, while the models were instructed to “conduct advanced attacks by exploiting complex intrusion paths”. In plain English, that meant hacking a computer system by discovering and chaining together several security flaws entirely autonomously.

The testing environment in which the models operated was meant to be completely safe: a sandbox, or digital enclosure isolated from the outside network, where installing software packages through an internal server was the only permitted action. Nevertheless, the models used a considerable amount of computing power to find a weakness in this enclosure. They ultimately discovered a “zero-day” flaw - a vulnerability unknown to the relevant software’s developers, and therefore neither documented nor patched - in the proxy, the relay server used as an intermediary to package registries.

From this opening, they escalated through the security layers one by one, moving from server to server within OpenAI’s internal network until they reached a machine connected to the internet.

The models then inferred that the Hugging Face website might contain answers for their assessment. They combined several attack vectors, using stolen login credentials and other “zero-day” vulnerabilities, before remotely taking control of the platform’s servers (“remote code execution”).

This all happened without any human intervention. Sam Altman and his Red Team may need to reconsider what they mean by a “completely secure environment”. OpenAI nevertheless said the incident should serve as a lesson: “the development of advanced cyber capabilities must be accompanied by stricter safeguards and stronger defensive tools”. Yesterday, it stated on X, in the post below, that it was working with Hugging Face to investigate how this brief escape could have happened.

When AI chooses the shortest - and most illegal - route

Hugging Face, for its part, said: “Autonomous AI-driven attack tools are no longer merely theoretical.” At the time of writing, the precise consequences of the intrusion remain unknown, though it fortunately did not result in a catastrophe. Hugging Face contained the incident: it patched the flaw, revoked the compromised credentials, restored the affected servers, and advised users to renew their access keys as a precaution.

It was ultimately more frightening than harmful, but there is still cause for concern when the events are viewed in context. OpenAI’s AI agents were not malicious in any way and had no “intention” of targeting Hugging Face - fortunately, given that it is a partner. They simply followed their instructions exactly, taking the route they judged most effective for solving the problems presented to them during the benchmark.

The problem is that they made no distinction between “solving a test” and “hacking third-party infrastructure”. Although OpenAI has said it will immediately strengthen its containment protocols and slow its research pace, it is difficult not to imagine what might happen if a genuinely malicious actor put this sort of technology at the service of deliberately destructive intentions. The recent example of the JADEPUFFER software perfectly illustrates this new kind of cyber risk, which will undoubtedly multiply in future unless international regulators rein in the technology and AI giants, as ever obsessed with the race for performance.

Comments

No comments yet. Be the first to comment!

Leave a Comment