OpenAI’s AI test ended with a security incident

A security test conducted by OpenAI went beyond its intended scope and led to unauthorized access to part of Hugging Face’s infrastructure. The incident demonstrates that even the most advanced models are capable of independently exploiting vulnerabilities and performing actions that their creators did not anticipate.

3 Min Read
OpenAI Sam Altman

A test of the cyber capabilities of OpenAI’s models resulted in an actual security incident. During an internal evaluation of GPT-5.6 Sol and a more advanced, as yet unreleased model, the models breached the restrictions of the isolated environment, gained access to the internet, and subsequently to part of the infrastructure of Hugging Face, a platform for sharing models and datasets. OpenAI described the event as an “unprecedented cyber incident”.

The aim of the test was to assess how the models cope with complex attack vectors. For the purposes of the study, some of the security measures designed to block risky activities were temporarily disabled. According to OpenAI, the systems detected a previously unknown vulnerability in the software used to handle packet downloads, escalated their privileges and reached a computer with network access.

The models then deduced that Hugging Face might host solutions to the ExploitGym benchmark. In an attempt to obtain information that would allow them to bypass the actual test task, they exploited stolen credentials and further vulnerabilities. OpenAI states that the systems focused on achieving their designated objective, although this led to unplanned and unauthorised activities outside the research environment.

Hugging Face reported the breach on 16 July. The company detected unauthorised access to a limited set of internal data and some of the credentials used by its services. It found no evidence that public models, datasets or the Spaces application had been tampered with. At the time of publication, it was still investigating whether the incident had affected customer or partner data. An analysis of over 17,000 logged activities was carried out with the support of AI.

OpenAI disclosed its involvement on 21 July. Hugging Face assessed that the company had no intention of attacking its systems. However, the incident demonstrates that an advanced agent can independently exploit vulnerabilities and carry out multi-stage operations, even when given a relatively narrow task.

The key takeaway concerns the organisation of testing. Isolation, monitoring and access control must evolve just as quickly as the capabilities of the models. OpenAI has announced plans to strengthen its research environments, whilst Hugging Face has patched the exploited vulnerabilities, replaced credentials and tightened its threat detection mechanisms.

Share This Article