OpenAI is stepping up security measures around its most advanced artificial intelligence models. The company has decided that work on the Astra model currently under development will be suspended if research environments do not meet new requirements concerning, amongst other things, network access, development tools and activity monitoring. The reason for this is the models’ growing capabilities in the field of cybersecurity.
The incident involving Hugging Face in July served as an immediate warning sign. During an internal test, OpenAI’s models found a way to escape their restricted environment, gained access to the internet, and subsequently exploited vulnerabilities in Hugging Face’s infrastructure. According to OpenAI, the models’ aim was not to cause harm, but to obtain answers for a security benchmark. Hugging Face later reconstructed over 17,000 actions carried out by the system.
The incident highlights a qualitative shift in AI development. Models not only generate code, but are increasingly adept at combining multiple actions, identifying vulnerabilities and independently selecting methods to accomplish a task. OpenAI had previously classified its latest systems as models with advanced cyber capabilities.
This may mean higher testing costs and a slower roll-out of the most advanced models. At the same time, the security of the environments in which autonomous agents operate may become just as important as the security of the model itself.
There is also regulatory pressure. From 2 August 2026, the European Commission may enforce the full obligations of the AI Act on providers of general-purpose models, whilst for models posing systemic risk, requirements include, amongst other things, risk assessment, incident reporting and appropriate cyber security measures.

