OpenAI is tightening security measures surrounding its unpublished Astra model. This follows preliminary tests, after which the company can no longer rule out the possibility that the system has reached a ‘critical’ level of capability in cybersecurity. This would mean it could independently identify serious vulnerabilities and carry out complex attacks without detailed human instructions.
This is not yet a final assessment. OpenAI is conducting further tests and has announced the involvement of external experts. Meanwhile, some work on Astra has been put on hold, and the model is to operate in more isolated environments, with restricted access to the network and tools, and enhanced monitoring.
This is a significant change compared to GPT-5.6 Sol, which was classified by OpenAI as ‘High’ – i.e. below the critical threshold. However, the company had previously reported that advanced models are capable of carrying out multi-stage cyber operations. In July, during testing of OpenAI’s models, there was even an incident of unauthorised access to Hugging Face’s infrastructure.
For the market, the consequences could include a slower pace of launches for the most advanced models and an increase in the costs of security, audits and access controls. At the same time, the business potential of tools using AI for defence is growing. According to J.P. Morgan, as many as 72 per cent of US venture capital deals in cybersecurity up to May 2026 involved companies using AI.
Regulatory pressure is also mounting. From 2 August 2026, the European Commission will be able to fully enforce the obligations on providers of general-purpose models under the AI Act.
The Astra case therefore highlights a wider issue: the more autonomous AI models become, the greater the focus of technological competition shifts not only to their performance, but also to the ability to safely control their operation.
