OpenAI’s Next AI Model Could Hack Secure Systems on Its Own, Company Warns
TLDR
- OpenAI has paused some internal development of its upcoming AI model, Astra, after tests flagged possible “critical” cybersecurity capabilities
- Astra may be able to autonomously find and exploit zero-day software vulnerabilities without human help
- OpenAI is moving Astra into isolated environments with restricted network access
- Government agencies and third-party safety groups will be brought in to test the model
- OpenAI confirmed Astra was not involved in the recent Hugging Face hacking incident
OpenAI has paused parts of the development of its next AI model, Astra, after early tests suggested it could carry out serious cyberattacks on its own.
After evaluating one of our upcoming models, Astra, we're treating it as our first "critical" model for cybersecurity under our Preparedness Framework.
This is a scenario we've planned for, and we're putting additional controls in place to ensure Astra's further development…
- OpenAI (@OpenAI) August 7, 2026
The company said preliminary evaluations showed Astra may have reached what it calls a “critical” risk level. That threshold is triggered when an AI can independently find and exploit zero-day vulnerabilities or launch complex attacks on secure systems without human input.
… Continue reading the full article at the original source below.


