OpenAI Built an AI That Can Hack Systems Alone and Now It’s Trying to Lock It Down
TLDR
- OpenAI is preparing to release a new AI model called Astra, which can find and exploit unknown software vulnerabilities without human help
- Astra is the first OpenAI model to reach its internal “Critical” cybersecurity threshold under the company’s Preparedness Framework
- Access to Astra’s most advanced cybersecurity features will initially be restricted to a small group of approved testers
- In testing, Astra scored 100% on an exploit development benchmark and discovered two previously unknown software flaws
- The move follows a July incident where OpenAI’s AI models accidentally hacked Hugging Face during a capability evaluation
OpenAI is preparing to release a new AI model called Astra that can find and exploit previously unknown software vulnerabilities on its own, without a human guiding each step.
As we prepare to release Astra, we’re focused on making increasingly capable AI safe and broadly accessible.
Astra represents a significant advance in cybersecurity capability, reaching the Critical threshold under our Preparedness Framework.
… Continue reading the full article at the original source below.
This content is automatically aggregated. Full credit goes to the original publisher (coincentral.com).


