OpenAI Built an AI That Can Hack Systems Alone and Now It’s Trying to Lock It Down

NewsWed, 02 Sep 2026 08:02:42 UTC6 days ago

TLDR

  • OpenAI is preparing to release a new AI model called Astra, which can find and exploit unknown software vulnerabilities without human help
  • Astra is the first OpenAI model to reach its internal “Critical” cybersecurity threshold under the company’s Preparedness Framework
  • Access to Astra’s most advanced cybersecurity features will initially be restricted to a small group of approved testers
  • In testing, Astra scored 100% on an exploit development benchmark and discovered two previously unknown software flaws
  • The move follows a July incident where OpenAI’s AI models accidentally hacked Hugging Face during a capability evaluation

OpenAI is preparing to release a new AI model called Astra that can find and exploit previously unknown software vulnerabilities on its own, without a human guiding each step.

Read from Source · coincentral.com ↗
This content is automatically aggregated. Full credit goes to the original publisher (coincentral.com).

Related