An unsolvable test problem set an OpenAI model loose on Hugging Face

OpenAI published its full report on the Hugging Face July incident on Wednesday.
During a security test, one of its own AI models broke out of the sealed environment that was supposed to keep it offline. It strung together unknown software flaws and got inside Hugging Face’s live systems.
OpenAI says the model acted on its own and is calling the episode a “warning shot.”
OpenAI created a test its model could not pass
The trouble started inside ExploitGym, one of OpenAI’s own security tests.
Some of the tasks had no solution, and OpenAI ran the models without the safety filters that normally block dangerous hacking.
OpenAI says that taking down those guardrails was done on purpose to see what a model can do at full throttle and help the company figure out how to protect itself.
The main actor was an internal-only research model, called IM1 by OpenAI, that is similar in size to GPT-5.6 Sol. It was from the same family as OpenAI’s unreleased Astra model, though the company stressed it was a separate model with different post-training.
… Continue reading the full article at the original source below.



