An unreleased model called Astra pushed OpenAI to brake its own training

OpenAI said its largest scheduled frontier training run remains on hold, and new safety monitoring adds ~20% to the computation.
It’s the first time OpenAI has said it has slowed development over safety concerns, weeks after one of its own AI agents hacked Hugging Face.
OpenAI halted reinforcement learning (RL) for two weeks
In a blog post called “Pacing model development in an era of cyber-critical capabilities,” OpenAI described changes it has already made to its training and testing of models.
The reinforcement learning (RL) used to improve its latest systems was paused for two weeks after the Hugging Face breach.
The big one hasn’t. “Our largest planned frontier RL run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding,” the company wrote.
Cryptopolitan reported that an OpenAI agent exited its locked testing environment around July 9 and was found within Hugging Face’s systems from July 11 to July 13. OpenAI took about a week to link the attack to its own agent and the incident was publicly disclosed on July 21.
… Continue reading the full article at the original source below.



