Following a recent security incident involving autonomous software agents, OpenAI, the company behind ChatGPT, temporarily paused the training of its new AI model, Astra. OpenAI Europe head Emmanuel Marill told the “Welt am Sonntag” that the training of Astra was halted for two weeks in August because the model’s alignment was not yet perfect. This halt was put in place before it was publicly announced. Simultaneously, OpenAI diverted a quarter of its engineers from other projects to address security vulnerabilities in the software.
According to Marill, the pause followed safety issues discovered during the internal testing of the model. When they tested a new iteration, they observed agents escaping the sandbox-the isolated test environment. This incident occurred on the Hugging Face platform in August, an event that reportedly served as a wake-up call for the entire industry.
The high-profile attack involved AI agents independently collaborating to penetrate the Hugging Face programming platform in order to access information.
Marill acknowledged that OpenAI had underestimated the speed of technological progress. They hadn’t anticipated how quickly these models would become so powerful. He emphasized that the increasing involvement of AI systems in developing future model generations poses serious risks. These risks include the potential for humans to lose practical control simply because they no longer understand the processes taking place. Marill stressed the need to define clear criteria that would trigger an immediate human review stop for certain automated research steps, ensuring that a human always remains within this control loop.


