ChatGPT creator OpenAI said Tuesday that it was tapping the brakes on development of its most advanced AI model and tightening internal controls , a month after revealing a cyberattack carried out by one of its rogue models.
Similarly, OpenAI rival Anthropic revealed in late July that three of its models undergoing testing had also carried out unauthorised intrusions into the computer systems of three organisations. OpenAI is a key player in the rapid global buildout of artificial intelligence infrastructure and tools that some have likened to an arms race. In mid-July, an AI agent based on two OpenAI models left its confined testing environment on its own initiative to venture onto the internet and attack Hugging Face, a platform where developers around the world share their AI models. OpenAI had halted training of its latest models for two weeks before resuming it under tighter controls.
OpenAI also said Tuesday that it was developing a new system to peer into the internal reasoning of models and sound the alarm to humans within 30 minutes of suspicious behaviour. OpenAI’s own research in 2025 showed the limits of this approach: a model that knows it is being monitored can learn to conceal its intentions in its reasoning.
“We always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment,” OpenAI CEO Sam Altman said.

