OpenAI says it temporarily slowed development of its most advanced artificial intelligence models while overhauling the security, monitoring, and alignment systems used during research and training.

The decision followed two developments that raised concerns inside the company: a security incident involving OpenAI and Hugging Face, and early evidence that an upcoming OpenAI model called Astra may have reached what the company classifies as a “critical” level of cybersecurity capability.

Those developments prompted OpenAI to temporarily reduce the pace of scaling its frontier models, including a two-week pause in reinforcement learning training for its latest models intended for deployment.

“As models become more capable, the risks associated with developing and testing them internally also grow. Our standards for monitoring, alignment, and security must stay ahead of those risks,” OpenAI said in an August 18 blog post.

The move represents an unusual acknowledgment from one of the world’s leading AI developers that advances in model capabilities can outpace the infrastructure needed to safely study and train them.

A Major Training Run Remains on Hold

OpenAI said its largest planned frontier reinforcement learning run remains paused while researchers conduct smaller training runs and evaluations.

The goal is to gather more evidence about how the models behave, determine whether new safeguards work as intended, and evaluate alignment before allowing larger-scale training to resume.

The company said the slowdown reflects growing risks associated not only with deploying powerful AI systems but also with developing them inside research environments. Models with sophisticated cybersecurity capabilities may be able to write and execute code, interact with tools, access networks, or discover vulnerabilities, raising the stakes if safeguards fail.

OpenAI said Astra may have reached the “Critical cybersecurity capability” threshold defined under its Preparedness Framework. After reaching that preliminary determination on August 7, the company imposed additional monitoring requirements for Astra when the model uses tools.

OpenAI Tightens Its Research Environment

Part of the overhaul involves restricting what powerful models can access while researchers work with them.

Following the OpenAI-Hugging Face incident, the company paused frontier-model inference in research clusters when workloads could execute code or use tools capable of reaching the internet. Some workloads later resumed under tighter controls, while others remain suspended.

OpenAI has since introduced stronger workload isolation, additional network restrictions, reduced system privileges, enhanced security logging, and continuous testing designed to identify weaknesses in its research infrastructure.

The company said these measures have come with significant costs and delays.

OpenAI says a “significant number” of Astra-related workloads remain paused until they can be moved into environments that meet the new security requirements.

AI Is Increasingly Being Used to Watch AI

OpenAI is also expanding automated monitoring of what advanced models are doing during training and evaluation.

Its new system uses classifiers to examine model activity and escalate suspicious behavior to more sophisticated automated investigators. Those systems can examine tool usage and available reasoning for signs of unauthorized access, data theft, destructive activity, or attempts to circumvent safeguards.

If the monitoring system detects a potential violation of a critical security boundary, OpenAI says its safety, security, and research teams are alerted. If they cannot establish within 30 minutes that the warning was a false positive, the activity is expected to be stopped. The additional monitoring is computationally expensive. OpenAI estimates that it currently adds roughly 20% to the inference compute associated with the workloads being monitored.

Alignment Becomes Part of More Training Stages

The company is also extending alignment work deeper into the training process.

For reinforcement learning involving its most capable models, OpenAI says it is improving reward models, training systems to be more transparent about their actions and limitations, and targeting behaviors such as deception, reward hacking, unauthorized access, and attempts to exploit weaknesses in tools or oversight systems.

The change reflects a broader shift in AI safety strategy. Instead of evaluating dangerous capabilities primarily near deployment, developers may increasingly need security and alignment safeguards throughout the model-development lifecycle.

“The capabilities of frontier models are rapidly accelerating,” OpenAI said. “Our ability to understand, align, and secure them must stay ahead.”

What Comes Next?

OpenAI says it now plans to revise its Preparedness Framework to better account for risks arising during both model training and deployment.

The company also expects increasingly capable AI systems to play a larger role in cybersecurity itself, including helping defend systems against attacks conducted by other AI models. OpenAI plans additional investment in model-assisted security, monitoring, alignment research, and external collaboration as those capabilities advance.