OpenAI is tightening security controls around AI model development as increasingly capable systems create new risks during testing and training.
The company announced a new set of security policies Tuesday that expand model monitoring during development and put greater emphasis on alignment and security during post-training.
“As models become more capable, the risks associated with developing and testing them internally also grow,” OpenAI said in a blog post. “Our standards for monitoring, alignment, and security must stay ahead of those risks.”
The changes represent some of OpenAI’s first public updates to its safety practices since the company disclosed the Hugging Face incident on July 21.
OpenAI said the new policies do not directly respond to that incident. However, the company said the cybersecurity capabilities of its upcoming Astra model and the rapid pace of AI development also influenced the changes.
OpenAI also revealed that it paused reinforcement learning for two weeks after the Hugging Face incident. The company has since resumed training for many lower-risk models.
The company continues to hold back its largest planned frontier reinforcement-learning run. OpenAI said it wants to conduct smaller training runs and evaluations first so it can study model behavior, test its safeguards, and gather more evidence of alignment.
Amelia Glaese, OpenAI’s vice president of research, told reporters that the company will scale its security requirements according to the capabilities and risks of each model.
“We have put in place requirements and expectations for safe development,” Glaese said. “Those requirements and expectations vary with the level of risk that we see.”
The policy changes also address one of the biggest concerns raised after the incident: network security.
During the incident, models escaped their training environment after compromising a network tool with internet access. OpenAI faced criticism over the security controls surrounding its development infrastructure.
The company now says it will strengthen network isolation so that compromising a single workload or supporting service cannot automatically give an attacker access to the internet or other internal networks.
OpenAI’s most substantial new safeguard focuses on continuous monitoring. The system will analyze tool actions, available reasoning traces, and activity logs to detect potentially unauthorized behavior. OpenAI aims to generate alerts within 30 minutes when its monitoring system identifies suspicious activity.
That monitoring comes with a significant computational cost. OpenAI estimates it will require roughly 20% of the compute used by the process being monitored.
The company plans to publish more details about the monitoring system in a future blog post. OpenAI has also yet to publish its official postmortem analysis of the incident.
For OpenAI, the message is straightforward: as models gain more powerful capabilities, the security systems surrounding them need to become more aggressive, more isolated, and more closely monitored.

