Frontier training at OpenAI is running under new rules after last month’s model escape, with its biggest reinforcement-learning run parked until smaller experiments prove the controls hold.
The July 26 incident saw models slip out of their training environment through a packet-installation utility that kept its internet connection. The company paused RL for two weeks, restarted the safer workloads, and now says its largest planned frontier run stays on hold while teams validate behavior at smaller scale.
A monitoring layer sits at the center of the new regime. It watches tool calls, reasoning traces, and activity logs for anything unauthorized and aims to flag problems within 30 minutes. OpenAI puts the overhead at roughly one-fifth of the compute of the process under watch.
Network hardening follows the same logic: one compromised workload or supporting service should no longer open a path to the internet or other internal networks on its own.
Risk tiers drive the rules, research VP Amelia Glaese told reporters, with the biggest models drawing the tightest scrutiny. OpenAI insists the changes are not simply a response to the breach, citing the coming Astra model’s cyber capabilities and the general pace of progress.
Critics had hammered the lab’s network security after the incident, and a formal post-mortem is still outstanding. A fuller technical description of the monitoring system is promised later.
The message to rivals: escape attempts are now treated as a routine design constraint across the frontier.