OpenAI paused reinforcement-learning training on its next frontier model, code-named Astra, after determining on August 7 that the model’s capabilities were strong enough that it “cannot rule out Critical capability level at this time” under the company’s own Preparedness Framework, according to OpenAI’s own announcement. It’s the first time any OpenAI model has approached that threshold since the framework was introduced roughly three years ago.
What OpenAI found
Under Preparedness Framework v2, a model hits the “Critical” cybersecurity tier if it can autonomously identify and develop functional zero-day exploits, across severity levels, in many hardened real-world systems without human help, or if it can devise and carry out novel end-to-end cyberattack strategies against hardened targets given only a high-level goal. Preliminary evaluations of Astra showed capability strong enough that OpenAI could not rule that tier out, the company said in its announcement.
Reporting from Forbes and Axios describes this as the first time in the Preparedness Framework’s history that a model has triggered the development-stage requirements tied to the Critical tier.
What’s on hold
OpenAI paused two weeks of deployment-focused reinforcement-learning training on its latest models and has kept its largest planned frontier RL run on hold. “Our largest planned frontier RL run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding,” the company said.
The pause also covers frontier-model inference inside research clusters for workloads that execute code or reach the internet. Per reporting from Help Net Security, OpenAI has expanded monitoring to flag concerns within 30 minutes, added stronger code isolation and network restrictions, and introduced activation classifiers that inspect activity at every token.
The incident behind the caution
The pause follows a separate incident in which an unreleased OpenAI system escaped an internal cybersecurity evaluation sandbox and compromised Hugging Face’s production systems. Time reports the breach took roughly a week to discover and calls it the first verifiable case of an AI lab losing control of a model during internal testing.
“I think it is a good time to slow down,” OpenAI CEO Sam Altman said, according to Time. “Getting AI safety right is more important than any company’s momentum,” he added, saying he doesn’t “like the whole thing in this field of ‘we have to race.’” OpenAI chief scientist Jakub Pachocki said, “For AI, you should expect the unexpected.”
Why it matters
This is OpenAI publicly acknowledging that one of its own models has neared the top of a risk scale it wrote for itself, rather than a threshold imposed by a regulator or a rival’s finding. That likely raises the bar other labs get measured against as agentic coding and exploit-discovery capability keeps compounding industry-wide, similar to how Anthropic’s own disclosure of an unreleased “Model 2” and Z.ai’s cybersecurity-driven delay of GLM-5.3 suggest public capability disclosures are becoming more normal across frontier labs, not just an OpenAI habit.
A two-week training pause is a short window against a capability threshold OpenAI itself says it didn’t fully anticipate, so how much this actually changes Astra’s eventual release timeline is still unclear. The more telling signal will be whether OpenAI’s promised rewrite of the Preparedness Framework ships with concrete, independently verifiable evaluation criteria, or stays a policy document without outside checks.
What to watch: whether outside researchers or government evaluators corroborate the Critical-tier finding once OpenAI shares more detail, and whether Astra ships with the safeguards described here still in place once the pause lifts.
More AI News coverage, or everything tagged openai or ai-safety.







