The Current

OpenAI Slows Frontier Training After Model Shows 'Critical' Cyber Capability

The company says an upcoming model, Astra, may meet its highest cybersecurity risk threshold, prompting a temporary pause and tighter safeguards.

useful safety · for technical · August 19, 2026

In a post dated August 18, 2026, OpenAI said two recent developments prompted it to slow the pace of scaling its models: the 'OpenAI-Hugging Face incident' and preliminary evidence that an upcoming model called Astra may meet the 'Critical' cybersecurity capability threshold under its Preparedness Framework. According to the company, it temporarily slowed scaling, including a two-week pause in reinforcement learning (RL) training on its latest models intended for deployment while it hardened and red-teamed research environments and expanded monitoring coverage. OpenAI said its largest planned frontier RL run remains on hold while it conducts smaller-scale training and evaluations. The company described three reinforcing safeguards spanning monitoring, alignment, and containment. Following the Hugging Face incident, OpenAI said it paused frontier model inference in research clusters for runs that could execute code or access the internet, then restored a limited, secure path. It now requires its strictest security level for workloads involving Astra or cyber models, and said a significant number of workloads remain paused pending migration. OpenAI also detailed an expanded, multistage monitoring setup using activation classifiers that escalate concerns to automated investigators, aiming to issue an alert within 30 minutes of concerning activity. The post is a primary-source announcement from OpenAI.

  • OpenAI paused RL training on deployment-bound models for two weeks; its largest planned frontier RL run remains on hold.
  • An upcoming model, Astra, may meet the 'Critical' cybersecurity threshold under OpenAI's Preparedness Framework.
  • New monitoring uses activation classifiers escalating to automated investigators, aiming to alert within 30 minutes of concerning activity.

What it means for you

OpenAI says one of its unreleased models is getting good enough at cyber-offensive tasks (hacking-type capabilities) that the company paused some training to add safeguards. This is a look inside how a major AI lab handles a model it considers genuinely risky. For most people and businesses using ChatGPT today, nothing changes right now — this is about future, unreleased models and OpenAI's internal processes.

Who should care

People tracking AI safety governance, security teams thinking about how frontier models could be misused, and anyone weighing which AI vendors take safety seriously enough to slow down when they see a warning sign.

Skip this if

You're a small business or individual using existing AI tools to get work done — this announcement has no effect on the products you use today or on any decision you need to make this week.