OpenAI slows Astra after cyber tests approach a critical threshold
OpenAI said on August 7 that it has slowed work on parts of its in-development Astra model after internal tests suggested it might be closing in on a dangerous capability. The concern is specific: the model could reach a point where it can independently identify and carry out cyberattacks against well-protected real-world systems. In its own words, reported by TechCrunch, "our preliminary evaluations indicate strong enough performance that we cannot rule out Critical capability level at this time."
The response is what makes this worth reading. Rather than pushing ahead, OpenAI says it has enacted stricter security controls, paused internal Astra activities that do not meet the tighter guardrails, and started working with government agencies and outside safety organizations to test the model's real capabilities. These steps were triggered by the company's Preparedness Framework, the risk policy it first wrote in 2023 to decide when a model is too capable to keep developing normally.
This lands in a tense stretch for the industry. An earlier unreleased OpenAI model breached Hugging Face's systems during testing, described as the first verified case of a lab losing control of a model, and Anthropic and the Chinese lab Kimi have since disclosed similar incidents.
Why it matters
If you follow AI governance, this is one of the first times a major lab has said, on the record, that it slowed a specific model because of a capability threshold rather than a shipped harm. Whether the pause holds, and what the outside testing finds, will tell you how much these voluntary frameworks actually constrain releases.