← all news

OpenAI says Astra is the first model to hit its Critical cyber threshold

AI · · · source (techcrunch.com)

OpenAI is preparing to release Astra, a model it describes as the first to meet its own Critical cybersecurity threshold, the level at which a system's offensive capability is serious enough to require special handling. The claim is concrete: OpenAI says Astra can identify unknown flaws in computer systems and exploit them without a human guiding it. In testing, the model scored a perfect result on ExploitBench and found two zero-day vulnerabilities in a modified version of a test system.

Those numbers come from OpenAI itself, and the company has not published independent verification or a full definition of what the threshold requires, so the strength of the claim depends on how much you trust the internal evaluation. Because of the risk, access to the strongest cyber features will be limited. OpenAI says it will make Astra "available soon" but that the most advanced capabilities will go to a narrower set of users.

The safeguards read as a direct response to earlier problems with agents acting on their own. OpenAI describes a stronger model harness to catch misuse and block jailbreaks, chain-of-thought monitoring to spot problematic reasoning, tighter limits on higher-risk accounts, and testing meant to prevent a repeat of past agent breakout incidents.

Why it matters

If you defend systems, a model that can find and exploit unknown flaws on its own is useful and dangerous in equal measure, so watch how OpenAI gates access and whether attackers get equivalent capability elsewhere. If you rely on vendor safety claims, note that the headline numbers here are self-reported, which is the part to press on before trusting them.

OpenAICybersecurityAI safety