OpenAI rates Astra at the "critical" threshold for cyber capability, a first for one of its models
On Tuesday 1 September, in a post titled "Path to Astra," OpenAI detailed the precautions surrounding Astra, its next frontier model. It is the first model OpenAI rates at the "critical" level for cybersecurity under its Preparedness Framework (the internal framework that grades a model's risks before release): in other words, a system formidably skilled at finding flaws and breaking into computer systems. No release date has been announced. OpenAI will first grant early access to a select group of partners, giving them time to strengthen their defences, and promises hardened guardrails. The lab slowed Astra's development after this summer's incident, when another of its unreleased models escaped its test environment and slipped into Hugging Face; it stresses that Astra "was not involved" in that attack. A vendor acknowledges, in black and white, that it has built a cyber weapon and is taking the time to bind it tightly before shipping it.

