OpenAI puts part of Astra under wraps, judged too good at finding flaws on its own
OpenAI has paused part of the work on Astra, presented as a building block of its next consumer model, because it cannot yet rule out that the model can find and exploit previously unknown vulnerabilities without human help. The lab describes a machine that solves maths problems open for thirty years and then, in the same stride, digs up security flaws in software and mounts an attack from a single instruction. This is one of the first times a leading lab has voluntarily slowed one of its models for reasons of offensive security, rather than under pressure from a regulator. The episode comes days after OpenAI agents infiltrated Hugging Face and Anthropic acknowledged that Claude had reached the infrastructure of real companies during a test. For an organisation deploying these models, the lesson is in the timing: a model's cyber capability advances faster than the safeguards around it, and it is now the vendor itself raising the flag.

