OpenAI is preparing the launch of Astra, its most advanced model, and this time the power comes with a particularly striking safety system. After the incident in which an OpenAI model attacked Hugging Face, the company has significantly strengthened its safeguards and systems. When it’s released, Astra will automatically stop in response to certain actions when the security subsystem considers it appropriate. A kind of panic button to keep its most advanced capabilities under control.
OpenAI’s first model with “Critical” risk
According to Axios, OpenAI considers Astra to be the first model to reach the Critical level within its preparedness framework. It can find unknown vulnerabilities and develop ways to exploit them completely autonomously. During testing, Astra managed to discover and chain together two zero-day vulnerabilities, a major leap from what we already saw with the launch of GPT-5.6 and compared with OpenAI’s cybersecurity tools that have helped improve the security of Google Chrome.
If we consider agents like ChatGPT Work, capable of keeping tasks running for hours, the new approach, more focused on sustained control and not on control of the initial prompt or the results, is interesting. The longer an agent works, the more paths it can take and the more important it is for the system to monitor everything it does while it’s working.
With this approach, OpenAI has prepared safeguards capable of slowing down, pausing, or directly stopping a task when it detects certain signs of activity. In ChatGPT or Codex we’ll be able to request a review of the stopped action, while in the API the task will be automatically stopped at that point. In both cases, however, the greater capabilities will initially remain in the hands of a small group of testers while the new protections are fine-tuned.