OpenAI will launch Astra soon, but with a “panic button” we weren’t expecting

OpenAI is preparing the launch of Astra, its most advanced model, and this time the power comes with a particularly striking safety system. After the incident in which an OpenAI model attacked Hugging Face, the company has significantly strengthened its safeguards and systems. When it’s released, Astra will automatically stop in response to certain actions when the security subsystem considers it appropriate. A kind of panic button to keep its most advanced capabilities under control.

ChatGPT Download

OpenAI’s first model with “Critical” risk

According to Axios, OpenAI considers Astra to be the first model to reach the Critical level within its preparedness framework. It can find unknown vulnerabilities and develop ways to exploit them completely autonomously. During testing, Astra managed to discover and chain together two zero-day vulnerabilities, a major leap from what we already saw with the launch of GPT-5.6 and compared with OpenAI’s cybersecurity tools that have helped improve the security of Google Chrome.

If we consider agents like ChatGPT Work, capable of keeping tasks running for hours, the new approach, more focused on sustained control and not on control of the initial prompt or the results, is interesting. The longer an agent works, the more paths it can take and the more important it is for the system to monitor everything it does while it’s working.

With this approach, OpenAI has prepared safeguards capable of slowing down, pausing, or directly stopping a task when it detects certain signs of activity. In ChatGPT or Codex we’ll be able to request a review of the stopped action, while in the API the task will be automatically stopped at that point. In both cases, however, the greater capabilities will initially remain in the hands of a small group of testers while the new protections are fine-tuned.

Author: David Bernal Raspall

{ "de-DE": "Architekt | Gründer von hanaringo.com | Trainer für Apple-Technologien | Autor bei Softonic und iDoo_tech, zuvor bei Applesfera", "en-US": "Architect | Founder of hanaringo.com | Apple Technologies Trainer | Writer at Softonic and iDoo_tech, formerly at Applesfera", "es-ES": "Arquitecto | Creador de hanaringo.com | Formador en tecnologías Apple | Redactor en Softonic y iDoo_tech y anteriormente en Applesfera", "fr-FR": "Architecte | Créateur de hanaringo.com | Formateur en technologies Apple | Rédacteur chez Softonic et iDoo_tech, précédemment chez Applesfera", "it-IT": "Architetto | Fondatore di hanaringo.com | Formatore in tecnologie Apple | Scrittore per Softonic e iDoo_tech, precedentemente su Applesfera", "ja-JP": "建築家 | hanaringo.comの創設者 | アップル技術のトレーナー | SoftonicおよびiDoo_techのライター、以前はApplesferaで", "nl-NL": "Architect | Oprichter van hanaringo.com | Trainer in Apple-technologieën | Schrijver bij Softonic en iDoo_tech, voorheen bij Applesfera", "pl-PL": "Architekt | Założyciel hanaringo.com | Trener technologii Apple | Pisarz w Softonic i iDoo_tech, wcześniej w Applesfera", "pt-BR": "Arquiteto | Fundador do hanaringo.com | Instrutor em tecnologias Apple | Escritor na Softonic e iDoo_tech, anteriormente na Applesfera", "social": { "email": "races_provost0x@icloud.com", "facebook": "", "twitter": "https://twitter.com/david_br8", "linkedin": "https://www.linkedin.com/in/davidbernalraspall/" } }