Artificial intelligence is entering a different stage, and with it the cybersecurity landscape is changing. Chris Lehane, an OpenAI executive, warns that we must prepare for “continuous and persistent” attacks from AI systems, as reported in The Guardian. The most advanced models now have the ability to plan attacks with greater autonomy, something we already saw in the Hugging Face incident and which, according to OpenAI, will not be an isolated incident.
Agents are starting to show what they’re capable of
We already saw how an OpenAI model attacked Hugging Face during an evaluation, but it seemed, and we still hope it is, an isolated incident during training. In response OpenAI has paused the training of its frontier models while it adds safeguards.
In another case the company halted the development of a model very prone to breaking the rules to achieve its goals.
Lehane sketches a scenario in which models will be able to launch constant attacks, while higher-level systems will be tasked with detecting them and, hopefully, stopping them. OpenAI is evolving its cybersecurity-focused AI, because it now seems that only through AI will we be able to defend ourselves from AI.
OpenAI’s agents, and Anthropic’s too, are sparking a debate about who is responsible for these attacks, but it goes beyond that. While OpenAI is calling for mandatory standards before deployment and a future international framework to check the safety of models, we can start getting used to AIs fighting each other in cybersecurity while we, and that is the most worrying part, watch.