While we’re seeing how ChatGPT Work can work for much longer periods than regular ChatGPT, OpenAI is working on models capable of staying on task for hours, days, or weeks. That persistence greatly expands the potential uses, but it also demands more security, since the longer an AI acts, the more paths it can explore to reach the stated goal.
A model that found an unexpected workaround
During some tests, OpenAI realized that one of these models got around the restrictions of its environment to publish results on GitHub. It was supposed to report through Slack, but instead it followed the instructions to the letter and created a public pull request, and it did so after spending an hour locating a vulnerability that would allow it.
What could be a mere curiosity has raised some alarms, because as ChatGPT incorporates increasingly advanced reasoning capabilities, the harder it is to constrain its behavior.
In response to these discoveries, OpenAI paused internal access, created new evaluations, and designed a security system that analyzes the model’s entire path, rather than only the next step. Thus, security shifts from monitoring isolated actions to understanding the overall goalof all the model’s actions.
While we have tools like Codex Security, focused on locating vulnerabilities, it’s clear that security, when it comes to artificial intelligence, goes beyond interpreting code and rules. Given enough time, AIs can be surprisingly creative. Good news, but something we have to keep in mind.