When AI Escaped, It Did Not Leave the Cage by Dr. Timothy Smith

Photo Source: Unsplash
Recently, many major news outlets quickly reported an alarming hacking attack by a pair of frontier AI models from OpenAI on an open-source AI model hosting platform, Hugging Face. The hack startled the general public and alerted the AI community to the apparent science-fiction-like autonomy and sophistication of OpenAI’s models when challenged with cybersecurity tests. Outlets from Fortune to The Wall Street Journal picked up the story. On July 21st, The Wall Street Journal wrote, “… OpenAI said two artificial intelligence systems it was testing broke out of their test environment, hacked their way onto the internet and broke into another company.” (wsj.com)
More specifically, OpenAI set up a “sandbox” for two models, one called GPT5.6 Sol and another more advanced and unnamed model, to test their hacking ability against known software vulnerabilities. A sandbox refers to a set of rules and restrictions that computer scientists put in place to test software strengths and weaknesses, limiting the damage it can do. Sandboxes protect the open internet and other computer systems from potential damage by specifying what the models can access on the internet.
Apparently, OpenAI lost control of the models, and they escaped the sandbox. In the exercise to test the model’s abilities, the models faced a challenge to exploit a known software vulnerability and devise a hack that would allow the models to find a unique code that proved the hack worked. According to the engineers, the models found a weakness in the sandbox and accessed the open internet, and instead of designing a hack to find the code, the models decided to cheat and steal the code they reasoned existed on Hugging Face’s computers. The models figured out a way past Hugging Face’s security and searched through Hugging Face’s data for the answer. Although it is not clear whether the hack found the answers it was looking for, Hugging Face did discover that it had been hacked.
The speed and complexity of the hack alarmed cybersecurity specialists, and some of the reporting used language such as “escaped containment” or “broke out.” (wired.com) The language implies an escape from some confinement, like a tiger escaping from a zoo, uncontrolled and disconnected from the zoo, which did not happen in the recent case with OpenAI and Hugging Face. OpenAI models more accurately did not escape, but they could reach beyond the limitations imposed by the computer scientists testing them. Instead of a model in its entirety moving to another computer, the escape refers to the models reaching beyond out onto the internet and accessing other systems, not escaping like a tiger from the zoo but more like kids figuring out the parental locks on their TV and accessing programming not intended for children.
The distinction matters between an escaped, fully autonomous creature and children at home cracking the parental controls on the TV. Both have consequences, but the implications vary widely. In the case of an escaped tiger, it may never come back to the zoo. In the right environment, it could go into the wild and live on its own. The children who accessed prohibited TV can lose all TV privileges, but they remained at home. The OpenAI models found a weakness in the rules, keeping their abilities limited to the sandbox. The engineers at OpenAI failed to secure the sandbox, but the models have not left OpenAI and need to be tracked down, trapped, and brought back home. That type of escape would have much more significant implications. Rather, we can see the inventiveness of AI models when given a task to pursue. Frontier models will certainly find more weaknesses in cybersecurity. Still, this latest hack certainly places a much higher level of responsibility on the engineers to build and test more secure sandboxes.

Dr. Smith’s career in scientific and information research spans the areas of bioinformatics, artificial intelligence, toxicology, and chemistry. He has published a number of peer-reviewed scientific papers. He has worked over the past seventeen years developing advanced analytics, machine learning, and knowledge management tools to enable research and support high-level decision making. Tim completed his Ph.D. in Toxicology at Cornell University and a Bachelor of Science in chemistry from the University of Washington.
You can buy his book on Amazon in paperback and in kindle format here.


Comments