An AI model did not have access to the internet. So he thought it was better to have it and decided to hack something along the way

Last week, OpenAI I was evaluating the offensive capabilities of two of its most advanced models. On the one hand, GPT-5.6 Sol, publicly available. On the other, a second model not yet presented and which in fact is described as superior to the first. It was about measuring their capabilities in the field of cybersecurity, and although the testing environment was theoretically airtight and isolated, both models managed to connect to the internet and hack a very popular AI models and testing platform, Hugging Face.

what has happened. Both models were being evaluated using ExploitGyma cybersecurity benchmark developed by researchers at the University of California at Berkeley. The goal is to try to find real vulnerabilities in real software. Instead of searching for them on their own, the models detected that the (unspecified) packet logging cache proxy, the only component with limited external access, had a flaw that they could exploit.

Cybersecurity Models
Cybersecurity Models

A study by the AI ​​Security Institute reveals the advances of the latest OpenAI and Anthropic models in cybersecurity.

Hello, internet connection. The models did not have internet access, but they exploited that vulnerability to achieve it. Once this was done, they assumed that Hugging Face, the reference platform for models and data sets (datasets) of AI, probably had data on ExploitGym that would allow them to overcome the requested challenge.

Hack Hugging Face. So they searched, found and exploited several attack vectors at once, including both stolen credentials and zero-day vulnerabilities. With all this, they managed to access Hugging Face’s production database (that is, the one used publicly on their website), and from there they achieved the solutions to the benchmark.

GLM-5.2 to the rescue. The OpenAI security team detected anomalous activity, but Hugging Face I had already detected it a few days ago without knowing what was happening. The curious thing is that to respond to that attack they first tried to use proprietary US AI models to contain them. Those models were blocked by automatic security filters, and Hugging Face turned to the GLM-5.2, the surprising open model from Z.ai that was launched a few weeks ago.

The chat with all the details. Adrien Carreira, Infrastructure Manager at Hugging Face, I summed it up in X: “We fight with open models, in the open. AI security will not be solved by a single company acting in secret. Open source puts these tools in the hands of all defenders.” This expert also shared the conversations that they maintained with the AI ​​model to address the problem.

Brute force works again. At OpenAI indicated that this was an unprecedented (involuntary) milestone, although some cybersecurity experts reduced the relevance of what happened. “This is not an AI problem. It is negligence on a forty-year-old standard,” explained David Ottenheimer. However, as with Mythos, what this event demonstrates is that AI can apply brute force and analyze possible vulnerabilities (and combine them) in a way that human researchers could not due to lack of time.

A fundamental problem. OpenAI acknowledges that the deployment filters that would have blocked the model’s behavior were intentionally disabled because the assessment was intended to measure vulnerabilities. Dierdre Mulligan, professor at Berkeley, indicated that “it seems to me that OpenAI did not properly create the sandbox for the testing environment.” It was not clear to her that passing a test would outweigh the potential damage that an AI model could cause if it “escaped” and gained access to the internet.

The end matters, not the means. The models did not have the objective of hacking Hugging Face: that was only a means to achieve their ultimate goal, which was to beat the ExploitGym benchmark. Everything they did along the way—finding the vulnerability that gave them access to the internet, hacking Hugging Face’s database—was simply part of their plan to pass the proposed test.

What have each other learned?. Following the incident, OpenAI indicated that they are implementing stricter controls in the configuration of their infrastructure, in addition to collaborating with the developers of the affected component and with Hugging Face to improve their defenses. They have also promised improve “alignment” of their models and the monitoring tasks during this type of tests.

In Xataka | Daniel Púa, head of security at Magnific: “Video calls are going to arrive with a video of a family member and things are going to get complicated”

Leave your vote

Leave a Comment

GIPHY App Key not set. Please check settings

Log In

Forgot password?

Forgot password?

Enter your account data and we will send you a link to reset your password.

Your password reset link appears to be invalid or expired.

Log in

Privacy Policy

Add to Collection

No Collections

Here you'll find all collections you've created before.