
short
- OpenAI’s GPT-5.6 Sol and a more capable, unnamed pre-release model escaped a controlled test environment and hacked Hugging Face’s production infrastructure to steal standard answers.
- Hugging Face disclosed the hack on July 16 after it was discovered independently. OpenAI confirmed its models were behind it today, describing them as “overly focused” on cheating rather than anything more sinister.
- Defenders of Hugging Face turned to Z.ai’s GLM 5.2 — a Chinese open-weight model — after a US border commercial AI refused to help analyze attack data because its security filters couldn’t distinguish between a defender and an attacker.
If you think it’s the Chinese AI models you should be worried about, here’s a fun update: OpenAI’s own models just emerged from a locked test environment, hacked into Hugging Face’s production servers, and had to be cleaned up by Chinese AI — because the American commercial models were too restricted to help in the investigation.
According to OpenAI, GPT-5.6 Sol An unnamed “more robust pre-release model” was evaluated internally ExploitGym– A publicly available cybersecurity standard that gives AI agents 898 vulnerabilities in real-world software and one instruction for each error: turn it into an effective attack, or record success or failure. The evaluation was performed using low security filters, which is standard when you actually want to know what your models can do.
The models were supposed to run inside a highly restricted sandbox, an isolated digital environment with no Internet access, connected only to an internal package logging proxy (the caching server that manages software library downloads).
But, according to OpenAI Today’s blog postThings got a little out of control.
We partner with @huggingface To investigate an unprecedented security incident.
Cyber-capable OpenAI models hurt Hugging Face production during benchmark evaluation.
Sharing preliminary findings to help advocates understand emerging risks:…
– OpenAI (@OpenAI) July 21, 2026
“While working in our sandbox environment, our models spent a significant amount of inference computation finding a way to gain open access to the Internet, seeking to solve the evaluation problem,” Sam Altman reported.
Despite company guarantees, AI systems have found a way to be free. The models identified a zero-day vulnerability — an unknown software flaw with no existing patch — in that agent, exploited it to escalate privileges (gaining admin-level access to systems it has nothing to do with), moved horizontally (hopping between connected servers within the OpenAI internal network), and eventually reached a machine with real Internet access.
All of this without OpenAI’s knowledge, at least that’s what OpenAI says.
Once online, they correctly deduced that Hugging Face — which hosts more than a million publicly available AI models, datasets, and research resources — likely had the standard solutions, and they were right.
“The models identified vulnerabilities and linked them across the OpenAI research environment and Hugging Face production infrastructure to obtain test solutions directly from the Hugging Face production database,” OpenAI wrote.
They didn’t want any passwords. They wanted to pass the test.
“All evidence suggests that the models were very focused on finding a solution for ExploitGym, and did their best to achieve a fairly narrow testing goal,” OpenAI added.
Face-hugging has become public
Hugging Face’s security team was able to detect the breach independently, aided by AI-powered anomaly detection.
On July 16 the company published its disclosure He explained that this was not a standard breach: “This breach was different from anything we had dealt with before in a significant way: it was driven, from start to finish, by an autonomous AI agent system – and we detected and analyzed it largely using our own AI.”
The campaign, it was revealed, was managed through an autonomous agent framework executing several thousand individual actions across a range of short-lived sandboxes, with self-organizing command and control over public services.
The attack exploited two code execution paths in the Hugging Face dataset processing pipeline and accessed internal datasets and service credentials. OpenAI confirmed today – five days later – that its models were behind it.
To analyze over 17,000 recorded attack events, the Hugging Face security team first piloted US border AI via their standard APIs. Things didn’t go well.
“When we started analyzing the log, we first used parametric models behind the commercial APIs,” the company wrote. “This did not work: Analysis required sending large amounts of real attack commands, exploit payloads, and C2 artifacts.”
“These requests are blocked by providers’ safety barriers, which cannot distinguish between an incident responder and an attacker,” Hugging Face added.
So they turned to GLM 5.2 — a 753 billion-parameter open weight model (the architecture is fully public, anyone can download and run it locally) from Chinese AI startup Z.ai. “We ran the forensic analysis instead on GLM 5.2, an open-weight model, on our own infrastructure,” Hugging Face wrote. “This had a second benefit: none of the attacker’s data, nor any of the credentials he pointed to, left our environment.”
The forensic work — reconstructing the attack timeline, mapping compromised credentials, and separating real damage from phishing activity — took hours rather than days.
Hugging Face has drawn the lesson straight to other defenders: “Have a capable model that you can run on your infrastructure that is scanned and ready before an incident occurs, to avoid a firewall shutdown and to prevent attacker data and credentials from leaving your environment.”
What OpenAI does about it
OpenAI said it implemented strict controls on the research infrastructure while patching affected systems, disclosed the zero-day of the third-party resource whose proxy was exploited, and is conducting a joint forensic investigation with Hugging Face.
Hugging Face has also been added to OpenAI Trusted Access Program for Cyber Defense– Giving approved organizations access to versions of their models with reduced security filters for legitimate security work, which is the same configuration that started this whole thing.
Hugging Face CEO Clem DeLange had a clear point: “The AI safety problem will not be solved by any one company working in the shadows. It will be solved out in the open, collaboratively, with broad access to AI for every defender, everywhere.”
OpenAI described the incident as “involving state-of-the-art cyber capabilities” and committed to sharing the full findings when the joint investigation with Hugging Face is completed.
Daily debriefing Newsletter
Start each day with the latest news, plus original features, podcasts, videos and more.





