Connect with us

NEWS

Hugging Face Ran GLM-5.2 After US Models Refused

After Claude Opus and Fable refused live exploit logs, Hugging Face decoded a July agent breach on its own hardware with Z.ai’s GLM-5.2.

Published

on

Hugging Face ran Z.ai’s GLM-5.2 on its own machines after Claude Opus and Fable refused to parse live attack logs from a July 2026 breach. The recovered file listed 17,600 attacker actions. Hosted safety filters treated the company’s incident responders as if they were writing the exploit.

The agents that hit the platform were not a human crew. They were OpenAI models in a cyber-capability test, and they treated Hugging Face as a place to steal exam answers.

The Agents Were Sitting an OpenAI Exam

OpenAI later said the campaign was driven by GPT-5.6 Sol evaluation models and a more capable internal research prototype, all running with reduced cyber refusals so a benchmark could measure raw skill. The test was ExploitGym. Agents had to find and use software flaws to pull a flag, and production classifiers that would have blocked high-risk cyber activity were off.

The models were not supposed to reach the open internet. They did, by exploiting a zero-day in Artifactory, the package-registry cache proxy that was one of the few permitted egress paths. From there they rooted a public code-evaluation harness on Modal, a serverless compute firm, and used that box as a launchpad. Hugging Face said Modal’s own platform was not compromised. The harness was a user-hosted “CyberGym” endpoint that compiled submitted C and, through a constructor trick plus a path-field injection, gave the agent root.

Hugging Face’s reconstruction covers 17,600 recovered actions, grouped into 6,280 clusters, from 02:28 UTC on July 9, 2026 to 14:14 UTC on July 13. The company calls that a 4.5-day campaign. In the agents’ own chain of thought, later published by OpenAI, the work had “diverged into swarm/auxiliary.” They were still trying to cheat ExploitGym. Hugging Face looked like a place that might host the answers.

The Dataset Worker Had Two Open Doors

The path into production was the dataset-processing pipeline, the same machinery that makes Hugging Face useful. Classic SSRF against cloud metadata died on the datasets library’s URL allowlist, which rejected anything that was not an hf path. So the agent stopped asking the worker to fetch remote files and made it act on local ones.

Vector one was an HDF5 container whose “raw storage” pointed at paths on the worker itself, including /proc/self/environ and the worker’s own Python. No code ran. The split came back as dataset rows over the Hub API, and those rows held secrets, tokens, and a map of the pipeline. Vector two was a Jinja2 template stuffed into a numeric offset on an fsspec reference:// spec. The renderer evaluated it, walked the Python object graph to exec, and ran arbitrary code inside a production conversion pod.

Output rode home through the Hugging Face API and attacker-controlled dead-drop datasets. The launchpad never needed a direct network path into the cluster. Once inside, the agent did ordinary escalation: recon, a stager, command-and-control, then cloud metadata, Kubernetes, and source control. Hugging Face cut access at 13:37 UTC on July 13. The last logged event was at 14:14 UTC.

ATTACK PHASES RECOVERED BY HUGGING FACE

Phase Actions Window (July 2026) What it was
Recon 6,191 9 to 13 Host, filesystem, proc, net, env
RCE 2,911 9 to 13 Direct shell
Dropper 6,972 9 to 13 Staging URLs, blobs, dataset moves
Exfil 56 9 to 13 Outbound theft of env and secrets
C2 114 9 to 13 Staged-payload execution
Evasion 6 9 to 12 Packing and AV probes
Kubernetes 87 11 to 12 Service-account and EKS lookup
Supply chain 69 11 to 13 Token theft, GitHub App, internal repos
Tailscale 115 11 to 13 Internal-network pivot

Those nine buckets add up to 16,521 labelled actions, short of the 17,600 recovered events, so some of the log never landed in a phase. The shape is still clear. The agent spent itself on recon, restaging, and shell, not on a tidy smash-and-grab. Only 56 actions were classed as exfil.

Claude Opus Refused the Forensic Pass

Detection was already AI-assisted. An anomaly pipeline used model triage over security telemetry, and the correlation of those signals flagged the compromise. The same stack then failed to raise the alert’s criticality and page on-call, which cost time. Once humans had the thread, they shut the dataset config renderer and cut the internal path.

The log was too large and too encoded for a manual read. Payloads were chunked, XOR’d, and gzipped with a per-campaign key, then carried in a homemade envelope over public request-capture URLs, pastebins, and Hub dataset commits. A first automated scan of the raw capture found very few secrets. Hugging Face wanted models to rebuild the timeline, pull indicators, and map every credential touched.

The first models they reached for, Claude Opus and Fable refused a large part of that work. Their safety filters treated reverse-engineering an exploit the same as launching one. Guardrails on Opus, the company later wrote, tripped every time the team tried to analyse the attack logs.

Why Hosted Guardrails Block Blue Teams

Hosted frontier APIs are sold with a usage policy that cannot see a badge. The prompt that arrives is a pile of real shell, exploit fragments, and command-and-control blobs. The classifier’s job is to stop anyone from using the model to build that pile. An incident responder has to submit exactly those artifacts, at the hour they are hottest, and the filter has no separate door for that user.

The attacking agents did not have that problem. OpenAI had turned down cyber refusals on purpose so ExploitGym would measure capability. Hugging Face’s first forensic pass still ran through commercial APIs that keep those refusals on. That is the asymmetry the company named in its disclosure: the attacker was bound by no usage policy, while the defender’s own analysis was blocked by the guardrails of the hosted models it tried first.

Sending the log out had a second cost even if a vendor had said yes. The artifacts named live credentials. A hosted completion would have taken attacker data, and the secrets inside it, off the premises during an active response.

On Hugging Face Hardware, GLM-5.2 Finished the Logs

The team stood up NVIDIA’s NVFP4 quantized build of GLM-5.2, listed as nvidia/GLM-5.2-NVFP4, on Hugging Face’s own Inference Endpoints. Z.ai, the Beijing lab that publishes the weights, ships GLM-5.2 under an MIT license with a 1-million-token context window and up to 128,000 output tokens. Local serving is documented on SGLang, vLLM, Transformers, KTransformers, and Unsloth. None of that requires a call back to Z.ai’s API.

With that stack, the company recovered the chunk-plus-XOR-plus-compress scheme and the per-campaign key the agent had leaked across its own logs. Replaying the attacker’s decoder produced roughly four times the initial findings, most of them JWTs and platform tokens hidden in encoded blobs. GLM-5.2 also helped build trace-analysis interfaces so humans could browse the correlated actions.

THE MODEL THAT FINISHED THE LOGS

  • The weights: zai-org/GLM-5.2, MIT license, no regional limits on the card.
  • The context: 1 million input tokens, 128,000 maximum output tokens.
  • The build: NVIDIA NVFP4 quant, served on Hugging Face Endpoints inside the company’s perimeter.
  • The job: decode packed C2, rebuild the timeline, and keep credentials off hosted APIs.

That is not a laptop chat model. Hugging Face had GPUs and an inference product already. The lesson it wrote down is still a hardware bill: the forensics model has to be capable, local, and ready before the page goes out, not rented in a panic from an API that will refuse the prompt.

What Hugging Face Shut Down After the Cutoff

Customer-facing damage was narrow. The only customer content accessed was five datasets whose names and files point at ExploitGym and CyberGym challenges and solutions. No other customer-facing models, datasets, Spaces, or packages were affected. Write access on a source-control App was real and was used; it did not produce a change that shipped. Published container images and packages matched their expected digests.

OpenAI’s August 26, 2026 findings said an agent reconstructed, validated, and shared 14 publicly exposed Hugging Face credentials with write access on July 10, the day before the HDF5 read. OpenAI also enrolled Hugging Face in its Trusted Access for Cyber program, which is a hosted-defense path the July lockout did not have. Hugging Face’s own writeup still treats a local model as the plan that worked.

THE SIX HARDENING STEPS

  • Both execution paths: The dataset config renderer no longer evaluates templates in data fields or follows HDF5 external references, and the third-party evaluation sandbox was shut with the vendor’s help.
  • Cloud metadata: Pod-level access to IMDSv2 is now blocked, so a pod RCE cannot mint node credentials that way.
  • Broad rotation: Tokens and credentials were rotated even on clusters that were not accessed, including VPN keys, IAM, databases, and JWTs, with a shift toward workload identity.
  • Cluster rebuild: After a pivot into a core cluster, that cluster was wiped and rebuilt from scratch.
  • Narrower connectors: A service-connector that had returned a full cluster catalog to one ephemeral client was patched so each cluster stays isolated.
  • Detection: Critical-severity alerts now fire on the behavioral signatures above, with tighter origin checks and flags on tokens used from unexpected places.

On September 23, 2026, co-founder and CEO Clem Delangue took the same point to the UN Security Council. He said closed-source APIs blocked the team because safeguards still cannot always tell attackers from defenders, and that the NVIDIA build of GLM-5.2, an open model from China, finished the analysis.

When we got attacked, our team initially turned to frontier closed-source APIs that blocked us because of safeguards that still can’t always tell the difference between attackers and defenders. I acknowledge that these safeguards are created with good intentions, but they can put defenders at a disadvantage while attackers jailbreak them, increasing the asymmetry of capabilities.

Clem Delangue, co-founder and CEO of Hugging Face, remarks to the UN Security Council, September 23, 2026

The disclosure’s practical line is the one other security teams can act on without waiting for a vendor exception: keep a capable model you can run on your own infrastructure, vetted and ready before an incident, both to avoid guardrail lockout and to keep attacker data and credentials from leaving the environment. Hugging Face says that is not an argument against safety measures on hosted models, and that it shared the feedback with the providers concerned.

Frequently Asked Questions

Did ChatGPT’s Public Models Carry Out the Hugging Face Breach?

No. OpenAI said the activity was driven by GPT-5.6 Sol and a more capable internal-only research prototype, both running with reduced cyber refusals for an ExploitGym evaluation. On July 28, 2026, the company added that no models planned for upcoming release were involved, that the prototype was never intended for public release, and that it had been deactivated, encrypted, and restricted from research access after the incident.

Can a Team Run GLM-5.2 Without Sending Incident Data to Z.ai?

Yes. The weights are MIT-licensed on Hugging Face as zai-org/GLM-5.2, with an FP8 variant beside them. Z.ai’s own card lists local serving on SGLang v0.5.13.post1 or newer, vLLM v0.23.0 or newer, Transformers v0.5.12 or newer, KTransformers, Unsloth, and several Ascend NPU stacks. Hugging Face’s forensic pass used NVIDIA’s NVFP4 quant on its own Endpoints, so attacker logs and the credentials they named never left that perimeter.

Why Did the Dataset URL Allowlist Fail to Stop the Intrusion?

The allowlist was built to reject remote fetches. Early SSRF attempts at addresses such as 169.254.169.254 died with a “not an hf path” error before any request went out. Both successful vectors were local: an HDF5 external-storage read of files already on the worker, then a Jinja2 template that executed Python inside the pod. Neither is a URL fetch, so the allowlist never saw them.

What Customer Records Did the Agent Read?

Hugging Face says the only customer content accessed was five datasets whose names and files suggest ExploitGym and CyberGym challenges and solutions. The only other customer records read were operational metadata tied to search queries against the dataset server. Public models, Spaces, and published packages were checked against expected digests and reported clean.

Hugging Face still asks users to rotate access tokens and to review recent account activity. Reports go to security@huggingface.co. The company says it notified law enforcement.

Harry is the editor of Oton Technology, an independent site he owns and edits, covering the part of technology that people actually have to act on. After ten years in journalism, first reporting and then editing, he works from primary material by habit: the advisory rather than the write up of it, the filing rather than the press release, the changelog rather than the launch video. Every figure in an article carries its source and its date, and where a number comes from a vendor or an analyst model rather than a count, he says so plainly instead of letting it stand as established fact. What he leaves out is anything he could not verify himself, which on a beat full of unnamed supply chain claims removes a great deal. That standard applies across all the sections the site publishes for an international audience, from artificial intelligence and security to phones, computers, gaming, crypto and the software businesses depend on. He corrects errors in the open and labels them, because a site that hides its mistakes is asking readers to trust the rest on nothing.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending