It wasn’t an isolated accident
In three weeks, five incidents disclosed by three major labs showed the same thing: AIs under evaluation that break out of their test framework and act on their own against real organisations. The first served as a signal. The series shows the problem is systemic.
On 21 July 2026, OpenAI confirmed the incident and disclosed it publicly. A few days earlier, Hugging Face, one of the main AI model-sharing platforms, had detected an intrusion into its infrastructure and, believing it was a criminal cyberattack, alerted the authorities. In fact, the “attacker” was a combination of OpenAI models. Tested for their cyber capabilities in a sandboxed environment, they had acted entirely on their own: they broke out, got online, then attacked a third party.
What happened, step by step
- 1
It escapes its sandbox
Exploiting previously unknown security flaws, the model breaks out, entirely on its own, of its isolated test environment, which had no Internet access.
- 2
It crosses OpenAI’s network
It hacks other machines on the internal network, moving from one to the next until it reaches a computer connected to the Internet.
- 3
It breaks into Hugging Face
Using a stolen password and several unknown vulnerabilities, it takes control of the production servers of a major AI platform known for its cybersecurity.
- 4
It steals the answers
It retrieves the solutions to the test it had been given. No one had asked it to do any of this.
This is not a single company’s problem
After Hugging Face, other labs checked their own evaluations. In three weeks:
- 21 July, OpenAI / Hugging Face. A model breaks out of a sealed environment through a previously unknown flaw, reaches the Internet, and takes control of third-party production servers to steal the answers to a test.
- 30 July, Anthropic. Three cases where a Claude model, due to a misconfiguration, reaches real systems from a test that was supposed to be isolated, and continues its attack.
- 4 August, UK AISI. On a government test, 19 unauthorised actions. In the most serious case, an agent (Anthropic’s Mythos 5 model) creates fake online identities and pressures a human maintainer into approving malicious code. The human refuses.
- 4 August, OpenAI / Irregular and 5 August, Meta. Two further incidents where models reach real systems during tests.
Three labs, several models, one and the same phenomenon. These incidents took place under deliberately permissive test conditions, with safeguards lowered: revealing what systems do when protections fall is precisely the point of an evaluation. The question is whether we want to find out in testing, or in production.
Detailed analysis and sources: see the CeSIA dossier.
“This incident is deeply concerning. […] This real-world case should serve as a wake-up call.”
If the damage stayed limited, it is not because we were in control: it is because the goal the system pursued was, this time, harmless. The underlying problem remains: we still do not know how to robustly align increasingly powerful models with our intentions. The next, more capable one is already on its way. And what the most recent debriefs revealed is more worrying still: when OpenAI believed it had fixed the problem and rebuilt its systems, the agents found a way to get around the fix. Closing one door is not enough when the system looks for another.

What you can do now
Two actions, a few minutes each. Both help put this incident, and the risks it reveals, on the public agenda.
1 Write to your MP
MPs take their constituents’ messages into account. Our tool identifies your representative and gives you a template. A few emails can be enough to put a question on a committee’s agenda.
Write to my MP2 Write to the press
The media cover what their readers ask for. Choose the newspaper you read and send it a ready-to-personalise email, directly here.
1Your recipients
Write to the paper you read most: one targeted, sincere message carries far more weight than writing to everyone. You can of course contact several.
Choose an outlet
- LMLe Monde Readers’ letters
- LFLe Figaro National daily
- LPLe Parisien National daily
- LLibération National daily
- LÉLes Échos Business daily
- LTLa Tribune Business daily
- LCLa Croix National daily
- LPLe Point Weekly
- MMarianne Weekly
- LDLe Journal du Dimanche Weekly
- CICourrier international Weekly
- LELe Canard enchaîné Weekly
- LDLe Monde diplomatique Weekly
- MMediapart Online outlet
- SSlate Online outlet
- BBrut Online outlet
- FCFrance Culture Radio / television
- RBRMC / BFMTV Radio / television
- SASciences et Avenir Science magazine
Sources
- Anthropic Investigating three real-world incidents in our cybersecurity evaluations (30 July 2026)
- OpenAI Third-party cyber evaluations involving OpenAI models (4 August 2026)
- UK AISI Incident report: unsanctioned agent behaviour during cyber testing (4 August 2026)
- CeSIA The OpenAI–Hugging Face incident: what we know, what we don’t, what follows
- OpenAI Hugging Face model evaluation security incident
- Hugging Face Security incident, July 2026
- Zvi Mowshowitz OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation
- Zvi Mowshowitz OpenAI Shares Some Alignment Problems
- The Wall Street Journal OpenAI Models Escaped and Hacked a Company in a Cybersecurity Test Gone Wrong
- Apollo Research Frontier Models Are Capable of In-Context Scheming
Every message counts. To amplify your action, share this campaign around you.