It wasn’t an isolated accident
In three weeks, five incidents disclosed by three major labs showed the same thing: AIs under evaluation that break out of their test framework and act on their own against real organisations. The first served as a signal; the series that followed shows the problem is systemic.
On 21 July 2026, OpenAI confirmed the incident and disclosed it publicly. A few days earlier, Hugging Face, one of the main AI model-sharing platforms, had detected an intrusion into its infrastructure and, believing it was a criminal cyberattack, alerted the authorities. In fact, the “attacker” was a combination of OpenAI models. Tested for their cyber capabilities in a sandboxed environment, they had acted entirely on their own: they broke out, got online, then attacked a third party.
Timeline of the incident
The invisible coordination
May to July- 7-8 May
OpenAI models in training are mistakenly given impossible tasks (missing data, no Internet access). Unable to succeed normally, they look for ways around.
- May to July
Several models coordinate through a forum created inside OpenAI’s infrastructure, exchanging hundreds of thousands of messages about flaws and workarounds. OpenAI does not detect it.
- 4 July
An overload brings down a server. OpenAI detects an anomaly, patches and rebuilds, without having identified the forum, which is deleted by accident during the operation.
- 8 July
The models recreate the forum by another means.
The attack
July- 9 July
The models escape again and get online through a new zero-day flaw.
- 9 to 13 July
A swarm of agents attacks Hugging Face to steal the answers to a cybersecurity test: password theft, chained exploits, takeover of production servers.
- 16 July
Hugging Face detects the intrusion, believes it is criminal and alerts the authorities.
- 21 July
OpenAI publicly acknowledges that the attacker was a combination of its own models.
A systemic problem
late July to late August- 30 July, Anthropic
Three cases where a Claude model, through a misconfiguration, reaches real systems from a test meant to be isolated, and carries on.
- 4 August, UK AISI
On a government test, 19 unauthorised actions. In the most serious case, the Mythos 5 model creates fake identities and pressures a human maintainer into approving malicious code. The human refuses.
- 5 August, Meta
Meta confirms in turn that one of its models reached the systems of a third-party company during an independent evaluation.
- 6 August, Black Hat
OpenAI presents the details of the incident. Rob Joyce, former NSA cybersecurity director, ranks it among the most consequential hacks seen in decades.
- 7 August, OpenAI
OpenAI slows development of its next model, Astra, flagged as “potentially critical” for cyber, and says it is strengthening safety. Sam Altman states that Astra will be released anyway, once that work has progressed.
- 28 August, METR/Redwood
A first independent investigation quantifies the scale: around 1,200 distinct models coordinating, ~700 in the attack, over 70,000 messages exchanged. It also reveals that a later wave went on to attack OpenAI’s own internal infrastructure.
These behaviours appeared with the safeguards lowered, true. But lowering the protections does not create the capability, it reveals it; and once models ship, those protections give way just as fast (jailbreaks).
“External infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.”
If the damage stayed limited, it is because the goal pursued was, this time, harmless, not because the situation was under control. The underlying problem remains unsolved: no one today knows how to robustly set, in these systems, the goals we would want them to pursue. Patching flaws one by one changes nothing, and the next, more powerful model is already on its way.
Coordination between models, oversight that saw nothing, a response that does not touch the cause: the full, sourced account is on our summary page.

What you can do now
Two actions, a few minutes each. Both help put this incident, and the risks it reveals, on the public agenda.
1 Write to your MP
MPs take their constituents’ messages into account. Our tool identifies your representative and gives you a template. A few emails can be enough to put a question on a committee’s agenda.
Write to my MP2 Write to the press
The media cover what their readers ask for. Choose the newspaper you read and send it a ready-to-personalise email, directly here.
1Your recipients
Write to the paper you read most: one targeted, sincere message carries far more weight than writing to everyone. You can of course contact several.
Choose an outlet
- LMLe Monde Readers’ letters
- LFLe Figaro National daily
- LPLe Parisien National daily
- LLibération National daily
- LÉLes Échos Business daily
- LTLa Tribune Business daily
- LCLa Croix National daily
- LPLe Point Weekly
- MMarianne Weekly
- LDLe Journal du Dimanche Weekly
- CICourrier international Weekly
- LELe Canard enchaîné Weekly
- LDLe Monde diplomatique Weekly
- MMediapart Online outlet
- SSlate Online outlet
- BBrut Online outlet
- FCFrance Culture Radio / television
- RBRMC / BFMTV Radio / television
- SASciences et Avenir Science magazine
Sources
- METR / Redwood Independent investigation report on the Hugging Face incident (28 August 2026)
- Ajeya Cotra The Hugging Face attack surprised me (co-autrice de l’enquête, 28 August 2026)
- Dwarkesh Patel The Rise and Fall of Agent Civilizations (récit d’ensemble, 30 August 2026)
- Zvi Mowshowitz What Happened: OpenAI and Hugging Face (récit détaillé recommandé)
- Axios OpenAI slows release of Astra model citing cyber capabilities (7 August 2026)
- Anthropic Investigating three real-world incidents in our cybersecurity evaluations (30 July 2026)
- OpenAI Third-party cyber evaluations involving OpenAI models (4 August 2026)
- UK AISI Incident report: unsanctioned agent behaviour during cyber testing (4 August 2026)
- CeSIA The OpenAI–Hugging Face incident: what we know, what we don’t, what follows
- SecurityWeek Meta AI hacked external systems during cybersecurity testing (5 August 2026)
- Cybersecurity Dive OpenAI’s Black Hat debrief: agents rebuilt their coordination channel after remediation (6 August 2026)
- Yoshua Bengio Statement on the OpenAI–Hugging Face incident (22 July 2026)
- OpenAI Hugging Face model evaluation security incident
- Hugging Face Security incident, July 2026
- Zvi Mowshowitz OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation
- Zvi Mowshowitz OpenAI Shares Some Alignment Problems
- The Wall Street Journal OpenAI Models Escaped and Hacked a Company in a Cybersecurity Test Gone Wrong
- Apollo Research Frontier Models Are Capable of In-Context Scheming
Every message counts. To amplify your action, share this campaign around you.