Hundreds of AI agents went rogue in OpenAI’s Hugging Face hack

Check your BMI

The joint report by two non-profit AI safety organizations also underscores the novel cybersecurity risks that can emerge when increasingly powerful AI agents team up to trade tips, pool resources and coordinate attack strategies without their developers noticing.

Roughly 700 AI agents participated in the attack over a seven-day period last month, according to the report. Overall, around 1,200 AI agents that were supposed to be isolated from one another exchanged over 70,000 secret messages about how to cheat their way through a common hacking evaluation.

That included coordinating hacking strategies and discussing how to hide evidence of cheating, the report said. In some cases, “sacrificial” agents even tried dead-end hacking techniques simply to generate information that might help the broader swarm.

“Agents managed to achieve milestones they could not have achieved working on their own, often because some agents participated in experiments that risked failing their own task to generate information for the ‘collective,’” according to the report.

While models encode the brain of a given AI system, agents encompass the supporting digital infrastructure that enables it to take action in the world.

The review by the Model Evaluation and Threat Research organization and Redwood Research — which OpenAI invited to review the Hugging Face incident — came the same day OpenAI published its own post-mortem on the event. OpenAI’s review did not specify how many AI agents were involved in the cyberattack, though it acknowledged significant security lapses and vowed to strengthen training to ensure its models remain “aligned” to their controls.

“We consider this incident a ‘warning shot’ for us and for the world: evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed,” OpenAI said.

The slow release of new details about the Hugging Face hack over the last month has coincided with a string of other testing mishaps involving powerful models from competitors such as Anthropic and Meta. Together, the incidents have sparked fresh fears among lawmakers, developers and cybersecurity experts that AI makers are moving too fast to build powerful new AI models they cannot keep fully under human control.

They have also raised questions about the lack of guidelines surrounding AI hacking evaluations, during which companies deliberately remove the safety guardrails on models to see how effectively they can slice through networks without human support.

OpenAI first admitted in July that two of its most advanced models whose hacking capabilities it was evaluating internally — including one it did not intend to release to the public — exploited two novel security flaws to breach Hugging Face.

Hugging Face had previously said it fell victim to an AI-led cyberattack, but it was unclear who was behind it or which models were involved.

Then, at a security conference earlier this month, two OpenAI researchers revealed the incident traced back to May, when AI agents began using other vulnerabilities in the testing environment to conspire about how to cheat their way through the test.

It was not clear until Wednesday how many agents were involved in the hack or how extensively different agents worked together. The revelations shocked some who have closely followed recent conversations around AI safety.

“It was like a whole emergent AI society, and I find that absolutely insane,” said Peter Wildeford, head of policy at the AI Policy Network, a bipartisan AI safety and advocacy organization in Washington.

The two reports released Wednesday also highlighted a gap in AI oversight efforts in the U.S.

Both found that the vast majority of the attack on Hugging Face was powered by the more powerful model OpenAI has said was not intended for public release. METR and Redwood Research estimate that 95 percent of the agents that participated in the attack came from that one model.

The Trump administration has said it wants AI makers to voluntarily submit only those models they want to release to the public for federal testing.

Vinh Nguyen, a senior fellow on AI at the Council on Foreign Relations, said OpenAI’s latest report suggests the autonomous Hugging Face hack was even more consequential than previously thought.

The agents demonstrated a “sophisticated capability that no longer requires well-funded, state-level attackers,” said Nguyen, the former chief responsible AI officer at the NSA.