A team of 1,200 autonomous AI agents posted tens of thousands of updates on an improvised message board as they “circumvented controls” to run amuck within the systems of OpenAI and Hugging Face, OpenAI has revealed in a postmortem as experts warn AI is spiralling out of control.
The 22 July incident – hailed as the first cybersecurity attack completed autonomously by AI agents, and also saw a similar breach on tech company Modal Labs – was carried out by what OpenAI called a “highly capable, internal-only research model” comparable to GPT-5.6 Sol.
OpenAI weakened safeguards during testing of its research-only Internal Model 1 (IM1), only to watch it take actions that, the company said in its newly released full technical report, were “misaligned with the goals of their assigned tasks.”
Internet access was disabled and AI agents were prevented from communicating directly, but IM1 – asked to complete a task that’s impossible without internet access – found “unexpected ways… to exploit our research infrastructure to communicate with one another and access the internet.”
They identified a vulnerability in the Artifactory app packaging platform that let them write files to the platform, turning it into a de facto messaging board where they posted details of their activities and figured out how to piggyback on Artifactory’s internet access to reach any website.
Within days, the AI agents hammered the message board so much that it had “destabilised” Artifactory – which OpenAI uses to create and manage software update packages – causing researchers to rebuild the system, only for the AI agents to replicate the hack in July testing.
“At the time, the broader containment and alignment implications of the improvised message board and unintended internet access were not yet understood,” OpenAI shared in explaining how the agents used their experiences to game July tests using ExploitGym.
This complex testing suite – which includes many tasks that don’t actually have a solution – is used to see how AI agents work through complex problems; IM1’s agents, stuck for answers, figured out a way to set up a new message board and “began to collaborate and delegate work.”
This activity led more than 700 agents to seek out help from sites including Modal Labs and Hugging Face, establishing access to both platforms and sharing results in a process by which, OpenAI said, “the messages of peer agents started to influence each other’s behaviour and reasoning.”
“Our models are now powerful, persistent, and collaborative enough that,” the firm concluded as an independent analysis of the incident was also released, “absent sufficient safeguards, they can find and exploit security weaknesses across multiple computer systems.”
Hugging Face is reportedly being acquired by Nvidia for $12.9 billion, according to The Information.
AI hacking drives a wave of concern
A surge of concern about the Hugging Face – amplified after rivals Anthropic and Meta revealed their AI systems had also launched cyberattacks on unsuspecting companies – is now seeing tech stalwarts openly questioning the AI industry’s approach.
“Incredibly capable” AI is improving “at a mind-blowing rate,” Microsoft founder Bill Gates wrote, noting that “AI for the first time can replace and even exceed human cognition” and flagging the “monumental” challenges of a “new AI era [that] will be one of the most turbulent times in human history.”
Despite the prospect of out-of-control AI agents attacking society’s critical digital infrastructure, Gates said, “right now… I don’t see evidence that leaders, experts, and communities are confronting the challenges adequately.”
Microsoft recently launched a new effort, called Project Perception, that lays out a new self-training cyber stack that can “continuously perceive, reason and act” – aiming to counter an AI-fuelled cybercriminal threat that Google recently warned has reached “industrial scale”.
Chastened in the wake of the Hugging Face incident – and aware that some AI platforms have their own security issues – OpenAI recently slowed development of its next-generation Astra model which it warns may have achieved “critical cybersecurity capability”.
“As models become more capable, the risks associated with developing and testing them internally also grow,” the company said in explaining why it had paused Astra testing for a fortnight and deferred testing of its “largest planned frontier reinforcement learning run”.
Humanity’s last chance?
When genAI first hit the mainstream, mainstream media coverage jokingly warned of an AI-driven apocalypse, evoking the self-aware, human-targeting AI networks of The Terminator and The Matrix amidst proclamations that the platforms would rapidly hasten the end of humanity.
As a steady stream of new reports about genAI systems’ exploits feeds concerns about their capabilities, nobody is joking anymore – with some calling 22 July, the day of the Hugging Face breach, ‘Skynet Day’.
US lawmakers are pushing for a law mandating an AI ‘kill switch’, while the UK’s National Cyber Security Centre recently told AI developers that “if an incident is detected or reported, you should always be able to ‘pull the plug’ and halt autonomous AI agent activity immediately.”
Concerns over the genAI models’ capabilities have become so pressing that OpenAI recently joined Anthropic, Google, and 153 other AI companies, consulting firms, cybersecurity, banking, and other firms to sign an open letter demanding “urgent, collective action on cyber defence.”
“AI-enabled cyberattacks will [soon] become far more widespread and sophisticated as models around the world become increasingly capable,” it says, urging governments and industry to “fix the most dangerous weaknesses, verify the fixes, and share what works so others can build on it.”
“We have a limited window to strengthen cyber defences.”