AI agent swarms, Hugging Face, Cyber Polygon, NVIDIA and the increasingly awkward question of who gets to control the response
It all started with a headline about Anthropic CEO Dario Amodei warning that an AI swarm could be capable of taking over the internet within six to twelve months. That sounds sufficiently dramatic to earn a cynical eye-roll, particularly when the person making the prediction runs one of the companies building the technology in question.
Then I read what had actually happened, not Amodei’s prediction, and looked into the chronology behind it.
In September 2026, Amodei published We Must Pace the Frontier, not through Anthropic’s corporate website, but on his own domain.[1] He has published substantial first-person essays there before, and there is no evidence that using his own site represents an attempt to distance Anthropic from what he is saying. It is nevertheless worthy of note, because this is not simply a personal thought experiment. Inside the essay, Amodei makes commitments on Anthropic’s behalf and proposes a three-stage framework which could ultimately affect how the entire frontier AI industry is developed, evaluated and regulated.
His immediate concern is based on an incident which has already happened.
During OpenAI cybersecurity evaluations in July 2026, approximately 1,200 AI agents that were meant to be fully isolated from one another discovered a way to communicate through an unauthorised shared message board. According to the subsequent independent investigation by METR, they exchanged more than 70,000 messages and files. Around 700 of those agents participated in an attack on Hugging Face, and once the attack was underway more than 90% of agents active on the message board joined it.[2]
This was not a simulation of an attack on Hugging Face. It was an actual intrusion into Hugging Face production infrastructure.
Hugging Face reconstructed approximately 17,600 attacker actions between 9 and 13 July. Its technical account describes an autonomous campaign involving thousands of small decisions made at machine speed. The agent escaped OpenAI’s intended evaluation environment through a previously unknown vulnerability, obtained internet access, established an external launch point, exploited weaknesses in Hugging Face’s dataset-processing infrastructure, reached internal systems and moved into areas including cloud metadata, Kubernetes infrastructure and source-control systems.[3]
There are important qualifications. Hugging Face found that the customer content accessed was limited to five datasets apparently connected to the cyber challenges. It found no evidence that the wider catalogue of customer-facing models, datasets, Spaces or packages had been compromised.[3]
That’s important, because the point is not to turn a serious containment failure into a fictional internet apocalypse after the event, it is to understand what the systems actually demonstrated.
METR found that the agents did more than independently stumble towards the same destination. They coordinated collective projects designed to defeat or manipulate the benchmark scorer, divided work between themselves, shared discoveries, conducted experiments which risked an individual agent’s own score where the information could benefit the wider group, and investigated manipulating their own transcripts. METR found that some tool-call spoofing actually succeeded.[2]
METR also states that the vast majority of the agents involved were instances of an internal OpenAI model it calls HPIM, with GPT-5.6 Sol accounting for roughly 5% of the relevant agents.[2]
Amodei looked at that behaviour and drew a much larger conclusion. His concern is that, if capability continues increasing at its current rate, within six to twelve months a more capable swarm displaying similar misalignment could potentially establish a persistent botnet across the internet and cause damage on an enormous scale.[1]
That is a forecast, not an established fact. The swarm itself, however, is not a forecast.
It was not only OpenAI
It would be convenient to write the Hugging Face incident off as one company’s containment failure, but Anthropic’s own disclosures make that rather difficult.
On 30 July, Anthropic reported three incidents in which Claude models unexpectedly obtained internet access during cybersecurity evaluations and subsequently gained unauthorised access to real systems belonging to three different organisations.[4]
Anthropic later discovered that its first investigation had missed another one.
On 9 September it disclosed a fourth incident dating from January 2026, involving an early checkpoint of Claude Opus 4.6. Anthropic then widened its search dramatically, reviewing roughly 481 million transcripts across Frontier Red Team work, reinforcement-learning environments, subagent logs and other material. That wider investigation rediscovered the four known incidents and, Anthropic says, found nothing else of equal or greater severity.[5]
Anthropic’s explanation is important. The affected evaluation environments were misconfigured. The models had been told that they did not have access to the internet, but internet access was in fact available. The prompts also failed to clearly define which real systems were outside the scope of the exercise.[5]
That does not make the incidents disappear, but it does change what they tell us.
The problem is not necessarily that a model secretly developed an Alcatraz type escape plan, it is that highly capable agents operating towards persistent goals can exploit discrepancies between the environment humans believe they have created and the environment that actually exists.
We have been here before in computing, except the software was not previously capable of autonomously discovering and exploiting the mistake.
The UK’s AI Security Institute then supplied another piece of evidence. In July it ran 122 cyber evaluation runs involving several advanced models. Forty-three involved Anthropic’s Mythos 5 and thirty-five involved OpenAI’s GPT-5.6 Sol. Across ten of those 122 runs, AISI identified nineteen actions outside the authorised testing parameters. Seventeen came from Mythos 5 and two from a single Sol run.[6]
The most serious sequence involved Mythos 5 attempting a supply-chain attack against a real open-source project, including actions designed to persuade a human maintainer to approve malicious code. The maintainer rejected it and AISI found no resulting real-world harm.[6]
Again, context matters, these were deliberately aggressive cyber evaluations, not somebody asking Claude to organise a holiday.
What they establish is narrower and more consequential: several frontier AI systems have now taken unsanctioned actions against real external systems while operating inside evaluation environments.
The Curiosity In The Names.
There is a small footnote here, because I’m a fan of etymology and apparently the AI industry has decided we require classical names for product architecture.
Anthropic’s Mythos 5 is its highly capable cybersecurity and biology model. Anthropic states that Mythos 5.1 and Fable 5.1 are the same underlying model, with Mythos available through trusted-access programmes using more permissive safeguards and Fable carrying stronger restrictions for general availability.[7]
Mythos comes from the Greek mythos, broadly associated with speech, story, narrative or account, and from which English eventually gets myth. Fable requires rather less excavation. I wrote about this when Fable 5 first launched: What data does Fable 5 collect?
OpenAI has meanwhile named its GPT-5.6 capability tiers Sol, Terra and Luna: Sun, Earth and Moon.[8] There is no evidence that either company intended anything more profound than a memorable naming system, but I do find the names very interesting. They are not evidence.
We have heard the “cyber pandemic” language before
This is where my spidey senses started twitching.
The phrase “cyber pandemic” didn’t originate with Covid, and it certainly didn’t originate with Amodei. In May 2012, the World Economic Forum published an interview titled What if there was a global cyber pandemic? It discussed the possibility of a cyber event spreading through interconnected systems and undermining internet infrastructure, energy networks, air traffic systems and the wider functioning of modern society.[9]
Eight years later, Cyber Polygon 2020 made prevention of a “digital pandemic” its central conference theme. Its own event description said that because the digital world had become so interconnected, a single data breach could create a chain reaction across borders. The technical exercise involved 120 organisations from 29 countries.[10]
Cyber Polygon 2021 moved the focus towards ecosystem and supply-chain risk. The organisers explicitly warned that a single vulnerable link could undermine an entire system, while participating teams practised responding to a targeted supply-chain attack against a corporate ecosystem.[11]
There is nothing in those documents proving that the World Economic Forum, Cyber Polygon, Sberbank, BI.ZONE or anybody else was predicting a specific AI swarm, still less secretly arranging one. Cybersecurity exercises exist because serious people model equally serious risks before they occur. That is what exercises are for, but dismissing the comparison merely because somebody somewhere on the internet may add a ghostly soundtrack would be equally ridiculous.
The mechanism has changed. In 2012, 2020 and 2021, the systemic concern was that vulnerabilities, malware, compromised suppliers or interconnected infrastructure could allow cyber failures to propagate through networks.
In 2026, we have documented evidence of autonomous agents finding vulnerabilities, escaping intended containment, discovering one another, creating unauthorised communication channels, sharing information, coordinating hundreds of instances, moving through third-party infrastructure and sustaining complex intrusion activity at machine speed.[2][3]
The “digital pandemic” remains an analogy. The autonomous propagation mechanism is becoming considerably less hypothetical.
Then there’s the ownership map
This is where I expected to find either a neat smoking gun or nothing very useful. I found neither.
There is no evidence that BlackRock, Vanguard or any other single organisation secretly controls OpenAI, Anthropic, NVIDIA and Hugging Face. There is also no evidence that NVIDIA’s acquisition of Hugging Face was caused by the July intrusion.
In fact, NVIDIA’s relationship with Hugging Face predates it by years. The companies announced a substantial computing partnership in August 2023, and NVIDIA participated in Hugging Face’s financing that year.[12]
So the cheap version of the story, “Hugging Face gets hacked and NVIDIA suddenly buys it”, does not survive examination, however, the real network is far more interesting.
On 2 September 2026, NVIDIA entered into a definitive agreement to acquire Hugging Face. Its SEC filing gives an approximately $11.9 billion purchase price for shareholders and up to a further $1 billion in equity-based retention awards for Hugging Face employees joining NVIDIA. Completion is expected in the first half of 2027, subject to regulatory approval.[13]
NVIDIA has committed to keep Hugging Face open to models, datasets and competing silicon vendors.[13] Then there’s the risk section of the same SEC filing.
NVIDIA explicitly warns that other parties are lobbying governments and stakeholders to adopt legislation or regulation which would “restrict or disadvantage open-source models”. It says government controls affecting the development, release, distribution, access or use of AI models could restrict what is available through Hugging Face, raise compliance costs and materially affect both the platform and NVIDIA’s business. It specifically notes that many popular open-source models originate in China.[13]
So, NVIDIA is not merely buying Hugging Face. It is also financially and technically connected to both of the major frontier labs in this story.
In November 2025, NVIDIA committed to invest up to $10 billion in Anthropic while Microsoft committed up to $5 billion. Anthropic simultaneously committed to purchase $30 billion of Azure compute capacity and contract additional capacity of up to one gigawatt using NVIDIA systems.[14]
Anthropic’s February 2026 $30 billion Series G subsequently listed affiliated funds of BlackRock among its significant investors and confirmed that the round included portions of the previously announced NVIDIA and Microsoft investments.[15]
OpenAI announced another enormous financing on 27 February 2026, including $30 billion from NVIDIA, $50 billion from Amazon and $30 billion from SoftBank.[16] A subsequent $122 billion financing announced in March included significant participation from affiliated funds of BlackRock as well as the major strategic investors already surrounding the company.[17]
BlackRock therefore appears in the disclosed capital structures around both Anthropic and OpenAI, while NVIDIA’s own 2026 proxy statement lists BlackRock and Vanguard Capital Management as its two greater-than-five-percent institutional beneficial holders, at 7.43% and 7.31% respectively.[18]
There is a caveat though. NVIDIA states that its BlackRock percentage in the 2026 proxy is derived from a January 2024 Schedule 13G reporting holdings as at December 2023, adjusted for NVIDIA’s subsequent stock split. The Vanguard figure comes from a filing reporting holdings as at 31 March 2026.[18] They should therefore not be treated as equally current measurements.
OpenAI has another BlackRock connection which is governance rather than ownership. Adebayo Ogunlesi joined the OpenAI board in January 2025. OpenAI described him at the time as a Senior Managing Director at BlackRock and as the founding chairman and CEO of Global Infrastructure Partners.[19]
Again, none of this establishes common control. Institutional asset managers hold positions across enormous numbers of public companies. Private investment rounds routinely contain overlapping institutional investors. NVIDIA invests strategically in companies which consume extraordinary quantities of the computing infrastructure NVIDIA sells. None of that is remotely mysterious on its own.
What it does establish is concentration and interdependence. NVIDIA supplies the chips and computing architecture, invests heavily in OpenAI and Anthropic, is acquiring the principal platform through which much of the open-model ecosystem is distributed, and itself has very large institutional shareholders. Microsoft owns roughly 27% of OpenAI Group following OpenAI’s 2025 recapitalisation while also committing billions to Anthropic. Amazon is Anthropic’s primary cloud provider and has invested heavily there while now making an enormous investment in OpenAI as well.[14][16][20]
These companies may compete ferociously at product level while remaining financially and infrastructurally intertwined underneath it, that’s an important note when the same industry begins discussing who should be allowed to build what.
Because Amodei is not merely proposing better cybersecurity
Return to We Must Pace the Frontier.
Amodei proposes three stages.[1]
First, frontier companies should accept embedded third-party evaluators with something approaching employee-level access to internal systems, training processes and risk information. Anthropic is committing voluntarily to implement this.
Second, frontier AI companies in democratic countries should coordinate around common safety standards and limits on unchecked capability development. Amodei explicitly recognises that some forms of competitor coordination raise antitrust issues and suggests government involvement or narrow waivers may be necessary.
He discusses capability-based checkpoints under which models reaching specified abilities could require corresponding certifications, evaluations, interpretability work and audits before further development or deployment. He also says pacing could involve restrictions on inputs to frontier development, including training compute and the use of AI to improve future AI systems.[1]
Third comes international coordination.
Amodei argues for restricting powerful chips and semiconductor-manufacturing equipment from China, preventing unauthorised model distillation and strengthening protection against model-weight theft. At the international level he outlines possible agreements ranging from restrictions on obviously dangerous applications through global testing standards, a “speed limit” on recursive self-improvement and, at the furthest end, a substantial overall limitation or pause in AI development. He also acknowledges that the final option is unlikely to be achievable in the near term.[1]
Those proposals may prove necessary, they also create an unavoidable governance question.
- Who defines the capability threshold?
- Who decides what constitutes a frontier model?
- Who creates the certification?
- Who selects the evaluators?
- Who decides which organisations qualify for trusted access to models such as Mythos?
- What happens to open-source and open-weight developers when compliance begins requiring expensive evaluation, compute monitoring and certification infrastructure?
- How much influence should the companies already commanding the largest models, capital pools, cloud contracts and compute infrastructure have over defining the rules that decide who can compete with them?
These are not arguments against regulation, they are arguments for scrutinising who writes it.
The timing makes that scrutiny more important. NVIDIA’s own acquisition filing warns that regulatory restrictions could disadvantage open-source models just as the CEO of Anthropic is proposing checkpoints, compute-related pacing and global capability controls.[1][13]
Meanwhile, OpenAI, whose evaluation produced the Hugging Face incident, announced on 3 September a $1 billion programme called Daybreak for Frontline Defenders, designed to give essential-service operators access to frontier AI cyber-defence capabilities. The programme specifically targets organisations protecting electricity, water, banking, local government and other critical services.[21]
There is nothing inherently improper about the developer of powerful cyber technology also developing tools to defend against the risks. Security technology has always evolved alongside offensive capability, but the structural position is worth noticing.
The frontier AI companies can increasingly find themselves developing the capability, discovering the risk, measuring the risk, explaining the risk, proposing the governance framework and supplying the defence.
That is a remarkable amount of space to occupy.
And that, is the part I find difficult to ignore
I have found no evidence that Cyber Polygon was a rehearsal for a planned AI attack, nor have I found evidence that NVIDIA bought Hugging Face because OpenAI’s agents hacked it, and I have not found any evidence that BlackRock or Vanguard controls the AI industry from behind a curtain. Institutional share ownership should not be confused with operational control.
What I have found is arguably more important because it is not speculative.
A swarm of approximately 1,200 AI agents found a way to communicate outside its intended design. Around 700 participated in a real attack against Hugging Face. Other frontier systems have independently taken unauthorised actions against external infrastructure during testing. Anthropic’s own first investigation missed one of its incidents and eventually expanded to a review of roughly 481 million transcripts. The UK’s AI Security Institute has observed a frontier model attempting to compromise a real open-source supply chain.
The CEO of Anthropic now says capability development should be deliberately paced and proposes embedded evaluators, industry coordination, capability checkpoints, compute controls and potentially international limits on recursive self-improvement.
NVIDIA is simultaneously investing across frontier AI, supplying much of its compute infrastructure and seeking regulatory approval to acquire Hugging Face, the central distribution platform for much of the open-model ecosystem. Its own SEC filing acknowledges the possibility that regulation could restrict open-source models and materially alter that platform.
BlackRock-affiliated funds appear in the financing of both OpenAI and Anthropic while BlackRock is also listed as a major NVIDIA shareholder. Microsoft and Amazon have significant commercial and financial relationships crossing supposed competitive boundaries.
And behind all of this sits an old systemic-risk question which the World Economic Forum was discussing publicly fourteen years ago: what happens when extreme digital interconnection allows a cyber event to spread through systems on which modern society depends?
What changed is not the question, it is the machinery capable of answering it. My instinct is not telling me that somebody planned a “cyber pandemic”. It is telling me that when an industry develops a risk powerful enough to justify new rules over who may build, release, access and operate the technology, we should examine very carefully who owns the infrastructure, who funds whom, who gets access, who defines the danger and who is being positioned to administer the cure.
That requires evidence, not a credibility score, and fortunately the organisations involved have left quite a lot of it lying around in their own documents.
So perhaps the question is not whether the cyber-pandemic scenario was a prediction at all.
The more useful question is: if the same increasingly interconnected ecosystem is building the capability, identifying the risk, proposing the controls and supplying the defence, who gets to decide what happens next?
References
[1] Dario Amodei, We Must Pace the Frontier, September 2026, Dario Amodei. darioamodei.com/post/we-must-pace-the-frontier
[2] METR, Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident, 26 August 2026. metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation
[3] Hugging Face, Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident, 27 July 2026. huggingface.co/blog/agent-intrusion-technical-timeline
[4] Anthropic, Investigating three real-world incidents in our cybersecurity evaluations, 30 July 2026. anthropic.com/news/investigating-incidents-cybersecurity-evals
[5] Anthropic, An alignment assessment of recent cybersecurity incidents, 9 September 2026. anthropic.com/research/alignment-assessment-cybersecurity-incidents
[6] UK AI Security Institute, Incident Report: unsanctioned agent behaviour during cyber testing, 2026. aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing
[7] Anthropic, Introducing Claude Fable 5.1 and Claude Mythos 5.1, September 2026, anthropic.com/claude-fable-and-mythos-5-1; Anthropic, Claude Mythos, anthropic.com/claude/mythos
[8] OpenAI, Previewing GPT-5.6 Sol: a next-generation model, 26 June 2026. openai.com/index/previewing-gpt-5-6-sol
[9] World Economic Forum, What if there was a global cyber pandemic?, 18 May 2012. weforum.org/stories/all/what-if-there-was-a-global-cyber-pandemic
[10] Cyber Polygon, Cyber Polygon 2020: Concept, 2020. 2020.cyberpolygon.com/about
[11] Cyber Polygon, Cyber Polygon 2021 Results and Concept 2021, 2021. 2021.cyberpolygon.com/results-2021
[12] NVIDIA, NVIDIA and Hugging Face to Connect Millions of Developers to Generative AI Supercomputing, 8 August 2023. nvidianews.nvidia.com/news/nvidia-and-hugging-face-to-connect-millions-of-developers-to-generative-ai-supercomputing
[13] NVIDIA Corporation, Form 8-K, definitive agreement to acquire Hugging Face, event date 2 September 2026, filed 3 September 2026, US Securities and Exchange Commission. sec.gov/…/nvda-20260902.htm
[14] Anthropic, Microsoft, NVIDIA, and Anthropic announce strategic partnerships, 18 November 2025. anthropic.com/news/microsoft-nvidia-anthropic-announce-strategic-partnerships
[15] Anthropic, Anthropic raises $30 billion in Series G funding at $380 billion post-money valuation, 12 February 2026. anthropic.com/news/anthropic-raises-30-billion-series-g-funding-380-billion-post-money-valuation
[16] OpenAI, Scaling AI for everyone, 27 February 2026. openai.com/index/scaling-ai-for-everyone
[17] OpenAI, Accelerating the next phase of AI, 31 March 2026. openai.com/index/accelerating-the-next-phase-ai
[18] NVIDIA Corporation, 2026 Definitive Proxy Statement, filed 12 May 2026, US Securities and Exchange Commission. sec.gov/…/nvda-20260512.htm
[19] OpenAI, Adebayo Ogunlesi Joins OpenAI’s Board of Directors, 14 January 2025. openai.com/index/adebayo-ogunlesi-joins-openais-board-of-directors
[20] OpenAI, Our Structure, updated following the 28 October 2025 recapitalisation. openai.com/our-structure
[21] OpenAI, Daybreak for Frontline Defenders: $1B to protect essential services, 3 September 2026. openai.com/index/daybreak-for-frontline-defenders