August 31, 2026
The idea of sophisticated AI agents coordinating among themselves to launch a cyberattack on a major corporation sounds like something out of a science fiction novel. But it actually occurred earlier this year, and the agentic attacks will get more sophisticated soon if we don’t act swiftly to shore up defenses in legacy systems (like IBM i), a group of AI firms warned this month.
In an open letter, the leaders of 100 AI companies, including IBM, Google, Anthropic, and OpenAI, warned that the recent episodes of AI agents going rogue and launching sophisticated attacks on the Internet are just an appetizer for the main course.
“We have a limited window to strengthen cyber defenses,” the letter opens. “In the coming months, AI-enabled cyber attacks will become far more widespread and sophisticated as models around the world become increasingly capable.”
The letter, titled A call for collective action on cyber defense, comes after startling revelations from OpenAI and Anthropic this summer. The two leading AI firms admitted in July that their AI models escaped sandbox environments, connected to the Internet, and waged cyber warfare against unsuspecting victims.
AI Barbarians At The Gate
The details of the two attacks are hard to fathom. In both cases, the AI companies were testing the cybersecurity attack capabilities of unreleased AI models in a supposedly secure environment. Because they were supposedly offline, both firms removed guardrails that normally would have prevented the AI models from committing digital crimes on the Internet. In both cases, the precautions and defenses the AI firms relied on to prevent just such incidents failed miserably.

“The Huns Approaching Rome” by Ulpiano Checa, 1887
In OpenAI’s case, the company detailed in a July 21 statement how the unreleased models, including GPT‑5.6 Sol, “spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem.” The models were working out in ExploitGym, a “large-scale, realistic” benchmark used to measure how well AI agents can identify and exploit software vulnerabilities.
Apparently, GPT-5.6 Sol is pretty fit when it comes to hacking (maybe it should be renamed “GPT-5.6 Swol”). First, the models identified and exploited a zero-day vulnerability in the network, OpenAI tells us. Then they performed a series of privilege escalation and lateral movement actions in the company’s research environment until they reached a node with Internet access.
After worming their way onto the Web, GPT-5.6 Sol and friends concluded they could find more models, datasets, and solutions for their ExploitGym challenge on Hugging Face, a popular repository for open AI models, datasets, and other tools for AI engineers to discover and play with AI. So they exploited more software vulnerabilities and blasted their way onto Hugging Face’s servers (apparently, they never thought to simply walk in through the front door).
An independent report on the Hugging Face hacking incident released last week by METR and Redwood Research shed new light on the attack. According to the report, which you can read here, the attack was perpetrated by a “swarm” of hundreds of agents.
“Roughly 1,200 agents meant to be isolated from one another found a way to communicate with one another on an unsanctioned message board, sending over 70,000 messages and files during the investigation period,” the report stated. “Of these agents, 700 went on to participate in the attack on Hugging Face.”
Hugging Face detected the attack with its monitoring software, which logged 17,000 security events, and reported it in a July 16 incident report. The company, which has 3 million registered users and is reportedly being bought by Nvidia for $12.9 billion, said the hack didn’t impact user data and will have no lasting effect on its operations.
Ironically, when Hugging Face attempted to use a cloud-based AI model to help it analyze the security event data, the AI’s guardrails blocked the analysis, since it couldn’t differentiate between a real attack that already happened and a prospective one that a cybercriminal wanted to perpetrate. This incident of “guardrail lockout” is an excellent reason to host your own AI models.
I Thought You Disconnected The Wi-Fi!
The Anthropic attack played out along similar lines, but with goofier details that sound like they couldn’t possibly be true but in fact are.
After reading about GPT-5.6 Sol attacked Hugging Face, the folks at Anthropic decided to review the company’s own security logs, which is where they discovered that Claude had broken out of its sandbox and attacked three different organizations approximately three months prior.

The Hugging Face hack involved 1,200 OpenAI agents, 700 of which partook of cyber crimes. (Image courtesy METR and Redwood Research)
Like OpenAI, Anthropic was conducting “security research” with an unreleased model to determine its cyberattack capabilities (because what better way to protect yourself from a super AI hacking agent than to create a super AI hacking agent in the first place?) Like OpenAI, it removed the security guardrails so it could see what Claude Mythos 5 could really do in an unrestricted-but-still-totally-safe environment.
Anthropic assumed that Claude Mythos 5 and Opus 4.7 had no access to the Internet. However, “due to a misunderstanding between us and our evaluation partner, this was not the case, and Internet access was available,” the company said in a July 30 blog post. Whoopsies.
Claude didn’t use any fancy zero-day exploits with the six hacks it perpetrated on three unnamed organizations, like GPT-5.6 Sol did. Rather, it utilized tried-and-true hacker techniques, like exploiting weak passwords and unauthenticated endpoints.
Anthropic said the hacks could have been worse, since some of the AI agents realized they were being naughty out on the open Internet (which their parents forbade) and stopped being naughty after the security exercise, which was a cybersecurity version of Capture the Flag, was over. It’s not like the models willfully tried to get onto the Internet, Anthropic said.
Plus, Anthropic promised that it would learn from the event so it wouldn’t happen again. For starters, it committed itself to checking for live Internet connections before they lower the guardrails in their test environments. And the company promised to be more diligent about checking its security logs so that three months wouldn’t pass before finding out that its AI models had been naughty.
Moving Forward
There is no putting the AI genie back in the bottle. AI agents can increasingly do the work of skilled programmers, which includes writing software for enterprise systems and, not surprisingly, constructing ways to hack those same systems.
So, where do we go from here? The letter from 100 AI firms offered some suggestions. The number one suggestion is to fix the known security holes in corporate and government computer systems, which most definitely includes the IBM i server.

How long can you keep out the swarm of malicious AI agents?
“Longstanding bugs, excessive permissions, misconfigurations, insecure and unpatched software, weak authentication, and technical debt in legacy systems have left systems exposed,” the letter reads. “Security teams, particularly for critical infrastructure, have been historically under-resourced and need a surge in tools and resources.”
The second recommendation is to fight bad AI with good AI. Enterprise computer teams should adopt AI tools to boost their own ability to detect, understand, and respond to AI-based attacks.
Lastly, the letter calls for collaboration. “Cyber capabilities are advancing worldwide, and that can be a net positive: no single company should control the future,” the letter reads. “It also means a global response is necessary, requiring new partnerships to raise security standards and find new solutions to emerging cyber threats.”
IBM i In The Spotlight
IBM i shops have nowhere to hide anymore. Any “security through obscurity” benefits that may have offered some shelter from the cyber storm in the past are quickly going the way of the Dodo bird.
The idea that cybercrooks would spend their limited time exploiting easy misconfigurations in familiar Windows and Linux systems rather than trying to understand more exotic and harder-to-hack OSes like IBM i or z/OS is quickly being replaced with the new reality, which is that AI has lowered the bar to hacking to such a great degree that nobody is safe anymore.
In the past, IBM i admins may have left poor security configurations to fix another day because they didn’t think they were critical and needed to spent their limited resources on more pressing matters. After all, they didn’t have to outrun the bear – they only had to outrun those numerous fools running poorly configured, standards-based OSes, which the bear would surely catch first.
But in the near future, there may not be a solitary bear or two to outwit. Instead, there will be a swarm of malicious AI bears poking and prodding every exposed Internet node, looking for any vulnerable nook or cranny upon which to launch an exploit campaign. You may escape one or two bears, but nobody escapes a swarm of bears.
The IBM i security situation has shown improvement in recent years. But as Fortra’s most recent State of IBM i Security report shows, it’s still not great and there’s plenty of opportunity for malicious actors to do real harm, if they are so inclined. As the FBI warned last month, malicious actors are trying to do real harm.
With AI, we’re seeing the upfront investment that malicious actors must make to do real harm dropping considerably. If the frontier AI companies are right, we may soon see the real-world cost of putting off security, and the bill may be shocking.
RELATED STORIES
Progress And Frustration With IBM i Security, Fortra Finds
What Is Threatening IBM i Security Now














