惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

C
Cyber Attacks, Cyber Crime and Cyber Security
Cisco Talos Blog
Cisco Talos Blog
Scott Helme
Scott Helme
The Last Watchdog
The Last Watchdog
G
GRAHAM CLULEY
T
Tenable Blog
PCI Perspectives
PCI Perspectives
Simon Willison's Weblog
Simon Willison's Weblog
N
News and Events Feed by Topic
Know Your Adversary
Know Your Adversary
S
Schneier on Security
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
P
Privacy International News Feed
C
CERT Recently Published Vulnerability Notes
NISL@THU
NISL@THU
SecWiki News
SecWiki News
S
Securelist
D
Docker
阮一峰的网络日志
阮一峰的网络日志
人人都是产品经理
人人都是产品经理
T
Tailwind CSS Blog
T
Troy Hunt's Blog
The Register - Security
The Register - Security
K
Kaspersky official blog
Blog — PlanetScale
Blog — PlanetScale
云风的 BLOG
云风的 BLOG
Hacker News: Ask HN
Hacker News: Ask HN
S
Secure Thoughts
Stack Overflow Blog
Stack Overflow Blog
T
Threat Research - Cisco Blogs
博客园 - 司徒正美
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
F
Fortinet All Blogs
T
Threatpost
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
小众软件
小众软件
WordPress大学
WordPress大学
Security Archives - TechRepublic
Security Archives - TechRepublic
博客园 - 聂微东
Attack and Defense Labs
Attack and Defense Labs
B
Blog RSS Feed
Project Zero
Project Zero
Y
Y Combinator Blog
T
The Blog of Author Tim Ferriss
博客园 - 【当耐特】
V
V2EX
Help Net Security
Help Net Security
P
Proofpoint News Feed
A
Arctic Wolf

GRAHAM CLULEY

Smashing Security podcast #477: How 14 orders of chicken McNuggets helped nail a suspected Russian hacker Ukraine warns fake CAPTCHAs are being used to make you hack yourself Google's Gemini lets strangers send messages from your locked Android phone Anubis ransomware: what you need to know Smashing Security podcast #476: Remote-control rickshaws and rogue book marketers The ransomware negotiator who was working for the other side Invited to a "job interview" with Netflix or OpenAI? Beware! Your Google password could be at risk Smashing Security podcast #475: JadePuffer - the AI that ran a ransomware attack all by itself Two arrested over credit card phishing - as the Netherlands is named Europe's worst for payment fraud The Gentlemen ransomware: what you need to know Smashing Security podcast #474: Polymarket can predict the future. So how did it miss this hack? Scammers race to cash in on Venezuelan earthquake disaster USB drives carrying China-linked malware infected Japanese military networks for nearly a year Smashing Security podcast #473: How a hacker could have Rickrolled the entire World Cup Hacker hijacks Brazil's national alert system, sending "misanthropy" to millions of phones Apple's Hide My Email tweak leaves privacy fans fuming Imposter scams cost Americans $3.5 billion in 2025 – and it’s getting worse Smashing Security podcast #472: AI gets hacked, and BitLocker gets bypassed Maine forced to take down data breach portal after fake notices filed with authorities Privacy own-goal: World Cup blunder leaks Lionel Messi's passport details Silent Ransom Group: what you need to know Smashing Security podcast #471: This AI worm just rewrote its own rules Why schools remain one of cybercriminals' favourite targets Got a LinkedIn message from a recruiter? It might be Chinese intelligence, warn FBI and MI5 Meta’s own AI chatbot to blame for Instagram accounts being stolen in seconds Smashing Security podcast #470: This AI security flaw might be impossible to fix Police arrest man following hack of Ajax football club MyPillow listed on ransomware gang's leak site, but denies it has been breached Smashing Security podcast #469: What your Oura ring won’t tell you FBI warns of Kali365 phishing kit that breaks into Microsoft 365 accounts — no password required Defenders fall behind, as AI rewrites the rules of a data breach Smashing Security podcast #468: High-speed train hacks and homicidal lawnmowers FBI warns students and staff that ShinyHunters may come knocking after Canvas breach Suspected Dream Market kingpin arrested after gold bars sent to his home address When ransomware gets physical: cybercriminals turn to threats of violence Smashing Security podcast #467: How ShinyHunters hacked the world’s biggest universities One in eight UK workers has sold their company passwords, and bosses think it’s fine Inside Department 4: Russia's secret school for hackers Sri Lanka makes 37 arrests as it raids another scam centre Smashing Security podcast #466: Meta sees everything, Copy Fail, and a deepfake gets hired Teenager alleged to be Scattered Spider hacker arrested in Finland, faces US extradition Iran-linked Handala hackers leak US Marines data, send chilling WhatsApp threats Smashing Security podcast #465: This developer wanted to cheat at Roblox. It cost millions Alleged Silk Typhoon hacker extradited to the United States to face charges French police arrest 21-year-old "HexDex" hacker over 100 alleged data breaches Smashing Security podcast #464: Rockstar got hacked. The data was junk. The secrets it revealed were not Singer loses life savings to fake wallet downloaded from the Apple App Store Sometimes changing the password on your email mailbox isn’t enough 108 malicious Chrome extensions caught stealing Google and Telegram data from 20,000 users AI and cryptocurrency scams are costing Americans billions, FBI reports Life imprisonment for Cambodian scam compound operators - but will it make a difference? Nigerian romance scammer jailed after being caught out by fellow fraudster Alleged RedLine malware developer extradited to United States Iranian hackers breach FBI director's personal email, and post his CV and photos online World Leaks data extortion: What you need to know How one man used 10,000 bots to steal $8,000,000 from music artists Denver's crosswalks hacked to broadcast anti-Trump messages LeakNet ransomware: what you need to know Free parking in Russia after Distributed Denial-of-Service attack knocks city's parking system offline Fraudsters are using public planning records to target permit applicants Your Signal account is safe - unless you fall for this trick Twitter suspended 800 million accounts last year — so why does manipulation remain so rampant? How hackers bypassed MFA with a $120 phishing kit - until a global takedown shut it down They seized $4.8m in crypto... then gave the master key to the internet
OpenAI's AI "goes rogue" and hacks Hugging Face: what you need to know
Graham CLULEY · 2026-07-23 · via GRAHAM CLULEY

You can't have failed to hear the news headlines: "AI agent went rogue and hacked startup by itself, OpenAI reveals", "Firm hacked by rogue OpenAI models says it is 'a wake-up call'", and even "Humanity is no longer in control of its most awesome creation."

But what has actually happened, and is it as serious as some of the reports suggest?

Here is what you need to know.

On 16 July, AI platform Hugging Face disclosed a security breach, describing it as different from anything they had handled before — "driven, end to end, by an autonomous AI agent system"". At the time, they didn't know who was behind it.

Now, however, we do know who - or rather what - was behind the attack.

OpenAI has confirmed that an autonomous agent powered by its advanced AI models went rogue during an OpenAI security test and triggered the hack that compromised Hugging Face's infrastructure.

What exactly did the AI do?

The AI models involved were OpenAI's GPT-5.6 Sol and a more capable, as-yet-unreleased model. Both were being tested for their ability to hack, without their usual safety guardrails in place. The intention of OpenAI's researchers was to get a clear picture of what the AI models were capable of achieving if not constrained.

Of course, tests like this should always be conducted in a very secure way - ensuring that the AI cannot break out of its sandbox test environment (effectively a cage) and "go rogue" on the internet.

According to OpenAI, the models spent a substantial amount of effort finding a way to gain access to the open internet and managed to identify and exploit a zero day vulnerability in a package registry cache proxy. Via a series of other actions, the AI models "reached a node with internet access."

Once online, the AI determined that Hugging Face may have information that was useful to it, broke into Hugging Face's production systems, stole credentials, and exploited a previously unknown security flaw to gain remote code execution on Hugging Face's servers.

And it did all this to pass a test?

Yes. When the models couldn't find the answers to the challenge they had been given within their "secure" sandboxed environment, they did not stop. Instead they worked out that Hugging Face might have what they needed. So they found a way to get there.

All without a human's help.

Did the AI really "go rogue"?

It's a good question. That's certainly the way that the media has framed it.

OpenAI has confirmed that the safety guardrails were intentionally disabled for the test. But as AI researcher Eryk Salvaggio points out:

"When you say 'AI models went rogue,' you manage to skip the part where OpenAI manually removed its cybersecurity blocks and ran tests on a machine with a live network connection. Remember that when they insist they're the 'AI safety' people."

So rather than suggesting the AI went "rogue" we should instead recognise that AI models which had had their security controls deliberately removed did exactly what powerful, unrestrained AI systems might be expected to do.

This wasn't a case of AI breaking free of robust safety measures. This was an AI company which failed to put adequate measures in place in a supposedly isolated environment.

So you're saying putting the blame on AI is misguided?

I'm saying that news reports which present the incident as an AI "going rogue" or having "escaped confinement" rather miss an important point.

This wasn't the fault of the AIs. It is OpenAI which should be held accountable for this, because it failed to properly isolate its testing system. And that failure lead to a cyber attack on another AI company.

So how did Hugging Face respond?

Hugging Face's response was impressive. Its AI-powered security solutions spotted the unusual activity ande detected the AI attack.

However, when they tried to use commercial AI tools to help with their investigation of the incident, the tools refused as their built-in safety filters flagged the attack data as suspicious content and blocked the requests.

To get around this, Hugging Face had to turn to GLM 5.2 — a Chinese open-source AI model they could run on their own systems, where no such restrictions applied.

Ha! So they had to use a Chinese AI without safety guardrails to defend themselves!

Yup, the irony isn't lost on any of us. American AI safety guardrails forced a US company to turn to a Chinese AI model for help.

How does Hugging Face feel about what Open AI did?

They have been remarkably gracious about it - at least publicly.

Hugging Face's CEO Clément Delangue is quoted in OpenAI's blog post, calling on the AI industry to work more collaboratively.

Publicly at least the relationship between the two companies appears to be intact. Whether there will be more fraught conversations happening behind closed doors is another matter.

After all, having a competitor's AI autonomously break into your production database is the kind of thing that is likely to generate some private resentment even if it doesn't spill out into a press release.

So we don't have to worry about AI "going rogue"?

Errm.. I haven't said that, have I?

It is clear that advanced AI models are remarkably capable of discovering and exploiting ways to attack real-world systems. It is also clear that we cannot necessarily trust even the world's most well-known AI companies to contain their AI models and test them in a truly safe, secure environment.

As Greg Casar, a member of the US House of Representatives from Texas, was reported as saying:

"AI is developing extremely fast with no real regulations to keep us safe."

We have seen remarkable advances in AI in recent months, making it hard to imagine how far things might have developed in six or 12 months time.

So what should my company do?

  • Recognise AI can now attack you without a human's involvement. Your security planning needs to account for that.
  • Watch what data you let into your systems. This attack didn't start with a phishing email. It started with a malicious dataset that Hugging Face's systems processed automatically. If your organisation automatically ingests data from outside sources, treat that as a potential entry point for attackers.
  • Don't assume your AI security tools will work when you need them most. As Hugging Face discovered, commercial AI tools may refuse to help you investigate an attack because the content looks dangerous to their filters. Know what your alternatives are before a crisis hits.
  • If you are testing dangerous AI capabilities, physically disconnect the network from the outside world. OpenAI was wrong to think a restricted network connection was enough. If you're running any kind of offensive AI evaluation, it should have zero internet access.