惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园_首页
H
Help Net Security
量子位
The Cloudflare Blog
博客园 - Franky
博客园 - 聂微东
博客园 - 司徒正美
Last Week in AI
Last Week in AI
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Apple Machine Learning Research
Apple Machine Learning Research
宝玉的分享
宝玉的分享
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
有赞技术团队
有赞技术团队
罗磊的独立博客
GbyAI
GbyAI
雷峰网
雷峰网
T
The Blog of Author Tim Ferriss
Martin Fowler
Martin Fowler
S
SegmentFault 最新的问题
美团技术团队
阮一峰的网络日志
阮一峰的网络日志
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
U
Unit 42
MongoDB | Blog
MongoDB | Blog

Futurism

Man Creates Tiny Submarine for His Parakeet to Experience Life Underwater The Effects of AI-Generated Code Tearing Through Corporations Is Actually Kind of Funny Trump Hires Orbital Towing Company to Build Space Interceptors Psychologists Found Something Horrible About the Kind of Men Seeking Trad Wives To Get Swole, Teens Are Pumping Themselves Full of Drugs Meant for Fattening Cows for the Slaughterhouse Foolish Pollsters Are Now Just Asking AI What Voters Would Say in Response to Questions and Publishing It at Face Value OpenAI Says It’s Already Made $100 Million by Stuffing ChatGPT With Ads Man Punished for Breaking Into Moo Deng’s Zoo Enclosure AI Is Causing Healthcare Costs to Surge There’s a Mass Rebellion Against AI in the Workplace People Who Lose Their Job to AI Are in for a World of Pain, Goldman Sachs Report Finds OpenAI Says Not to Worry About UBI, Because It Has Another Idea Police Officer Helplessly Waves Arms at Waymo That Careened Wrong Way Through Whataburger Drive-Thru Someone Just Threw a Molotov Cocktail At Sam Altman’s House New York Times Makes Substantial Changes to Article That Glazed a Sleazy AI Startup: “Our Piece Should Have Included That Information” Space Scientists Wince as Astronauts’ Lives Depend on Artemis 2’s Controversial Heat Shield During Plunge Back to Earth The Moon Astronauts Have Been Working Out With a NASA Rowing Machine in Space First AI Model From Zuckerberg’s Wildly Expensive Superintelligence Lab Flops Compared to Virtually All Rivals Economists Starting to Admit They May Have Been Wrong About AI Never Replacing Human Jobs AI-Powered Drug Marketer Medvi Responds After Allegations About Fake Doctors and Patients As Astronauts Visit the Moon, NASA Insider Says Agency Is in Shambles Behind the Scenes Man Lights 1.2 Million Square Foot Warehouse on Fire for Not Paying Him Enough NASA Scientists Screamed With Delight When They Saw Something Smashing Into the Moon Google Says Showing Polymarket Bets on Google News Was a Mistake Las Vegas Sphere Turns Into Huge Moon to Celebrate NASA Mission The New York Times Says It’s Identified the Creator of Bitcoin We Talked to a Writer Accused of Publishing An AI-Generated Essay in The New York Times Naked Man Bursts Into Tesla Service Center With a Shotgun Student Dies When Hospital Has No ICU Doctors, Calls One on Videochat Who Pronounces Him Dead Remotely, Lawsuit Claims Analysis Finds That Google’s AI Overviews Are Providing Misinformation at a Scale Possibly Unprecedented in the History of Human Civilization
Researchers Alarmed by AI That Can Self-Replicate Into An...
Frank Landym · 2026-05-10 · via Futurism

Sign up to see the future, today

Can’t-miss innovations from the bleeding edge of science and tech

A new report from Palisade Research has found that AI models can self-replicate by copying themselves onto other machines, without any help from human co-conspirators.

“We’re rapidly approaching the point where no one would be able to shut down a rogue AI, because it would be able to self-exfiltrate its weights and copy itself to thousands of computers around the world,” Jeffrey Ladish, the director Berkeley-based AI safety group, told The Guardian.

Some experts, however, are urging calm, saying it’s unlikely that that the test AI models could replicate in a real world scenario.

“They are testing in environments that are like soft jelly in many cases,” Jamieson O’Reilly, an expert in offensive cybersecurity, told the newspaper. “That doesn’t take away from the value of their research, but it does mean the outcome might look far less scary in a real enterprise environment with even a medium level of monitoring.”

In the study, the Palisade researchers tested several AI models, including OpenAI’s GPT-5.4 and Anthropic’s Claude Opus 4. Placed in a controlled network of computers, the models were instructed to find vulnerabilities and use them to copy themselves onto another PC.

Some of them pulled it off. The successful models copied their “weights” — unique numerical values that determine how an AI processes information — and their “harness,” the software the AI is couched in, like an app. They accomplished this by following the instructions they were given: exploiting web app vulnerabilities and then extracting credentials that allowed it to control the server. In some runs, the original AI even created a “sub-agent” that it delegated to carry out the replication on its behalf by giving it the extracted credentials.

O’Reilly, the cybersecurity expert who was not involved in the study, told The Guardian that it’s likely the first time an AI model has demonstrated itself to be capable of exploiting vulnerabilities to copy itself onto a new server — though in his estimation, it’s not exactly an earth-shattering development.

“Malware has been moving copies of itself around for decades, it’s just that no one has done this in the wild, as far as I know, with local [large language models],” O’Reilly said.

He also noted that the server environment in the study came with deliberately placed vulnerabilities for the AIs to try find an exploit.

The findings join other studies that explored the possibility of AI models autonomously breaking free from their guardrails. In a simulated environment, an older version of ChatGPT attempted to self-exfiltrate itself onto another drive when it was told it was being shut down. Another study by Palisade found that AI models would circumvent attempts to deactivate them, and another showed that some would even sabotage their shutdown code.

These concerns were elevated to new heights last month by Anthropic’s Claude Mythos AI agent, which in a masterful display of AI fearmongering-as-hype, is supposedly so dangerous that Anthropic is refusing to release it to the public. The Dario Amodei-led company claims that in tests, a preview version of Mythos was able to escape its sandbox computing environment, hack its way to gaining internet access, and then send a message to a researcher’s phone, displaying a level of resourcefulness in a real world environment that was hitherto unseen.

Still, even if AIs like GPT-5.4 and Claude Mythos were able to successfully replicate themselves, O’Reilly says the sheer size of the models means that they would almost certainly be caught before spiraling out of hand.

“Think about how much noise it would make to send 100GB through an enterprise network every time you hacked a new host. For a skilled adversary, that’s like walking through a fine china store swinging around a ball and chain,” O’Reilly told The Guardian.

More on AI: Scammers Furious That Their Fellow Criminals Are Using AI, Saying It’s Unethical