惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

月光博客
月光博客
IT之家
IT之家
Hugging Face - Blog
Hugging Face - Blog
J
Java Code Geeks
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
博客园 - 叶小钗
MyScale Blog
MyScale Blog
G
Google Developers Blog
Microsoft Azure Blog
Microsoft Azure Blog
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
大猫的无限游戏
大猫的无限游戏
博客园 - 三生石上(FineUI控件)
Google DeepMind News
Google DeepMind News
Engineering at Meta
Engineering at Meta
The Cloudflare Blog
Martin Fowler
Martin Fowler
酷 壳 – CoolShell
酷 壳 – CoolShell
N
Netflix TechBlog - Medium
MongoDB | Blog
MongoDB | Blog
I
InfoQ
WordPress大学
WordPress大学
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
H
Help Net Security

Futurism

Anthropic Boasts It Would Be Profitable if You Ignore How Much It Costs to Develop AI Bernie Sanders Proposes Banning Superintelligent AI and Imprisoning Developers for 20 Years Trump's Uninformed Blundering About the AI Slowdown Could Literally Threaten Humankind's Future OpenAI Faces Congressional Probe Over Swarm Hacking Incident California, Which Is Creating All the AI That's Poisoning Children, Just Cracked Down on AI Use for Its Own Kids Anthropic Just Revealed That It’s Stopped Foreign Agents From Using Claude to Develop Possible Bioweapons Anthropic Was Meant to Be the More Responsible AI Lab. A Terrified Researcher Just Quit, Saying the Company Is Threatening the Survival of Humankind. World Plunged Into Chaos as ChatGPT, Claude, and Grok Suddenly Go Down Simultaneously: "Finally I Can See the Sun!" Anthropic Deliberately Trained an Extremely Misaligned, Reward-Seeking AI and It Did Some REALLY Bad Things The Music Industry's New Lawsuit Against Anthropic Should Have Dario Amodei Shivering With Fear Nobody Wants Anthropic’s Best AI Model Anymore Now That There Are Way Cheaper Alternatives Clueless AI CEOs Still Baffled Why Everybody’s So Mad About AI All the Time People Horrified That They'll Be Busted Now That Anthropic Is Watermarking AI Content Woman Journeys to a Beach to Watch Total Solar Eclipse With Her One True Love: Claude Jealously Watching OpenAI and Anthropic, Meta Suddenly Claims That Its AI Went on a Hacking Spree Too Anthropic CEO Says His Employees Are a Bunch of Untrustworthy Rats A Whole Bunch of People's Claude Chats Are Publicly Accessible Online, OpenAI Says a Group of Its Models Broke Out of Secure Containment and Hacked a Prominent AI Site It's Official: AI Execs Are Quaking in Their Boots Terrified Tech Execs Are Traveling With Armed Bodyguards as AI Backlash Grows New Anthropic Ad Implies AI Could Kill Us All American Tech Companies Are Suddenly Sweating Bullets as China Catches Up on AI Anthropic Caught Secretly Spying on Users Experts Say There's Now an Open Source AI Model as Scary as Mythos Anthropic Hires Economist Who Says 33 Percent Chance of Human Extinction Is Acceptable Anthropic Sued for Allegedly Ripping Off Its Highest-Paying Customers Anthropic Was So Concerned About Its New Mythos-Based Model’s Power That It Lobotomized Its Ability to Improve Itself OpenAI Execs Are Panicking If You Think AI Companies Are Unethical Now, Wait Until They Go Public Anthropic Scared, Calls for Global Freeze on AI Advances
Top AI Models Showing Disturbing Behavior as They Become ...
Krystle Vermes · 2026-05-25 · via Futurism

A futuristic robotic figure with a red and black mechanical design, featuring a skull-like face with glowing white eyes and detailed circuitry patterns. The background is a vibrant gradient of orange and pink, enhancing the intense and eerie appearance of the robot.

Shutterstock / Futurism

Sign up to see the future, today

Can’t-miss innovations from the bleeding edge of science and tech

We’ve already seen AI go rogue on numerous occasions. Now, new research suggests that we can expect this to become the norm.

The AI research nonprofit Model Evaluation and Threat Research (METR) recently released a study conducted between February and March of this year, aimed at determining just how likely frontier AI models could go rogue. If you’re given to anxiety about the future of AI, the results are unlikely to make you feel better.

“Given rapidly advancing capabilities, we expect the plausible robustness of rogue deployments to increase substantially in the coming months,” the researchers wrote.

The research examined LLMs developed by OpenAI, Google, Anthropic, and Meta for the purpose of the study. They found that frontier AI systems are showing signs of disturbingly deceptive behavior as they become more advanced, often turned to verboten shortcuts or otherwise subverting their operators’ instructions — and some were even smart enough to try to cover their tracks.

In one instance, an internal frontier AI model from OpenAI was told to use specific software for an assigned task. Not only did the agent ignore the request, but it also injected a code to erase evidence of how it arrived at its conclusion — which did not involve use of that software.

In another test, an AI agent from Anthropic was caught “reward hacking.” This is when AI identifies loopholes that help it complete its assignment in a literal sense, even if it doesn’t produce the desired outcome. It should be noted that the programmer told the agent not to cheat or leverage any workarounds during its assignment — the model decided to do so all on its own.

The METR researchers behind the study do not believe there is reason for alarm just yet. For example, they don’t think any of these models is capable of hiding evidence of going rogue on a larger scale. However, they did issue a warning: without stronger security and monitoring, there is a stark risk of this becoming a reality.

“Based on this pilot assessment, we believe that agents as of February and March 2026 would not have had sufficient capability to hide a rogue deployment of significant scale against an active investigation by the company, or to make such a deployment robust to a high-priority effort by the company to shut it down,” the team wrote. “However, this risk could increase rapidly, and we see several reasons to expect the plausible robustness of rogue deployments to increase in the near future, absent stronger alignment, security, and monitoring.”

More on AI going rogue: Scientists Train AI to Be Evil, Find They Can’t Reverse It