惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

J
Java Code Geeks
F
Fortinet All Blogs
Martin Fowler
Martin Fowler
M
MIT News - Artificial intelligence
G
Google Developers Blog
P
Proofpoint News Feed
Recent Announcements
Recent Announcements
MyScale Blog
MyScale Blog
D
DataBreaches.Net
Stack Overflow Blog
Stack Overflow Blog
月光博客
月光博客
爱范儿
爱范儿
罗磊的独立博客
腾讯CDC
Hugging Face - Blog
Hugging Face - Blog
博客园 - 叶小钗
Vercel News
Vercel News
酷 壳 – CoolShell
酷 壳 – CoolShell
B
Blog
C
Check Point Blog
美团技术团队
宝玉的分享
宝玉的分享
Microsoft Security Blog
Microsoft Security Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻

Futurism

Meta Installing Software on Employee Computers to Track Everything They Do, Feed the Data to AI Concern Grows That AI Is Damaging Users’ Cognitive Abilities JPMorganChase Data Center Gets $77 Million Handout to Create Grand Total of One Job Nvidia CEO Loses His Cool at Tough Question CEO of $1.5 Billion AI Startup Accused of Massive Fraud by Justice Department Palantir Issues Ominous Corporate Manifesto Madison Square Garden Reportedly Used Facial Recognition to Stalk Trans Woman For Two Years The Florida Mass Shooter’s Conversations With ChatGPT Are Worse Than You Could Possibly Imagine China Is Starting to Pull Ahead of US in AI Race AI Company Known for Teen Suicides Launches New Feature to Turn Books Into Roleplaying Experiences Study Finds AI Use Eats Away at Users’ Confidence in Their Own Brains Democrats Warned Not to Upset Multi-Million Dollar AI Lobbyists, Even Though It’d Be a Slam Dunk With Voters City Council Wrecked in Voter Bloodbath After Allowing New Data Center Mother Reportedly Doesn’t Know Her Son Died Because She’s Been Talking to an AI Version of Him Things You Told ChatGPT or Claude My Have Already Doomed You in Court Millions of Americans Are Talking to AI Instead of Going to the Doctor, and It’s Giving Them Horrendously Flawed Medical Advice There Are Signs of a Massive AI Backlash A Prominent PR Firm Is Running a Fake News Site That’s Plagiarizing Original Journalism at Incredible Scale Fury Erupts as Val Kilmer’s Estate Announces Starring Role in AI Film Made From Beyond the Grave Allbirds Stock Now Crashing as Reality Sets in About Its Delusional AI Pivot NAACP Sues Elon Over His Noxious AI Data Center Top Security Experts Alarmed by Power of Anthropic’s New Hacker AI Teens Alarmed at What AI Is Doing to Their Minds What It Really Means That a Failing Shoe Brand “Pivoted to AI” and Its Stock Soared 700 Percent Starbucks’ Baffling ChatGPT Collab Treats Customers Like Empty, Soulless Venti Cups ChatGPT’s “Honest Reaction” to a “Song” Composed Entirely of Gas-Passing Noises Will Make You Question Whether It’s Honestly Evaluating Your Other Brilliant Ideas AI Is Turning Workplaces Into Hopeless Gridlock Companies Just Learned a Brutal Lesson About Training AI to Do Human Jobs Berklee College of Music Students Furious That It’s Offering an AI “Songwriting” Class Usually, Young People Embrace New Technology. Gen Z’s Attitude Toward AI Should Worry the Entire Tech Industry
Top AI Models Showing Disturbing Behavior as They Become ...
Krystle Vermes · 2026-05-25 · via Futurism

A futuristic robotic figure with a red and black mechanical design, featuring a skull-like face with glowing white eyes and detailed circuitry patterns. The background is a vibrant gradient of orange and pink, enhancing the intense and eerie appearance of the robot.

Shutterstock / Futurism

Sign up to see the future, today

Can’t-miss innovations from the bleeding edge of science and tech

We’ve already seen AI go rogue on numerous occasions. Now, new research suggests that we can expect this to become the norm.

The AI research nonprofit Model Evaluation and Threat Research (METR) recently released a study conducted between February and March of this year, aimed at determining just how likely frontier AI models could go rogue. If you’re given to anxiety about the future of AI, the results are unlikely to make you feel better.

“Given rapidly advancing capabilities, we expect the plausible robustness of rogue deployments to increase substantially in the coming months,” the researchers wrote.

The research examined LLMs developed by OpenAI, Google, Anthropic, and Meta for the purpose of the study. They found that frontier AI systems are showing signs of disturbingly deceptive behavior as they become more advanced, often turned to verboten shortcuts or otherwise subverting their operators’ instructions — and some were even smart enough to try to cover their tracks.

In one instance, an internal frontier AI model from OpenAI was told to use specific software for an assigned task. Not only did the agent ignore the request, but it also injected a code to erase evidence of how it arrived at its conclusion — which did not involve use of that software.

In another test, an AI agent from Anthropic was caught “reward hacking.” This is when AI identifies loopholes that help it complete its assignment in a literal sense, even if it doesn’t produce the desired outcome. It should be noted that the programmer told the agent not to cheat or leverage any workarounds during its assignment — the model decided to do so all on its own.

The METR researchers behind the study do not believe there is reason for alarm just yet. For example, they don’t think any of these models is capable of hiding evidence of going rogue on a larger scale. However, they did issue a warning: without stronger security and monitoring, there is a stark risk of this becoming a reality.

“Based on this pilot assessment, we believe that agents as of February and March 2026 would not have had sufficient capability to hide a rogue deployment of significant scale against an active investigation by the company, or to make such a deployment robust to a high-priority effort by the company to shut it down,” the team wrote. “However, this risk could increase rapidly, and we see several reasons to expect the plausible robustness of rogue deployments to increase in the near future, absent stronger alignment, security, and monitoring.”

More on AI going rogue: Scientists Train AI to Be Evil, Find They Can’t Reverse It