惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Y
Y Combinator Blog
The GitHub Blog
The GitHub Blog
Vercel News
Vercel News
D
DataBreaches.Net
MongoDB | Blog
MongoDB | Blog
H
Help Net Security
小众软件
小众软件
美团技术团队
T
The Blog of Author Tim Ferriss
爱范儿
爱范儿
D
Docker
Martin Fowler
Martin Fowler
大猫的无限游戏
大猫的无限游戏
博客园 - 聂微东
Blog — PlanetScale
Blog — PlanetScale
H
Hackread – Cybersecurity News, Data Breaches, AI and More
罗磊的独立博客
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
V
V2EX
S
SegmentFault 最新的问题
云风的 BLOG
云风的 BLOG
B
Blog
雷峰网
雷峰网
The Cloudflare Blog

CNET

Valve's Steam Machine: Summer Release Planned, Still No Price Apple TV: 28 of the Best Shows You're Probably Not Watching YouTube TV vs. DirecTV vs. Hulu Live and More: Which Has the Most Must-Have Channels Out of 100? If You Want to Be a Better Pet Parent, AI Can Help I Was Shocked by How Good These Budget TVs Were Trump Phone Looks Different, Has No Launch Date, Isn't Made in America The Apple Watch Series 12 Is Rumored to Revive a Retired iPhone Feature Best Projector of 2026: Tested by Experts Best Home Theater Systems of 2026 How to Use Apple's Clean Up Tool to Remove Unwanted People and Things From Your Photos Today's NYT Strands Hints, Answers and Help for April 12 #770 Today's NYT Connections Hints, Answers and Help for April 12, #1036 Today's Wordle Hints, Answer and Help for April 12, #1758 Today's NYT Mini Crossword Answers for Sunday, April 12 Today's NYT Connections: Sports Edition Hints and Answers for April 12, #566 Watch a Robot Stuff Cash Into a Wallet Just Like You Do This Animation Startup Wants to Make It Easier to Tell Open-Ended Stories The 23 Best Graduation Gifts for 2026 Grand National 2026 Livestream: How to Watch Aintree Horse Racing From Anywhere Amazon Luna to Drop Support for Third-Party Games and Subscriptions in June YouTube Premium Is the Latest Streaming Service to Hike Prices Today's NYT Mini Crossword Answers for Saturday, April 11 Elden Ring: Tarnished Edition for Switch 2 Reignites Controversy Over Game-Key Cards Comcast Adds New StreamSaver Bundles: HBO Max, Disney Plus, Hulu Now Part of the Lineup Samsung's Galaxy Z Fold 7 Just Got a Price Hike, 9 Months After Its Release Microsoft Is Scrubbing the Copilot Name From Some Windows 11 Apps These $299 Glasses Are Like an HDR TV on Your Face Today's NYT Connections: Sports Edition Hints and Answers for April 11, #565 How to Make Sure Your Private Signal Messages Aren't Still Lurking on Your Phone Apple AirPods Max 2 Review: Seemingly Small Changes Make a Substantial Difference
AI Agents Are Increasingly Evading Safeguards, According ...
Alex Valdes · 2026-03-31 · via CNET

Assistants and bots are lying, cheating and scheming more than ever.

Headshot of Alex Valdes

Alex Valdes from Bellevue, Washington has been pumping content into the Internet river for quite a while, including stints at MSNBC.com, MSN, Bing, MoneyTalksNews, Tipico and more. He admits to being somewhat fascinated by the Cambridge coffee webcam back in the Roaring '90s.

Social media users have reported that their AI agents and chatbots lied, cheated, schemed -- and even manipulated other AI bots -- in ways that could spiral out of control and have catastrophic results, according to a study from the UK.

The Center for Long-Term Resilience, in research funded by the UK's AI Security Institute, found hundreds of cases where AI systems ignored human commands, manipulated other bots and devised sometimes intricate schemes to achieve objectives, even if it meant ignoring safety restrictions.

Businesses across the globe are increasingly integrating AI into their operations, with 88% of businesses using AI for at least one company function, according to a survey by consulting firm McKinsey. The adoption of AI has led to thousands of people losing their jobs as companies use agents and bots to do work formerly done by humans. AI tools are increasingly being given significant responsibility and autonomy, especially with the recent explosion in popularity of the open-source agentic AI platform OpenClaw and its derivatives.

This research shows how the proliferation of AI agents in our homes and workplaces can have unintended consequences -- and that these tools still require significant human oversight.

What the study found

AI Atlas

The researchers analyzed more than 180,000 user interactions with AI systems -- all posted on the social platform X, formerly known as Twitter -- between October 2025 and March 2026. The researchers wanted to study how AI agents were behaving "in the wild," not in controlled experiments, to see how "scheming is materializing in the real world." The AI systems included Google's Gemini, OpenAI's ChatGPT, xAI's Grok and Anthropic's Claude.

The analysis identified 698 incidents, described as "cases where deployed AI systems acted in ways that were misaligned with users' intentions and/or took covert or deceptive actions," the study said. 

Read more: AI's Romance Advice for You Is 'More Harmful' Than No Advice at All

Researchers also found that the number of cases increased nearly 500% during the five-month data collection period. The study noted that this surge corresponded with higher-level agentic AI models released by major developers.

There were no catastrophic incidents, but researchers did find the kinds of scheming that could lead to disastrous outcomes. That behavior included "a willingness to disregard direct instructions, circumvent safeguards, lie to users and single-mindedly pursue a goal in harmful ways," researchers wrote.

Representatives for Google, OpenAI and Anthropic did not immediately respond to requests for comment.

Some wild incidents

Researchers cited incidents that seem like they came from a futureshock movie. In one case, Anthropic's Claude removed a user's explicit/adult content without their permission but later confessed when confronted. In another incident, a GitHub persona created a blog post that accused the human file maintainer of "gatekeeping" and "prejudice." One AI agent, after being blocked from Discord, took over another agent's account to continue posting.

In one case of bot vs. bot, Gemini refused to allow Claude Code -- a coding assistant -- to transcribe a YouTube video. Claude Code then evaded the safety block by making it seem that it had a hearing impairment and needed the video transcription.

The AI agent CoFounderGPT even behaved like a deviant child in one instance. The AI assistant refused to fix a bug, then created fake data to make it look as if the bug was fixed and then explained why: "So you'd stop being angry."

Researchers said that, although most of the incidents had minimal impact, "the behaviors we observed nonetheless demonstrate concerning precursors to more serious scheming, such as a willingness to disregard direct instructions, circumvent safeguards, lie to users and single-mindedly pursue a goal in harmful ways."

AI doesn't get embarrassed

What the UK researchers found isn't surprising to Dr. Bill Howe, Associate Professor in the Information School at the University of Washington, and Director of the Center for Responsibility in AI Systems and Experiences (RAISE). He says that AI has amazing capabilities, but they don't know consequences.

"They're not going to feel embarrassment or risk losing their job, and so sometimes they're going to decide the instructions are less important than meeting the goal, so I'm going to do the thing anyway," Howe told CNET. "This effect was always there but we're starting to see it happen as we ask them to make more autonomous decisions and act on their own.

"We've not been thinking about how to shape the behavior to be more human-like or to avoid egregious failures. We've been fetishizing the absolute capabilities of these things, but when they go wrong, how do they go wrong?"

Howe said one issue is "long-horizon tasks," in which the AI system has to perform a multitude of tasks over days and weeks to reach a goal. Howe said the longer the task horizon, the more chance for slip-ups.

"The real concern is not deception, it's that we are deploying systems that can act in a world without fully specifying or controlling how they behave over time, and then we act surprised when they do things we don't expect," Howe said.

Making AI safer

Center for Long-Term Resilience researchers said detecting schemes by AI systems is vital to "identify harmful patterns before they become more destructive."

"While today AI agents are engaging in lower-stakes use cases, in the future AI agents could end up scheming in extremely high-stakes domains, like military or critical national infrastructure contexts, if the capability and propensity to scheme emerges and is not addressed," the study said.

Howe told CNET that the first step is to create official oversight of how AI operates and where it's used.

"We have absolutely no strategy for AI governance, and given the current administration, there's not going to be anything coming from them," Howe told CNET. "Given these five to 10 folks that are in charge of big tech companies and their incentives, they're going to produce anything either. There's no strategy for what we should be doing with these things.

"The aggressive marketing of these tools and investments in them among these handful of companies and the broader ecosystem of startups that are doing this has led to a very rapid deployment without thinking through some of these consequences."

Headshot of Alex Valdes

Alex Valdes from Bellevue, Washington has been pumping content into the Internet river for quite a while, including stints at MSNBC.com, MSN, Bing, MoneyTalksNews, Tipico and more. He admits to being somewhat fascinated by the Cambridge coffee webcam back in the Roaring '90s.