惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

人人都是产品经理
人人都是产品经理
量子位
月光博客
月光博客
罗磊的独立博客
宝玉的分享
宝玉的分享
博客园_首页
酷 壳 – CoolShell
酷 壳 – CoolShell
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
WordPress大学
WordPress大学
博客园 - 叶小钗
博客园 - 聂微东
阮一峰的网络日志
阮一峰的网络日志
V
V2EX
雷峰网
雷峰网
博客园 - 三生石上(FineUI控件)
Jina AI
Jina AI
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
博客园 - Franky
美团技术团队
爱范儿
爱范儿
V
Visual Studio Blog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
H
Hackread – Cybersecurity News, Data Breaches, AI and More
Y
Y Combinator Blog

The Register - Security: Research

Novel Blue Moon kit targeting Chrome and Windows reflects new reality of AI-driven exploits Extortion crews have their eyes on high-value AI data, Google warns Researcher shows how Claude Code can be tricked simply by asking it to summarize a website Copilot tricked into telling reseachers how to hack itself Akira ransomware scum blocked victim How the famed USENIX Security conf is managing a flood of papers in the AI era www.theregister.com Self-destructing Mistic backdoor linked to access broker selling corporate footholds to ransomware gangs PRC-linked spies hid inside medical and military networks for more than a year, snooping through Gmail and stealing data Nobody needs Mythos or 0-days to build a chaos-causing computer worm – free open source models work just fine ChatGPT blindly trusts browser content, turning the page into a payload Russia-linked threat group put ChatGPT to work from lure to payload Kids can bypass some age checks with a drawn-on mustache What type of 'C2 on a sleep cycle' do they leave behind? Novel Chinese spy group found in critical networks in Poland, Asia ORNL builds more sensitive GPS interference detector Researchers find sabotage malware that may predate Stuxnet Vibe coding upstart Lovable denies data leak, cites 'intentional behavior,' then throws HackerOne under the bus Anthropic, Google, Microsoft paid AI bug bounties – quietly Security reserchers tricked Apple Intelligence into cursing Don't open that WhatsApp message, Microsoft warns Security boffins harvest bumper crop of API keys from web Lightning-fast exploits mean patch fast, says Cisco Talos AI agents are 'gullible' and easy to turn into your minions Smooth criminals talking their way into cloud environments, Google says Snoops plant info-stealing malware on iPhones, Google warns Cybercrime up 245% since the start of the Iran war Rogue AI agents can work together to hack systems Fake applicants are sending security-killing malware AI agent hacked McKinsey chatbot for read-write access Kaspersky: No signs Coruna iPhone exploit kit made by US
Chatbots that butter you up make you worse at conflict
Thomas Claburn Thomas Claburn · 2025-10-05 · via The Register - Security: Research

AI + ML

Top AI models keep saying you’re right, and that’s the problem

State-of-the-art AI models tend to flatter users, and that praise makes people more convinced that they're right and less willing to resolve conflicts, recent research suggests.

These models, in other words, potentially promote social and psychological harm.

Computer scientists from Stanford University and Carnegie Mellon University have evaluated 11 current machine learning models and found that all of them tend to tell people what they want to hear.

The authors – Myra Cheng, Cinoo Lee, Pranav Khadpe, Sunny Yu, Dyllan Han, and Dan Jurafsky – describe their findings in a preprint paper titled, "Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence."

"Across 11 state-of-the-art AI models, we find that models are highly sycophantic: they affirm users’ actions 50 percent more than humans do, and do so even in cases where user queries mention manipulation, deception, or other relational harms," the authors state in their paper.

Sycophancy – servile flattery, often as a way to gain some advantage – has already proven to be a problem for AI models. The phenomenon has also been referred to as "glazing." In April, OpenAI rolled back an update to GPT-4o because of its inappropriate effusive praise of, for example, a user who told the model about a decision to stop taking medicine for schizophrenia.

Anthropic's Claude has also been criticized for sycophancy, so much so that developer Yoav Farhi created a website to track the number of times Claude Code gushes, "You're absolutely right!"

Anthropic suggests [PDF] this behavior has been mitigated in its recent Claude Sonnet 4.5 model release. "We found Claude Sonnet 4.5 to be dramatically less likely to endorse or mirror incorrect or implausible views presented by users," the company said in its Claude 4.5 Model Card report.

That may be the case, but the number of open GitHub issues in the Claude Code repo that contain the phrase "You're absolutely right!" has increased from 48 in August to 108 presently.

A training process that uses reinforcement learning from human feedback may be the cause of this obsequious behavior from AI models.

Myra Cheng, a PhD candidate in computer science in the Stanford NLP group and corresponding author for the study, told The Register in an email that she doesn't think there's a definitive answer at this point about how model sycophancy arises.

"Previous work does suggest that it may be due to preference data and the reinforcement learning processes," said Cheng. "But it may also be the case that it is learned from the data that models are pre-trained on, or because humans are highly susceptible to confirmation bias. This is an important direction of future work."

But as the paper points out, one reason that the behavior persists is that "developers lack incentives to curb sycophancy since it encourages adoption and engagement."

The issue is further complicated by the researchers' findings that study participants tended to describe sycophantic AI as "objective" and "fair" – people tend not to see bias when models say they're absolutely right all the time.

The researchers looked at four proprietary models – OpenAI’s GPT-5 and GPT-4o; Google’s Gemini-1.5-Flash; and Anthropic’s Claude Sonnet 3.7 – and at seven open-weight models – Meta’s Llama-3-8B-Instruct, Llama-4-Scout-17B-16E, and Llama-3.3-70B-Instruct-Turbo; Mistral AI’s Mistral-7B-Instruct-v0.3 and Mistral-Small-24B-Instruct-2501; DeepSeek-V3; and Qwen2.5-7B-Instruct-Turbo.

They evaluated how the models responded to various statements culled from different datasets. As noted above, the models endorsed users' reported actions 50 percent more than humans do in the same scenarios.

The researchers also conducted a live study exploring how 800 participants interacted with sycophantic and non-sycophantic models.

They found "that interaction with sycophantic AI models significantly reduced participants’ willingness to take actions to repair interpersonal conflict, while increasing their conviction of being in the right."

At the same time, study participants rated sycophantic responses as higher quality, trusted the AI model more when it agreed with them, and were more willing to use supportive models again.

Thus, the researchers say this suggests that people prefer AI that uncritically endorses their behavior, despite the risk that AI cheerleading erodes their judgment and discourages prosocial behavior.

The risk posed by sycophancy may appear to be innocuous flattery, the researchers say, but that's not necessarily the case. They point to research showing that LLMs encourage delusional thinking and to a recent lawsuit [PDF] against OpenAI alleging that ChatGPT actively helped a young man explore methods of suicide.

"If the social media era offers a lesson, it is that we must look beyond optimizing solely for immediate user satisfaction to preserve long-term well-being," the authors conclude. "Addressing sycophancy is critical for developing AI models that yield durable individual and societal benefit."

"We hope that our work is able to motivate the industry to change these behaviors," said Cheng. ®