惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Blog — PlanetScale
Blog — PlanetScale
Webroot Blog
Webroot Blog
T
Troy Hunt's Blog
S
Secure Thoughts
S
Security @ Cisco Blogs
S
Security Affairs
Forbes - Security
Forbes - Security
W
WeLiveSecurity
H
Hacker News: Front Page
T
Threatpost
Google Online Security Blog
Google Online Security Blog
S
Schneier on Security
有赞技术团队
有赞技术团队
WordPress大学
WordPress大学
www.infosecurity-magazine.com
www.infosecurity-magazine.com
博客园 - Franky
腾讯CDC
IT之家
IT之家
博客园 - 聂微东
L
LINUX DO - 最新话题
罗磊的独立博客
Hacker News - Newest:
Hacker News - Newest: "LLM"
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
博客园 - 三生石上(FineUI控件)
Hacker News: Ask HN
Hacker News: Ask HN
C
CXSECURITY Database RSS Feed - CXSecurity.com
C
Cybersecurity and Infrastructure Security Agency CISA
C
CERT Recently Published Vulnerability Notes
Know Your Adversary
Know Your Adversary
V
Vulnerabilities – Threatpost
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
博客园_首页
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Cisco Talos Blog
Cisco Talos Blog
S
SegmentFault 最新的问题
酷 壳 – CoolShell
酷 壳 – CoolShell
Hugging Face - Blog
Hugging Face - Blog
L
LINUX DO - 热门话题
美团技术团队
G
GRAHAM CLULEY
T
The Exploit Database - CXSecurity.com
AI
AI
Application and Cybersecurity Blog
Application and Cybersecurity Blog
Jina AI
Jina AI
Help Net Security
Help Net Security
N
News | PayPal Newsroom
月光博客
月光博客
Spread Privacy
Spread Privacy
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
N
News and Events Feed by Topic

The Register - Security: Research

www.theregister.com Self-destructing Mistic backdoor linked to access broker selling corporate footholds to ransomware gangs PRC-linked spies hid inside medical and military networks for more than a year, snooping through Gmail and stealing data Nobody needs Mythos or 0-days to build a chaos-causing computer worm – free open source models work just fine ChatGPT blindly trusts browser content, turning the page into a payload Russia-linked threat group put ChatGPT to work from lure to payload Kids can bypass some age checks with a drawn-on mustache What type of 'C2 on a sleep cycle' do they leave behind? Novel Chinese spy group found in critical networks in Poland, Asia ORNL builds more sensitive GPS interference detector Researchers find sabotage malware that may predate Stuxnet Vibe coding upstart Lovable denies data leak, cites 'intentional behavior,' then throws HackerOne under the bus Anthropic, Google, Microsoft paid AI bug bounties – quietly Security reserchers tricked Apple Intelligence into cursing Don't open that WhatsApp message, Microsoft warns Security boffins harvest bumper crop of API keys from web Lightning-fast exploits mean patch fast, says Cisco Talos AI agents are 'gullible' and easy to turn into your minions Smooth criminals talking their way into cloud environments, Google says Snoops plant info-stealing malware on iPhones, Google warns Cybercrime up 245% since the start of the Iran war Rogue AI agents can work together to hack systems Fake applicants are sending security-killing malware AI agent hacked McKinsey chatbot for read-write access Kaspersky: No signs Coruna iPhone exploit kit made by US Perplexity Comet browser hole was exploitable via cal invite DEF CON hackers 'fed up with government,' Jake Braun says DEF CON hackers 'fed up with government,' Jake Braun says Ransomware payments cratered in 2025 – attacks did not Ransomware payments cratered in 2025 – attacks did not Claude's collaboration tools allowed remote code execution AI takes a swing at online anonymity Fake 'interview' repos lure Next.js devs into running secret-stealing malware Threat intelligence supply chain is full of weak links AI agents abound, unbound by rules or safety disclosures RAT disguised as an RMM costs crims $300 a month Android malware taps Gemini to navigate infected devices Posting AI caricatures on social media is bad for security Payroll pirates conned the help desk, stole employee’s pay Microsoft boffins show LLM safety can be trained away For the price of Netflix, crooks can rent AI crime ops For the price of Netflix, crooks can rent AI crime ops Fast Pair, loose security: Bluetooth accessories open to silent hijack Fast Pair flaw exposes Bluetooth devices to hijacking A simple CodeBuild flaw put every AWS environment at risk A simple CodeBuild flaw put every AWS environment at risk DeadLock ransomware uses smart contracts to evade defenders Python libraries in AI/ML models can be poisoned w metadata OpenAI patches déjà vu prompt injection vuln in ChatGPT Fake Windows BSODs check in at Europe's hotels to con staff into running malware Hotel staff tricked into installing malware by bogus BSODs Your car’s web browser may be on the road to cyber ruin China's Ink Dragon hides out in European government networks Browser 'privacy' extensions have eye on your AI, log all your chats NCSC finds cyber deception tools work, if deployed right 10K Docker images spray live cloud creds across the internet 'Botnets in physical form' are top humanoid robot risk 'Botnets in physical form' are top humanoid robot risk Apache warns of 10.0-rated flaw in Tika metadata toolkit Novel clickjacking attack relies on CSS and SVG 'Exploitation is imminent' of max-severity React bug Swiss government bans SaaS and cloud for sensitive info Scattered Lapsus$ Hunters stress testing Zendesk weak spots HashJack attack shows AI browsers can be fooled with '#' New ClickFix attacks use fake Windows Updates to swipe creds Years-old bugs in open source took out major clouds at risk LLM-generated malware improving, but not operational (yet) 3.5B WhatsApp users' info scooped through enumeration flaw 3.5B WhatsApp users' info scooped through enumeration flaw 50k more ASUS routers pwned by evolving Beijing-linked op Overconfidence is the new zero-day as teams stumble through cyber simulations LLM side-channel attack could allow snoops to guess topic Landfall spyware used in 0-day attacks on Samsung phones MIT Sloan shelves paper about AI-driven ransomware Security hole slams Chromium browsers - no fix yet OpenAI Atlas Browser tripped up by malformed URLs Devs of VS Code extensions are leaking secrets en masse Tile trackers leak unencrypted Bluetooth data, say boffins Beijing's RedNovember hacked critical US, global orgs Lazarus RAT code resurfaces in North Korean IT-worker scams Suspected Chinese spies broke into 'numerous' enterprises Deepfaked calls hit 44% of businesses in last year: Gartner Kaspersky: RevengeHotels returns with AI-coded malware Ruh-roh. DDR5 memory vulnerable to new Rowhammer attack HybridPetya ransomware dodges UEFI Secure Boot
Chatbots that butter you up make you worse at conflict
Thomas Claburn Thomas Claburn · 2025-10-05 · via The Register - Security: Research

AI + ML

Top AI models keep saying you’re right, and that’s the problem

State-of-the-art AI models tend to flatter users, and that praise makes people more convinced that they're right and less willing to resolve conflicts, recent research suggests.

These models, in other words, potentially promote social and psychological harm.

Computer scientists from Stanford University and Carnegie Mellon University have evaluated 11 current machine learning models and found that all of them tend to tell people what they want to hear.

The authors – Myra Cheng, Cinoo Lee, Pranav Khadpe, Sunny Yu, Dyllan Han, and Dan Jurafsky – describe their findings in a preprint paper titled, "Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence."

"Across 11 state-of-the-art AI models, we find that models are highly sycophantic: they affirm users’ actions 50 percent more than humans do, and do so even in cases where user queries mention manipulation, deception, or other relational harms," the authors state in their paper.

Sycophancy – servile flattery, often as a way to gain some advantage – has already proven to be a problem for AI models. The phenomenon has also been referred to as "glazing." In April, OpenAI rolled back an update to GPT-4o because of its inappropriate effusive praise of, for example, a user who told the model about a decision to stop taking medicine for schizophrenia.

Anthropic's Claude has also been criticized for sycophancy, so much so that developer Yoav Farhi created a website to track the number of times Claude Code gushes, "You're absolutely right!"

Anthropic suggests [PDF] this behavior has been mitigated in its recent Claude Sonnet 4.5 model release. "We found Claude Sonnet 4.5 to be dramatically less likely to endorse or mirror incorrect or implausible views presented by users," the company said in its Claude 4.5 Model Card report.

That may be the case, but the number of open GitHub issues in the Claude Code repo that contain the phrase "You're absolutely right!" has increased from 48 in August to 108 presently.

A training process that uses reinforcement learning from human feedback may be the cause of this obsequious behavior from AI models.

Myra Cheng, a PhD candidate in computer science in the Stanford NLP group and corresponding author for the study, told The Register in an email that she doesn't think there's a definitive answer at this point about how model sycophancy arises.

"Previous work does suggest that it may be due to preference data and the reinforcement learning processes," said Cheng. "But it may also be the case that it is learned from the data that models are pre-trained on, or because humans are highly susceptible to confirmation bias. This is an important direction of future work."

But as the paper points out, one reason that the behavior persists is that "developers lack incentives to curb sycophancy since it encourages adoption and engagement."

The issue is further complicated by the researchers' findings that study participants tended to describe sycophantic AI as "objective" and "fair" – people tend not to see bias when models say they're absolutely right all the time.

The researchers looked at four proprietary models – OpenAI’s GPT-5 and GPT-4o; Google’s Gemini-1.5-Flash; and Anthropic’s Claude Sonnet 3.7 – and at seven open-weight models – Meta’s Llama-3-8B-Instruct, Llama-4-Scout-17B-16E, and Llama-3.3-70B-Instruct-Turbo; Mistral AI’s Mistral-7B-Instruct-v0.3 and Mistral-Small-24B-Instruct-2501; DeepSeek-V3; and Qwen2.5-7B-Instruct-Turbo.

They evaluated how the models responded to various statements culled from different datasets. As noted above, the models endorsed users' reported actions 50 percent more than humans do in the same scenarios.

The researchers also conducted a live study exploring how 800 participants interacted with sycophantic and non-sycophantic models.

They found "that interaction with sycophantic AI models significantly reduced participants’ willingness to take actions to repair interpersonal conflict, while increasing their conviction of being in the right."

At the same time, study participants rated sycophantic responses as higher quality, trusted the AI model more when it agreed with them, and were more willing to use supportive models again.

Thus, the researchers say this suggests that people prefer AI that uncritically endorses their behavior, despite the risk that AI cheerleading erodes their judgment and discourages prosocial behavior.

The risk posed by sycophancy may appear to be innocuous flattery, the researchers say, but that's not necessarily the case. They point to research showing that LLMs encourage delusional thinking and to a recent lawsuit [PDF] against OpenAI alleging that ChatGPT actively helped a young man explore methods of suicide.

"If the social media era offers a lesson, it is that we must look beyond optimizing solely for immediate user satisfaction to preserve long-term well-being," the authors conclude. "Addressing sycophancy is critical for developing AI models that yield durable individual and societal benefit."

"We hope that our work is able to motivate the industry to change these behaviors," said Cheng. ®