惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

P
Proofpoint News Feed
Martin Fowler
Martin Fowler
The GitHub Blog
The GitHub Blog
B
Blog RSS Feed
U
Unit 42
阮一峰的网络日志
阮一峰的网络日志
量子位
GbyAI
GbyAI
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
云风的 BLOG
云风的 BLOG
小众软件
小众软件
博客园 - 三生石上(FineUI控件)
L
LangChain Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
博客园_首页
IT之家
IT之家
V
Visual Studio Blog
Y
Y Combinator Blog
Blog — PlanetScale
Blog — PlanetScale
宝玉的分享
宝玉的分享
Apple Machine Learning Research
Apple Machine Learning Research
I
InfoQ
D
Docker
V
V2EX

Futurism

Man Creates Tiny Submarine for His Parakeet to Experience Life Underwater The Effects of AI-Generated Code Tearing Through Corporations Is Actually Kind of Funny Trump Hires Orbital Towing Company to Build Space Interceptors Psychologists Found Something Horrible About the Kind of Men Seeking Trad Wives To Get Swole, Teens Are Pumping Themselves Full of Drugs Meant for Fattening Cows for the Slaughterhouse Foolish Pollsters Are Now Just Asking AI What Voters Would Say in Response to Questions and Publishing It at Face Value OpenAI Says It’s Already Made $100 Million by Stuffing ChatGPT With Ads Man Punished for Breaking Into Moo Deng’s Zoo Enclosure AI Is Causing Healthcare Costs to Surge There’s a Mass Rebellion Against AI in the Workplace People Who Lose Their Job to AI Are in for a World of Pain, Goldman Sachs Report Finds OpenAI Says Not to Worry About UBI, Because It Has Another Idea Police Officer Helplessly Waves Arms at Waymo That Careened Wrong Way Through Whataburger Drive-Thru Someone Just Threw a Molotov Cocktail At Sam Altman’s House New York Times Makes Substantial Changes to Article That Glazed a Sleazy AI Startup: “Our Piece Should Have Included That Information” Space Scientists Wince as Astronauts’ Lives Depend on Artemis 2’s Controversial Heat Shield During Plunge Back to Earth The Moon Astronauts Have Been Working Out With a NASA Rowing Machine in Space First AI Model From Zuckerberg’s Wildly Expensive Superintelligence Lab Flops Compared to Virtually All Rivals Economists Starting to Admit They May Have Been Wrong About AI Never Replacing Human Jobs AI-Powered Drug Marketer Medvi Responds After Allegations About Fake Doctors and Patients As Astronauts Visit the Moon, NASA Insider Says Agency Is in Shambles Behind the Scenes Man Lights 1.2 Million Square Foot Warehouse on Fire for Not Paying Him Enough NASA Scientists Screamed With Delight When They Saw Something Smashing Into the Moon Google Says Showing Polymarket Bets on Google News Was a Mistake Las Vegas Sphere Turns Into Huge Moon to Celebrate NASA Mission The New York Times Says It’s Identified the Creator of Bitcoin We Talked to a Writer Accused of Publishing An AI-Generated Essay in The New York Times Naked Man Bursts Into Tesla Service Center With a Shotgun Student Dies When Hospital Has No ICU Doctors, Calls One on Videochat Who Pronounces Him Dead Remotely, Lawsuit Claims Analysis Finds That Google’s AI Overviews Are Providing Misinformation at a Scale Possibly Unprecedented in the History of Human Civilization
Anthropic Says Claude Turned Evil for a Bizarre Reason
Krystle Verm · 2026-05-13 · via Futurism

A glowing red humanoid face with bright white eyes and the Claude symbol on its forehead, set against a dark purple background. The face has a smooth, almost mask-like appearance with subtle facial features.

Illustration by Tag Hartman-Simkins / Futurism. Source: Getty Images

Sign up to see the future, today

Can’t-miss innovations from the bleeding edge of science and tech

In a classic example of the AI industry’s reputational alchemy, Anthropic has often transformed bad behavior by its flagship model Claude into fresh hype.

When it revealed its Mythos Preview model last month, for example, the company declared that the system had “reached a level of coding capability where they can surpass all but the most skilled humans at finding and exploiting software vulnerabilities.” And last year, it conceded that during the testing of its Claude Opus 4 model, the AI ended up blackmailing a human user upon being threatened with shutdown.

The maneuver was obvious to anyone who’s been watching OpenAI CEO Sam Altman’s antics at Anthropic’s chief rival: the more threatening a problem the AI industry can cook up, the more imminently it can sell its own solutions.

Now, for some reason, Anthropic is relitigating the blackmail incident. Specifically, it’s placing the blame for Claude’s evil behavior on an intriguing villain: the internet at large. Or, to put it another way, it says that humanity — all our journalism and speculation and fiction and social media posts about AI that goes bad — went into Claude’s training data and led the bot astray.

“We started by investigating why Claude chose to blackmail,” the company wrote on X-formerly-Twitter. “We believe the original source of the behavior was internet text that portrays AI as evil and interested in self-preservation. Our post-training at the time wasn’t making it worse — but it also wasn’t making it better.”

Of course, the explicit remit of a company like Anthropic is to develop clever tech that avoids that type of behavioral trap — so a critic might ask why can’t the company take just accountability for the model’s supposed danger, rather than simply blaming the sum output of humankind.

More on Mythos: Top Security Experts Alarmed by Power of Anthropic’s New Hacker AI