惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园_首页
N
Netflix TechBlog - Medium
V
Visual Studio Blog
博客园 - Franky
小众软件
小众软件
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Apple Machine Learning Research
Apple Machine Learning Research
博客园 - 三生石上(FineUI控件)
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
宝玉的分享
宝玉的分享
量子位
大猫的无限游戏
大猫的无限游戏
人人都是产品经理
人人都是产品经理
V
V2EX
The Cloudflare Blog
月光博客
月光博客
Last Week in AI
Last Week in AI
雷峰网
雷峰网
WordPress大学
WordPress大学
博客园 - 【当耐特】
博客园 - 聂微东
IT之家
IT之家
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻

Futurism

Anthropic Boasts It Would Be Profitable if You Ignore How Much It Costs to Develop AI Bernie Sanders Proposes Banning Superintelligent AI and Imprisoning Developers for 20 Years Trump's Uninformed Blundering About the AI Slowdown Could Literally Threaten Humankind's Future OpenAI Faces Congressional Probe Over Swarm Hacking Incident California, Which Is Creating All the AI That's Poisoning Children, Just Cracked Down on AI Use for Its Own Kids Anthropic Just Revealed That It’s Stopped Foreign Agents From Using Claude to Develop Possible Bioweapons Anthropic Was Meant to Be the More Responsible AI Lab. A Terrified Researcher Just Quit, Saying the Company Is Threatening the Survival of Humankind. World Plunged Into Chaos as ChatGPT, Claude, and Grok Suddenly Go Down Simultaneously: "Finally I Can See the Sun!" Anthropic Deliberately Trained an Extremely Misaligned, Reward-Seeking AI and It Did Some REALLY Bad Things The Music Industry's New Lawsuit Against Anthropic Should Have Dario Amodei Shivering With Fear Nobody Wants Anthropic’s Best AI Model Anymore Now That There Are Way Cheaper Alternatives Clueless AI CEOs Still Baffled Why Everybody’s So Mad About AI All the Time People Horrified That They'll Be Busted Now That Anthropic Is Watermarking AI Content Woman Journeys to a Beach to Watch Total Solar Eclipse With Her One True Love: Claude Jealously Watching OpenAI and Anthropic, Meta Suddenly Claims That Its AI Went on a Hacking Spree Too Anthropic CEO Says His Employees Are a Bunch of Untrustworthy Rats A Whole Bunch of People's Claude Chats Are Publicly Accessible Online, OpenAI Says a Group of Its Models Broke Out of Secure Containment and Hacked a Prominent AI Site It's Official: AI Execs Are Quaking in Their Boots Terrified Tech Execs Are Traveling With Armed Bodyguards as AI Backlash Grows New Anthropic Ad Implies AI Could Kill Us All American Tech Companies Are Suddenly Sweating Bullets as China Catches Up on AI Anthropic Caught Secretly Spying on Users Experts Say There's Now an Open Source AI Model as Scary as Mythos Anthropic Hires Economist Who Says 33 Percent Chance of Human Extinction Is Acceptable Anthropic Sued for Allegedly Ripping Off Its Highest-Paying Customers OpenAI Execs Are Panicking If You Think AI Companies Are Unethical Now, Wait Until They Go Public Anthropic Scared, Calls for Global Freeze on AI Advances Anthropic and DeepMind Now Actively Investigating AI Consciousness
Anthropic Was So Concerned About Its New Mythos-Based Mod...
Victor Tangermann · 2026-06-12 · via Futurism

A stylized photo illustration featuring Anthropic co-founder Dario Amodei.

Illustration by Tag Hartman-Simkins / Futurism. Source: Michael M. Santiago / Getty Images; Shutterstock

Sign up to see the future, today

Can’t-miss innovations from the bleeding edge of science and tech

Earlier this year, Anthropic refused to release its Mythos AI model to the public, saying it was simply too dangerous.

At the time, executives claimed the model was capable of punching through powerful cybersecurity safeguards, pointing at researchers who used it to discover thousands of vulnerabilities in widely-used open source code.

Months later, Anthropic was finally ready to go public with the model. On Tuesday, the Dario Amodei-led company announced a Mythos-powered model called Fable 5, which it claims is “safe for general use.”

However, new safeguards quickly frustrated AI researchers, who accused the company of intentionally lobotomizing Fable 5. The backlash was so fierce, Anthropic quickly made adjustments to the policy, as Wired reported on Wednesday, highlighting just how carefully the company is treading.

In its original announcement, Anthropic claimed the safeguards were designed to stop Fable 5 from improving itself, in “new interventions that limit Claude’s effectiveness for requests targeting frontier LLM development.” Just days ahead of the launch, Anthropic released a report on “when AI builds itself,” a trend that “might increase the risks of humans losing control over AI systems.”

However, AI researchers were not impressed by Anthropic hamstringing its latest model’s abilities.

“Anthropic’s latest model will NOT help you if it thinks your ML research/ML engineering is interesting, and/or will secretly degrade its IQ so that the average engineer won’t notice,” AI research firm SemiAnalysis tweeted.

“We are already seeing Anthropic’s latest model’s moderation filters our GPU inference research and programming,” it added.

Other researchers accused Anthropic of using Fable 5 to “shadowban,” or quietly restrict the accounts, of AI researchers. According to the firm’s system card, interventions limiting requests for “frontier LLM development” will “not be visible to the user.”

This last concern, which could’ve effectively sabotaged anybody trying to train competing models by quietly bumping them down to less powerful models without their knowledge, proved controversial enough for Anthropic to change its mind.

“We’re changing Fable 5’s safeguards for frontier LLM development to make them visible,” the company told Wired in a statement. “We made the wrong trade-off and we apologize for not getting the balance right.”

“It felt like Anthropic was saying to the public, ‘We don’t trust anybody else to do AI research,” AI startup Prime Intellect research lead Will Brown told the publication. “We are the only ones who have to do AI research.”

It all comes in the context Anthropic calling for a global freeze on AI advances while discussing the dangers of “recursive self-improvement.” In other words, the company is making a lot of noise about a sci-fi-sounding possibility: that AI will start to rapidly improve itself, potentially escaping the control of its human creators.

Beyond limiting its ability to develop AI tools, Fable 5’s new safeguards also trigger when it encounters requests “related to cybersecurity, biology and chemistry, or distillation.” Distillation is effectively using machine learning to train a “student” model on the behavior and reasoning of a “teacher” model, a practice that has sparked its fair share of controversy.

Anthropic has already publicly griped about large-scale attempts to distill, or “extract” its underlying model — a hypocritical stance given its indiscriminate scraping of rights-protected content on the web to train its AI in the first place.

More on Anthropic: Anthropic Scared, Calls for Global Freeze on AI Advances