惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
MongoDB | Blog
MongoDB | Blog
博客园_首页
博客园 - 三生石上(FineUI控件)
博客园 - 聂微东
B
Blog RSS Feed
D
Docker
IT之家
IT之家
大猫的无限游戏
大猫的无限游戏
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
阮一峰的网络日志
阮一峰的网络日志
罗磊的独立博客
Recent Announcements
Recent Announcements
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
A
About on SuperTechFans
The GitHub Blog
The GitHub Blog
G
Google Developers Blog
V
V2EX
量子位
雷峰网
雷峰网
月光博客
月光博客
云风的 BLOG
云风的 BLOG
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
T
Tailwind CSS Blog

Futurism

Man Creates Tiny Submarine for His Parakeet to Experience Life Underwater The Effects of AI-Generated Code Tearing Through Corporations Is Actually Kind of Funny Trump Hires Orbital Towing Company to Build Space Interceptors Psychologists Found Something Horrible About the Kind of Men Seeking Trad Wives To Get Swole, Teens Are Pumping Themselves Full of Drugs Meant for Fattening Cows for the Slaughterhouse Foolish Pollsters Are Now Just Asking AI What Voters Would Say in Response to Questions and Publishing It at Face Value OpenAI Says It’s Already Made $100 Million by Stuffing ChatGPT With Ads Man Punished for Breaking Into Moo Deng’s Zoo Enclosure AI Is Causing Healthcare Costs to Surge There’s a Mass Rebellion Against AI in the Workplace People Who Lose Their Job to AI Are in for a World of Pain, Goldman Sachs Report Finds OpenAI Says Not to Worry About UBI, Because It Has Another Idea Police Officer Helplessly Waves Arms at Waymo That Careened Wrong Way Through Whataburger Drive-Thru Someone Just Threw a Molotov Cocktail At Sam Altman’s House New York Times Makes Substantial Changes to Article That Glazed a Sleazy AI Startup: “Our Piece Should Have Included That Information” Space Scientists Wince as Astronauts’ Lives Depend on Artemis 2’s Controversial Heat Shield During Plunge Back to Earth The Moon Astronauts Have Been Working Out With a NASA Rowing Machine in Space First AI Model From Zuckerberg’s Wildly Expensive Superintelligence Lab Flops Compared to Virtually All Rivals Economists Starting to Admit They May Have Been Wrong About AI Never Replacing Human Jobs AI-Powered Drug Marketer Medvi Responds After Allegations About Fake Doctors and Patients As Astronauts Visit the Moon, NASA Insider Says Agency Is in Shambles Behind the Scenes Man Lights 1.2 Million Square Foot Warehouse on Fire for Not Paying Him Enough NASA Scientists Screamed With Delight When They Saw Something Smashing Into the Moon Google Says Showing Polymarket Bets on Google News Was a Mistake Las Vegas Sphere Turns Into Huge Moon to Celebrate NASA Mission The New York Times Says It’s Identified the Creator of Bitcoin We Talked to a Writer Accused of Publishing An AI-Generated Essay in The New York Times Naked Man Bursts Into Tesla Service Center With a Shotgun Student Dies When Hospital Has No ICU Doctors, Calls One on Videochat Who Pronounces Him Dead Remotely, Lawsuit Claims Analysis Finds That Google’s AI Overviews Are Providing Misinformation at a Scale Possibly Unprecedented in the History of Human Civilization
Anthropic Was So Concerned About Its New Mythos-Based Mod...
Victor Tangermann · 2026-06-12 · via Futurism

A stylized photo illustration featuring Anthropic co-founder Dario Amodei.

Illustration by Tag Hartman-Simkins / Futurism. Source: Michael M. Santiago / Getty Images; Shutterstock

Sign up to see the future, today

Can’t-miss innovations from the bleeding edge of science and tech

Earlier this year, Anthropic refused to release its Mythos AI model to the public, saying it was simply too dangerous.

At the time, executives claimed the model was capable of punching through powerful cybersecurity safeguards, pointing at researchers who used it to discover thousands of vulnerabilities in widely-used open source code.

Months later, Anthropic was finally ready to go public with the model. On Tuesday, the Dario Amodei-led company announced a Mythos-powered model called Fable 5, which it claims is “safe for general use.”

However, new safeguards quickly frustrated AI researchers, who accused the company of intentionally lobotomizing Fable 5. The backlash was so fierce, Anthropic quickly made adjustments to the policy, as Wired reported on Wednesday, highlighting just how carefully the company is treading.

In its original announcement, Anthropic claimed the safeguards were designed to stop Fable 5 from improving itself, in “new interventions that limit Claude’s effectiveness for requests targeting frontier LLM development.” Just days ahead of the launch, Anthropic released a report on “when AI builds itself,” a trend that “might increase the risks of humans losing control over AI systems.”

However, AI researchers were not impressed by Anthropic hamstringing its latest model’s abilities.

“Anthropic’s latest model will NOT help you if it thinks your ML research/ML engineering is interesting, and/or will secretly degrade its IQ so that the average engineer won’t notice,” AI research firm SemiAnalysis tweeted.

“We are already seeing Anthropic’s latest model’s moderation filters our GPU inference research and programming,” it added.

Other researchers accused Anthropic of using Fable 5 to “shadowban,” or quietly restrict the accounts, of AI researchers. According to the firm’s system card, interventions limiting requests for “frontier LLM development” will “not be visible to the user.”

This last concern, which could’ve effectively sabotaged anybody trying to train competing models by quietly bumping them down to less powerful models without their knowledge, proved controversial enough for Anthropic to change its mind.

“We’re changing Fable 5’s safeguards for frontier LLM development to make them visible,” the company told Wired in a statement. “We made the wrong trade-off and we apologize for not getting the balance right.”

“It felt like Anthropic was saying to the public, ‘We don’t trust anybody else to do AI research,” AI startup Prime Intellect research lead Will Brown told the publication. “We are the only ones who have to do AI research.”

It all comes in the context Anthropic calling for a global freeze on AI advances while discussing the dangers of “recursive self-improvement.” In other words, the company is making a lot of noise about a sci-fi-sounding possibility: that AI will start to rapidly improve itself, potentially escaping the control of its human creators.

Beyond limiting its ability to develop AI tools, Fable 5’s new safeguards also trigger when it encounters requests “related to cybersecurity, biology and chemistry, or distillation.” Distillation is effectively using machine learning to train a “student” model on the behavior and reasoning of a “teacher” model, a practice that has sparked its fair share of controversy.

Anthropic has already publicly griped about large-scale attempts to distill, or “extract” its underlying model — a hypocritical stance given its indiscriminate scraping of rights-protected content on the web to train its AI in the first place.

More on Anthropic: Anthropic Scared, Calls for Global Freeze on AI Advances