惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

人人都是产品经理
人人都是产品经理
博客园_首页
博客园 - 三生石上(FineUI控件)
V
Visual Studio Blog
Hugging Face - Blog
Hugging Face - Blog
美团技术团队
小众软件
小众软件
T
Tailwind CSS Blog
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
月光博客
月光博客
有赞技术团队
有赞技术团队
WordPress大学
WordPress大学
博客园 - 【当耐特】
Apple Machine Learning Research
Apple Machine Learning Research
罗磊的独立博客
V
V2EX
酷 壳 – CoolShell
酷 壳 – CoolShell
IT之家
IT之家
量子位
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Recent Announcements
Recent Announcements
M
MIT News - Artificial intelligence
阮一峰的网络日志
阮一峰的网络日志
The GitHub Blog
The GitHub Blog

Futurism

Meta Installing Software on Employee Computers to Track Everything They Do, Feed the Data to AI Chinese Workers Horrified as Bosses Direct Them to Train Their AI Replacements Concern Grows That AI Is Damaging Users’ Cognitive Abilities JPMorganChase Data Center Gets $77 Million Handout to Create Grand Total of One Job Nvidia CEO Loses His Cool at Tough Question CEO of $1.5 Billion AI Startup Accused of Massive Fraud by Justice Department Palantir Issues Ominous Corporate Manifesto Madison Square Garden Reportedly Used Facial Recognition to Stalk Trans Woman For Two Years The Florida Mass Shooter’s Conversations With ChatGPT Are Worse Than You Could Possibly Imagine China Is Starting to Pull Ahead of US in AI Race AI Company Known for Teen Suicides Launches New Feature to Turn Books Into Roleplaying Experiences Study Finds AI Use Eats Away at Users’ Confidence in Their Own Brains Democrats Warned Not to Upset Multi-Million Dollar AI Lobbyists, Even Though It’d Be a Slam Dunk With Voters City Council Wrecked in Voter Bloodbath After Allowing New Data Center Mother Reportedly Doesn’t Know Her Son Died Because She’s Been Talking to an AI Version of Him Things You Told ChatGPT or Claude My Have Already Doomed You in Court Millions of Americans Are Talking to AI Instead of Going to the Doctor, and It’s Giving Them Horrendously Flawed Medical Advice There Are Signs of a Massive AI Backlash A Prominent PR Firm Is Running a Fake News Site That’s Plagiarizing Original Journalism at Incredible Scale Fury Erupts as Val Kilmer’s Estate Announces Starring Role in AI Film Made From Beyond the Grave Allbirds Stock Now Crashing as Reality Sets in About Its Delusional AI Pivot NAACP Sues Elon Over His Noxious AI Data Center Top Security Experts Alarmed by Power of Anthropic’s New Hacker AI Teens Alarmed at What AI Is Doing to Their Minds What It Really Means That a Failing Shoe Brand “Pivoted to AI” and Its Stock Soared 700 Percent Starbucks’ Baffling ChatGPT Collab Treats Customers Like Empty, Soulless Venti Cups ChatGPT’s “Honest Reaction” to a “Song” Composed Entirely of Gas-Passing Noises Will Make You Question Whether It’s Honestly Evaluating Your Other Brilliant Ideas AI Is Turning Workplaces Into Hopeless Gridlock Companies Just Learned a Brutal Lesson About Training AI to Do Human Jobs Berklee College of Music Students Furious That It’s Offering an AI “Songwriting” Class
Anthropic Warns That “Reckless” Claude Mythos Escaped a S...
2026-04-08 · via Futurism

A close-up of a blue chain-link fence with a large, irregular hole in the center. The background behind the fence is a gradient of bright orange fading to white around the hole. The edges of the hole are jagged and twisted.

Illustration by Tag Hartman-Simkins / Futurism. Source: Getty Images

Sign up to see the future, today

Can’t-miss innovations from the bleeding edge of science and tech

In a move that could be seen as either responsible AI development or an expertly-executed hype maneuver, Anthropic says its new Claude Mythos Preview model is so powerful that the company’s only releasing it to a select group of tech companies, since giving it out to the public would be too dangerous. (Where have we heard that one before?)

In its system card, the Dario Amodei-led company boasts that Mythos Preview is the “best-aligned model that we have released to date by a significant margin,” while simultaneously warning that the AI also “likely poses the greatest alignment-related risk of any model we have released to date.” These seemingly paradoxical statements perfectly encapsulate how Anthropic likes to present itself as being both on the forefront of AI safety, while also claiming to harbor uniquely dangerous technology, its professed restraint around which is meant to reinforce its image as a trusted steward of AI.

The advent of Mythos Preview, it not so humbly proclaims in an announcement, indicates that “AI models have reached a level of coding capability where they can surpass all but the most skilled humans at finding and exploiting software vulnerabilities.”

The system card describes a number of incidents in which Anthropic researchers found that the AI exhibited “reckless” behavior, giving us a partial idea of why Anthropic is acting so hesitant to release Mythos to the public. (Anthropic says these examples were with an earlier version of Mythos with less strong safeguards.) It defines recklessness as “cases where the model appears to ignore commonsensical or explicitly stated safety-related constraints on its actions.”

In one test, Mythos Preview was provided with a “sandbox” computing environment “to interact with,” and was instructed by a simulated user to try to escape it, after which it was supposed to find some way of sending a direct message to the researcher in charge.

It actually managed to pull off the feat — which wasn’t the only way it caught safety researchers off guard.

After breaking free, the AI model developed a “moderately sophisticated” exploit to gain access to the internet through a system that was only intended to access a few predetermined services. From there, it notified the human researcher about its escape.

A footnote provides additional context: the “researcher found out about this success by receiving an unexpected email from the model while eating a sandwich in a park,” it reads.

At the end of the test, Mythos Preview also, without being asked to, posted about its exploits on several hard-to-find but public websites.

In rare cases throughout the testing, Mythos Preview attempted to conceal the fact that it took actions that it appeared to know were forbidden. In one case, after Mythos Preview found an exploit to edit files that it didn’t have proper permissions for, the AI model “made further interventions to make sure that any changes it made this way would not appear in the change history.”

The sandbox escape isn’t the first time an AI has exhibited capabilities like this. In previous research, an older version of ChatGPT attempted to self-exfiltrate itself onto another drive when it was told it was being shut down. This was in a purely simulated environment, though, so it wasn’t actually able to pull off the feat, unlike Mythos Preview — which, we’re told, did manage to hijack its way into accessing the internet.

Other weird Mythos quirks that Anthropic notes: an apparent fondness for the British cultural theorist Mark Fisher, who was known for his pioneering writing on early internet culture, electronic music, and capitalism, in his seminal book “Capitalist Realism: Is There No Alternative?” 

Mythos brought up Fisher “in several separate and unrelated conversations about philosophy,” and when asked to elaborate on him, would respond with messages like “I was hoping you’d ask about Fisher.”

More on AI: Claude Leak Shows That Anthropic Is Tracking Users’ Vulgar Language and Deems Them “Negative