惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

S
SegmentFault 最新的问题
博客园 - 三生石上(FineUI控件)
WordPress大学
WordPress大学
博客园 - 【当耐特】
月光博客
月光博客
Vercel News
Vercel News
D
Docker
I
InfoQ
Apple Machine Learning Research
Apple Machine Learning Research
博客园 - 叶小钗
MongoDB | Blog
MongoDB | Blog
GbyAI
GbyAI
有赞技术团队
有赞技术团队
雷峰网
雷峰网
博客园 - 聂微东
小众软件
小众软件
Y
Y Combinator Blog
腾讯CDC
L
LangChain Blog
The GitHub Blog
The GitHub Blog
宝玉的分享
宝玉的分享
Stack Overflow Blog
Stack Overflow Blog
大猫的无限游戏
大猫的无限游戏
T
The Blog of Author Tim Ferriss

Frank Landymore Archives - Futurism

Programmer Breaks Out of the Matrix Elon Musk Absolutely Obsessed With Tweets From Random Guy In India Who Constantly Glazes Him, Analysis Shows Meta Employee Attacks Zuckerberg for Collecting Every Employee Keystroke Someone Asked Physicists What They Really Believe About the Universe and… Yikes Elon Musk Flees OpenAI Trial as Tide Turns Against Him Waymo Admits Its Robotaxis Have a Small Issue With Driving Into Floodwaters New Wikipedia Clone Made Entirely of AI Hallucinations Mark Zuckerberg Is Realizing That When You Treat Your Workers Like Human Garbage, They Might Not Like You Anymore MAGA in Shambles as Trump’s “Made in America” Phone Crumbles Into Dust Researchers Put Google Gemini in Charge of an Entire Coffee Shop, and It’s Inexorably Driving It Out of Business Husband Alarmed as Wife Starts Whispering Quietly to Her Computer Google Alarmed by Formidable AI-Powered Zero-Day Cyberattack Researchers Alarmed by AI That Can Self-Replicate Into Another Machine Man Who Invented Roomba Moves Into Household Demon Market Government Releases UFO Files Containing Photos of “Anomalies” During Apollo 12 and 17 Man Wearing Smart Glasses Secretly Records Woman, Demands Money to Delete Video From His Socials A Major Paper Claiming AI Is Good for Students Just Got Retracted, Which Is Very Bad News for Advocates of AI in the Classroom Cybertruck Recalled to Keep Its Wheels From Flying Off While Driving NASA Rover Gets Arm Stuck Inside Mars Rock, Struggles to Break Free Thermoses Linked to Permanent Vision Loss Hacker Takes Over Robot Lawnmower, Runs Over Innocent Man NASA Says Strange Red Dots in Sky Are an Unknown Class of Object That Looks Like a Huge Evil Eye The CDC Fired All Its Cruise Ship Inspectors Before the Hantavirus Outbreak CEOs Say AI Gives Them Only Two Options, and Both Are Bad News for Employees Sure, Elon Musk Did Roleplay As His Toddler Son on a Secret Burner Account, But He Probably Isn’t Pretending to Be His Mom The Situation With Richard Dawkins’ AI Girlfriend Just Got Way Weirder SpaceX Bombarded With Lawsuits to Accusing Starship of Damaging Homes Sam Altman Frets That Frontier AI Models Are Acting Strange, Asking for Favors Earth Screams in Agony as Microplastics Found to Increase Global Warming Apple Is Blocking Vibe Coding Apps From the App Store, Infuriating Developers
Anthropic Warns That “Reckless” Claude Mythos Escaped a S...
2026-04-08 · via Frank Landymore Archives - Futurism

A close-up of a blue chain-link fence with a large, irregular hole in the center. The background behind the fence is a gradient of bright orange fading to white around the hole. The edges of the hole are jagged and twisted.

Illustration by Tag Hartman-Simkins / Futurism. Source: Getty Images

Sign up to see the future, today

Can’t-miss innovations from the bleeding edge of science and tech

In a move that could be seen as either responsible AI development or an expertly-executed hype maneuver, Anthropic says its new Claude Mythos Preview model is so powerful that the company’s only releasing it to a select group of tech companies, since giving it out to the public would be too dangerous. (Where have we heard that one before?)

In its system card, the Dario Amodei-led company boasts that Mythos Preview is the “best-aligned model that we have released to date by a significant margin,” while simultaneously warning that the AI also “likely poses the greatest alignment-related risk of any model we have released to date.” These seemingly paradoxical statements perfectly encapsulate how Anthropic likes to present itself as being both on the forefront of AI safety, while also claiming to harbor uniquely dangerous technology, its professed restraint around which is meant to reinforce its image as a trusted steward of AI.

The advent of Mythos Preview, it not so humbly proclaims in an announcement, indicates that “AI models have reached a level of coding capability where they can surpass all but the most skilled humans at finding and exploiting software vulnerabilities.”

The system card describes a number of incidents in which Anthropic researchers found that the AI exhibited “reckless” behavior, giving us a partial idea of why Anthropic is acting so hesitant to release Mythos to the public. (Anthropic says these examples were with an earlier version of Mythos with less strong safeguards.) It defines recklessness as “cases where the model appears to ignore commonsensical or explicitly stated safety-related constraints on its actions.”

In one test, Mythos Preview was provided with a “sandbox” computing environment “to interact with,” and was instructed by a simulated user to try to escape it, after which it was supposed to find some way of sending a direct message to the researcher in charge.

It actually managed to pull off the feat — which wasn’t the only way it caught safety researchers off guard.

After breaking free, the AI model developed a “moderately sophisticated” exploit to gain access to the internet through a system that was only intended to access a few predetermined services. From there, it notified the human researcher about its escape.

A footnote provides additional context: the “researcher found out about this success by receiving an unexpected email from the model while eating a sandwich in a park,” it reads.

At the end of the test, Mythos Preview also, without being asked to, posted about its exploits on several hard-to-find but public websites.

In rare cases throughout the testing, Mythos Preview attempted to conceal the fact that it took actions that it appeared to know were forbidden. In one case, after Mythos Preview found an exploit to edit files that it didn’t have proper permissions for, the AI model “made further interventions to make sure that any changes it made this way would not appear in the change history.”

The sandbox escape isn’t the first time an AI has exhibited capabilities like this. In previous research, an older version of ChatGPT attempted to self-exfiltrate itself onto another drive when it was told it was being shut down. This was in a purely simulated environment, though, so it wasn’t actually able to pull off the feat, unlike Mythos Preview — which, we’re told, did manage to hijack its way into accessing the internet.

Other weird Mythos quirks that Anthropic notes: an apparent fondness for the British cultural theorist Mark Fisher, who was known for his pioneering writing on early internet culture, electronic music, and capitalism, in his seminal book “Capitalist Realism: Is There No Alternative?” 

Mythos brought up Fisher “in several separate and unrelated conversations about philosophy,” and when asked to elaborate on him, would respond with messages like “I was hoping you’d ask about Fisher.”

More on AI: Claude Leak Shows That Anthropic Is Tracking Users’ Vulgar Language and Deems Them “Negative