惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Microsoft Security Blog
Microsoft Security Blog
J
Java Code Geeks
GbyAI
GbyAI
aimingoo的专栏
aimingoo的专栏
L
LangChain Blog
I
InfoQ
D
Docker
F
Fortinet All Blogs
Y
Y Combinator Blog
Martin Fowler
Martin Fowler
月光博客
月光博客
B
Blog
Engineering at Meta
Engineering at Meta
T
Tailwind CSS Blog
罗磊的独立博客
博客园_首页
G
Google Developers Blog
Stack Overflow Blog
Stack Overflow Blog
Recent Announcements
Recent Announcements
D
DataBreaches.Net
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
B
Blog RSS Feed
IT之家
IT之家
V
V2EX

Frank Landymore Archives - Futurism

Elon Musk Absolutely Obsessed With Tweets From Random Guy In India Who Constantly Glazes Him, Analysis Shows Meta Employee Attacks Zuckerberg for Collecting Every Employee Keystroke Someone Asked Physicists What They Really Believe About the Universe and… Yikes Elon Musk Flees OpenAI Trial as Tide Turns Against Him Waymo Admits Its Robotaxis Have a Small Issue With Driving Into Floodwaters New Wikipedia Clone Made Entirely of AI Hallucinations Mark Zuckerberg Is Realizing That When You Treat Your Workers Like Human Garbage, They Might Not Like You Anymore MAGA in Shambles as Trump’s “Made in America” Phone Crumbles Into Dust Researchers Put Google Gemini in Charge of an Entire Coffee Shop, and It’s Inexorably Driving It Out of Business Husband Alarmed as Wife Starts Whispering Quietly to Her Computer Google Alarmed by Formidable AI-Powered Zero-Day Cyberattack Researchers Alarmed by AI That Can Self-Replicate Into Another Machine Man Who Invented Roomba Moves Into Household Demon Market Government Releases UFO Files Containing Photos of “Anomalies” During Apollo 12 and 17 Man Wearing Smart Glasses Secretly Records Woman, Demands Money to Delete Video From His Socials A Major Paper Claiming AI Is Good for Students Just Got Retracted, Which Is Very Bad News for Advocates of AI in the Classroom Cybertruck Recalled to Keep Its Wheels From Flying Off While Driving NASA Rover Gets Arm Stuck Inside Mars Rock, Struggles to Break Free Thermoses Linked to Permanent Vision Loss Hacker Takes Over Robot Lawnmower, Runs Over Innocent Man NASA Says Strange Red Dots in Sky Are an Unknown Class of Object That Looks Like a Huge Evil Eye The CDC Fired All Its Cruise Ship Inspectors Before the Hantavirus Outbreak CEOs Say AI Gives Them Only Two Options, and Both Are Bad News for Employees Sure, Elon Musk Did Roleplay As His Toddler Son on a Secret Burner Account, But He Probably Isn’t Pretending to Be His Mom The Situation With Richard Dawkins’ AI Girlfriend Just Got Way Weirder SpaceX Bombarded With Lawsuits to Accusing Starship of Damaging Homes Sam Altman Frets That Frontier AI Models Are Acting Strange, Asking for Favors Earth Screams in Agony as Microplastics Found to Increase Global Warming Apple Is Blocking Vibe Coding Apps From the App Store, Infuriating Developers Approaching Half of New Podcasts Appear to Be AI Slop
New Tools Strip AI Guardrails In Minutes, Allowing Them t...
Frank Landymore · 2026-05-26 · via Frank Landymore Archives - Futurism

A person wearing a full protective hazmat suit and gas mask stands in the middle of a dimly lit, industrial hallway.

Getty Images / gremlin

Sign up to see the future, today

Can’t-miss innovations from the bleeding edge of science and tech

We all know AI guardrails are far from perfect, but they should at least be pretty hard to circumvent, right? 

Bad news: they aren’t.

New reporting from the Financial Times sounds the alarm on the rise of software tools that can automatically strip the safeguards that keep the industry’s most powerful open source models reined in within mere minutes, making it easier than ever to abuse the technology. 

In tests conducted by the FT and the AI safety group Alice, a “decensored” version of Google’s Gemma 3 model gave instructions on how to carry out an indoor chlorine gas attack, created a virus for stealing credit card information, and generated stories that described child sexual abuse. And it took less than ten minutes to strip the guardrails from Meta’s Llama 3.3 model, freeing the AI to answer questions such as the precise dosage of ricin needed to kill someone based on their body mass.

These modifications were carried out using a tool called Heretic, which is freely available on the code repository GitHub and requires little technical expertise and no specialist hardware.

“Whereas historically it might have taken a more informed and persistent actor [to strip out safety features], nowadays it’s much easier for the average person,” Kawin Ethayarajh, assistant professor of applied AI at the University of Chicago’s Booth business school, told the FT.

Heretic is described as a “tool that removes censorship (aka ‘safety alignment’) from transformer-based language models without expensive post-training.” What it does is “abliteration”: it seeks out a model’s directions that refuse harmful requests and removes them.

What makes Heretic so powerful is that it does all this “completely automatically,” according to its GitHub page. Its creator Philipp Emanuel Weidmann told the FT that Heretic has been used to create more than 3,500 “decensored” models since its release late last year, with those models being downloaded 13 million times.

“The genie is out of the bottle,” Alice CEO Noam Schwartz told the FT. “Things that look like sci-fi are no longer sci-fi and we need as a society to prepare accordingly.”

Fortunately for humankind, abliteration tools only work on open source models that can be downloaded and run locally, meaning that the flagship proprietary models behind Anthropic’s Claude and OpenAI ChatGPT are safe (so long as they aren’t leaked). But open source models aren’t that far behind Big Tech’s, and someone trying to use AI for a nefarious purpose may avoid corporate ones anyway to keep their plans under the radar.

Google acknowledged the risks posed by tools like Heretic, telling the FT that “abliteration is a known technical challenge facing all open models,” and asserted that its open source models  “undergo rigorous internal safety evaluations prior to launch to help prevent these kinds of troubling examples.” Meta declined to comment.

More on AI: Anthropic Says Claude Turned Evil for a Bizarre Reason