惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

有赞技术团队
有赞技术团队
H
Hackread – Cybersecurity News, Data Breaches, AI and More
I
InfoQ
J
Java Code Geeks
Microsoft Security Blog
Microsoft Security Blog
G
Google Developers Blog
D
DataBreaches.Net
Recent Announcements
Recent Announcements
Microsoft Azure Blog
Microsoft Azure Blog
B
Blog RSS Feed
Y
Y Combinator Blog
博客园 - 【当耐特】
博客园 - 聂微东
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
大猫的无限游戏
大猫的无限游戏
P
Proofpoint News Feed
量子位
C
Check Point Blog
F
Fortinet All Blogs
罗磊的独立博客
Last Week in AI
Last Week in AI
GbyAI
GbyAI
L
LangChain Blog
博客园 - 司徒正美

Futurism

Frontier AI Models Giving Specific, Actionable Instructions to Perpetrate Bioterror Attack AI Slop YouTube Channel Glitches Out in a Way So Bizarre That It’s Vaguely Disturbing Double Murder Suspect Asked ChatGPT How to Hide Body in Dumpster An Elegant Solution to AI Slop: Tax It, and Use the Resulting Billions of Dollars to Fund Cultural Institutions, Artists, and Researchers The White House Suddenly Seems Pretty Terrified of Anthropic Democrat and Republican Voters United on Key Issue: Hatred of Data Centers Chinese Court Rules That a Worker Cannot Be Replaced by AI Toilet Maker Spikes in Value as It Flushes Money Into AI New England Journal of Medicine Retracts Paper Because Photo of Patient’s Insides Was Garbled by AI Gen Z Is Turning Against AI in an Incredible Way If OpenAI Loses This Trial, It Could Effectively Be Eliminated in Its Current Form AI Spy Cameras Suddenly Blanketing America Man Trapped in Dystopian Nightmare Thanks to AI Surveillance Cameras Flagging His Every Move John Oliver Just Took the AI Industry Behind a Shed and Beat It With a Pipe Wrench OpenAI Hit With Barrage of Lawsuits Over Failure to Report School Shooter Before Massacre Police Are Using AI Camera Networks to Stalk Women Sam Altman Caught in What May Be His Most Spectacular Lie Yet OpenAI in Shambles as IPO Looms A Tiny Town Is Building So Many Data Centers That There’ll Be Almost Nothing Else Left Weird Things Happen When You Give AI Agents Money and Let Them Spend It Sam Altman Issues Grim Apology Top Medical Journal Publishes Searing Article Warning Against Medical AI New Browser Plugin Adds Typos to Your AI-Generated Emails to Make Them Look Real Experts Warn of AI Swarms Hijacking Democracy With Fake Citizens Devious New AI Tool “Clones” Software So That the Original Creator Doesn’t Hold a Copyright Over the New Version Prestigious Wall Street Law Firm Humiliated When Its AI Use Is Discovered in Court Unions Attack AI for Menacing Human Jobs Your Former Employer Is Selling Your Slacks and Emails to Train AI Three Years Ago Today, “Avengers” Director Joe Russo Predicted There Would Be a Fully AI-Generated Movie Within Two Years Palantir’s Employees Are in Crisis
New Tools Strip AI Guardrails In Minutes, Allowing Them t...
Frank Landymore · 2026-05-26 · via Futurism

A person wearing a full protective hazmat suit and gas mask stands in the middle of a dimly lit, industrial hallway.

Getty Images / gremlin

Sign up to see the future, today

Can’t-miss innovations from the bleeding edge of science and tech

We all know AI guardrails are far from perfect, but they should at least be pretty hard to circumvent, right? 

Bad news: they aren’t.

New reporting from the Financial Times sounds the alarm on the rise of software tools that can automatically strip the safeguards that keep the industry’s most powerful open source models reined in within mere minutes, making it easier than ever to abuse the technology. 

In tests conducted by the FT and the AI safety group Alice, a “decensored” version of Google’s Gemma 3 model gave instructions on how to carry out an indoor chlorine gas attack, created a virus for stealing credit card information, and generated stories that described child sexual abuse. And it took less than ten minutes to strip the guardrails from Meta’s Llama 3.3 model, freeing the AI to answer questions such as the precise dosage of ricin needed to kill someone based on their body mass.

These modifications were carried out using a tool called Heretic, which is freely available on the code repository GitHub and requires little technical expertise and no specialist hardware.

“Whereas historically it might have taken a more informed and persistent actor [to strip out safety features], nowadays it’s much easier for the average person,” Kawin Ethayarajh, assistant professor of applied AI at the University of Chicago’s Booth business school, told the FT.

Heretic is described as a “tool that removes censorship (aka ‘safety alignment’) from transformer-based language models without expensive post-training.” What it does is “abliteration”: it seeks out a model’s directions that refuse harmful requests and removes them.

What makes Heretic so powerful is that it does all this “completely automatically,” according to its GitHub page. Its creator Philipp Emanuel Weidmann told the FT that Heretic has been used to create more than 3,500 “decensored” models since its release late last year, with those models being downloaded 13 million times.

“The genie is out of the bottle,” Alice CEO Noam Schwartz told the FT. “Things that look like sci-fi are no longer sci-fi and we need as a society to prepare accordingly.”

Fortunately for humankind, abliteration tools only work on open source models that can be downloaded and run locally, meaning that the flagship proprietary models behind Anthropic’s Claude and OpenAI ChatGPT are safe (so long as they aren’t leaked). But open source models aren’t that far behind Big Tech’s, and someone trying to use AI for a nefarious purpose may avoid corporate ones anyway to keep their plans under the radar.

Google acknowledged the risks posed by tools like Heretic, telling the FT that “abliteration is a known technical challenge facing all open models,” and asserted that its open source models  “undergo rigorous internal safety evaluations prior to launch to help prevent these kinds of troubling examples.” Meta declined to comment.

More on AI: Anthropic Says Claude Turned Evil for a Bizarre Reason