惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Hugging Face - Blog
Hugging Face - Blog
腾讯CDC
阮一峰的网络日志
阮一峰的网络日志
博客园_首页
Last Week in AI
Last Week in AI
月光博客
月光博客
D
DataBreaches.Net
WordPress大学
WordPress大学
雷峰网
雷峰网
酷 壳 – CoolShell
酷 壳 – CoolShell
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
博客园 - 叶小钗
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
U
Unit 42
Recent Announcements
Recent Announcements
宝玉的分享
宝玉的分享
MyScale Blog
MyScale Blog
C
Check Point Blog
F
Fortinet All Blogs
B
Blog
小众软件
小众软件
Vercel News
Vercel News
罗磊的独立博客
有赞技术团队
有赞技术团队

Futurism

Sam Altman Now Trying to Gain Control of Electric Grid Top Chinese Court Issues Sweeping Legal Crackdown on AI Deepfakes Meta Releases Uber-Creepy AI Chatbot as Its Platforms Crumble Under Grotesque Child Abuse LG TVs Caught Secretly Recording Users and Scanning Their Homes For Other Devices, Even When Disconnected From the Internet People Are Telling Their Darkest Thoughts to AI Without Realizing They Can Easily Become Public Hackers Are Selling Stolen Scans of 153 Million US and Canadian Drivers Licenses, Which Very Likely Include Yours FBI Now Allowing History of Bestiality Among New Recruits The Transcripts of OpenAI Models Plotting Together to Commit an Actual Crime Is Pretty Chilling Flock Is Quietly Selling Powerful Drones That Scan License Plates From the Sky McDonald's Has Hundreds of Pages of Intel on Its Repeat Customers, and You Can Get a Copy of Yours Man Wearing Pervert Glasses Films Himself Harassing Famous Female Comedian in the Middle of TV Shoot Sensing He's in Deep Trouble, Flock Safety CEO Says It's All Been a Big Misunderstanding Hackers Created a Device That Can Take Over a Boeing 737 Jet's Autopilot Without Anyone Noticing Man Covers Car in Special Wrap That Breaks Flock Cameras' Electronic Brains Scammers Tremble as AI Comes for Their Jobs Why Aren't Any AI Companies Watching Their Frontier Models to Make Sure They Don't Go on Hacking Sprees? Jealously Watching OpenAI and Anthropic, Meta Suddenly Claims That Its AI Went on a Hacking Spree Too OpenAI's Escaped Models Were Allegedly Rampaging More Extensively Than Previously Reported If You AI-Generate Code, Hackers Just Found a Devious Method to Install Malware Directly on Your Computer This New Meta "Advertisement" Is Absolutely Brutal Suspicion Grows About OpenAI's Tale About Its Rogue Hacker AI OpenAI Says a Group of Its Models Broke Out of Secure Containment and Hacked a Prominent AI Site AI Browsers Can Basically Be Hypnotized Into Turning Against Their User and Carrying Out Devastating Hacks Meta’s AI Support Bot Is Giving Hackers Access to Other People’s Instagram Accounts Just by Asking Websites Are Spying on Your Solid State Drive The MyPillow Guy’s Entire Business is Being Held Hostage by Hackers Riot Games Denies Using Anti-Cheat Software That Bricks Hackers’ Computers The Trump Phone Appears to Have Already Leaked Its Customers’ Personal Information Through a Glaring Exploit College Kid Shuts Down High Speed Trains With a Laptop and a Radio Google Alarmed by Formidable AI-Powered Zero-Day Cyberattack
It's Laughably Easy to Poison Open-Weight AI Models, Rese...
Frank Landymore · 2026-07-19 · via Futurism

A glitched bitmap silhouette of a hooded hacker-type.

Igor Kyrlytsya via Shutterstock

Sign up to see the future, today

Can’t-miss innovations from the bleeding edge of science and tech

Good news, everyone: “open-weight” AI models that are available to anyone to download and run are comically easy to poison.

Katie Paxton-Fear, a cybersecurity researcher at Semgrep, demonstrated this in an attack that took less than an hour and cost less than $100 to carry out, The Register reports, successfully manipulating the AI’s behavior by feeding it malicious data.

By training the model on just ten examples of poisoned material, the model started churning out new code that’s exposed to remote code execution, a vulnerability that allows hackers to run code on a person’s machine. 

“I did a proper backdoor,” she triumphantly shared on social media.

Backdoors are a particularly dangerous type of attack that involves training an AI in a way that introduces hidden phrases into the underlying model. A hacker can use these to quietly trigger the model into carrying out a specific action, lying dormant and unseen until they’re called into action. Last year, Anthropic published research conducted with the UK AI Security Institute and the Alan Turing Institute that showed that both small and large AI models are vulnerable to the attack using just a few hundred documents, suggesting that these attacks could remain cheap to carry out.

The findings will throw some cold water on the enthusiasm around open-weight models, which are praised for the control and transparency they provide over closed sourced models that run chatbots like ChatGPT and Claude, as well as their lower cost to use. But while their parameters may be visible, open-weight models don’t reveal their training data or their code — meaning they can still be black boxes to security researchers.

“Even when model weights are public (‘open weight’), we have almost no ability to predict its behavior,” Paxton-Fear’s colleagues at Semrep wrote in a post last week. “This is a major change: a typical computer program, in binary form, can still be analyzed with reverse engineering tools to arrive at a total description of its behavior. With models, we have nowhere close to this capability.”

Moreover, this is all fairly uncharted waters in cybersecurity. LLMs are incredibly complex and still new, so it’s difficult to uncover sophisticated attacks. And AI models can be compromised in much more subtle ways than software.

“If a software dependency contains malicious code, we have mature practices for discovering it, tracking its provenance, and reducing its impact,” the Semgrup researchers argued. “AI models are different. A compromised or subtly manipulated model doesn’t need to ‘break’ to create business risk, it only needs to influence decisions in ways that are difficult to detect.”

“So can we trust open weight models, fine-tuned online, and marketed as the solution to our AI token spend woes?” Paxton-Fear asked in a thread sharing her findings. “Well, we probably need something better than benchmarks and ‘and don’t write any insecure code.'”

More on AI: AI Bubble Fears Are Starting to Spill Over