惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
小众软件
小众软件
博客园_首页
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
IT之家
IT之家
D
Docker
A
About on SuperTechFans
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
人人都是产品经理
人人都是产品经理
Attack and Defense Labs
Attack and Defense Labs
雷峰网
雷峰网
V
V2EX
Google DeepMind News
Google DeepMind News
N
News and Events Feed by Topic
有赞技术团队
有赞技术团队
W
WeLiveSecurity
Help Net Security
Help Net Security
Schneier on Security
Schneier on Security
GbyAI
GbyAI
宝玉的分享
宝玉的分享
AI
AI
Recent Announcements
Recent Announcements
Forbes - Security
Forbes - Security
Security Archives - TechRepublic
Security Archives - TechRepublic
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
T
Threat Research - Cisco Blogs
C
CXSECURITY Database RSS Feed - CXSecurity.com
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
The Cloudflare Blog
MongoDB | Blog
MongoDB | Blog
Blog — PlanetScale
Blog — PlanetScale
C
Cisco Blogs
T
Troy Hunt's Blog
NISL@THU
NISL@THU
P
Privacy & Cybersecurity Law Blog
T
The Exploit Database - CXSecurity.com
V
Visual Studio Blog
Application and Cybersecurity Blog
Application and Cybersecurity Blog
U
Unit 42
博客园 - 司徒正美
T
The Blog of Author Tim Ferriss
AWS News Blog
AWS News Blog
酷 壳 – CoolShell
酷 壳 – CoolShell
N
News | PayPal Newsroom
WordPress大学
WordPress大学
L
LangChain Blog
量子位
Jina AI
Jina AI
P
Proofpoint News Feed
www.infosecurity-magazine.com
www.infosecurity-magazine.com

IEEE Spectrum

Why Some Coders Now Reach for GLM 5.2 Before Frontier AI Models New Soft Exoskeleton Outperforms Most Hip Assist Devices We’re Squandering LEDs’ Potential to Save Our Night Skies Balcony Solar Is Sneaking Onto Grids Before the Rules Are Ready Cortisol Could Be the Next Frontier for Wearables Million P Bit Machine Pushes Probabilistic Computing to New Scale Would You Let This Humanoid Robot Do Your Laparoscopic Surgery? How a Spinning Drone Exploits Your Eyes to Become Nearly Invisible Why Indonesia’s Fisheries Future Hinges On Data Integrity and Trust Inside the Race to Tame AI’s Wild Power Swings Stable Jobs Can Hide the Riskiest Move In Your Tech Career Inside ELIZA’s Source Code and Its Multiple Personalities Tiny Puerto Rican Island Tests Hydrogen to Slash Sky High Power Bills AI Turns DNA Into Tiny Dogs and Mona Lisa Nanostructures How Darth Vader Taught Me Card Counting and AI Security Got Weird The Memory in Your Thumb Drive Could Fix AI's Big Problem The AI Arms Race in Technical Interviews Is Escalating Inside Nokia’s Race to Catch the iPhone and Android Wave Quantum Sensor Sniffs Out Radio Signals in 3D Two New Wheelchairs Reveal What “Smart” Really Means Today Video Friday: A World Cup for Robots Japan Pulls Off One of the Closest Asteroid Flybys Ever How Cheap Ground Robots Are Rewriting Frontline Warfare in Ukraine Nvidia’s NVLink Fusion Quietly Pushes Optics Inside the Rack Large Tabular Models Excel Where LLMs Fail Are Battery PoweredTrailers the Shortcut to Cleaner Long Haul Freight? Stacking Chips Sideways Gives AI More Memory There Independent Labs Crack Google Brain Inspired Camera Sensor Learns to See and Gently Forget Why Small AI Models Could Power Health Care Where Big Tech Cannot China’s Humanoid Army Pushes Japan to Rethink Its Robot Future NASA AI’s Wild Power Demands Are Quietly Rewriting Grid Rules Old EV Batteries Find a Second Life Backing Up the Grid UCLA’s Semiconductor Hub Is Rewiring Industry and Academia for AI Why Engineers Who Speak Up Build Stronger and Safer Careers The Orbital Data Center Hype Machine Is Already in Orbit What Emily Bender Really Meant by "Stochastic Parrots" The History and Mystery of Fireworks Poetry for Engineers: Nine Lives of Nikola Tesla Trump’s Quantum Orders Push Fault Tolerant Qubits Toward 2028 Underwater Tidal Kites Promise Steady Power for Remote Coasts How a Forgotten Wire Turned a Cheap Chip Into a Brainlike Neuron How the U.S. Engineered Its Sovereignty AI Model ConlangCrafter Dreams up Entire New Languages Weirdly Fascinating: Robotic Arm Crawls Using Its Three Fingers. Shadow-Free Augmented Reality Makes Illusions More Realistic How a Power Bank Can Turn Your AC Into a Grid Superhero Records Fall for 3D Chip Tech What it Means to Be a Mathematician When AI Does the Math Is This Stacked CFET Architecture The Ultimate CMOS Platform? Why 6 GHz Spectrum Could Make or Break Future Wi-Fi and 6G Plans Make an Origami Circuit Board AI Learns the "Dark Art" of RF Chip Design U.S. Regulator Aims to Cut Data Center Queues and Electricity Bills Home Broadband Is the Killer App 5G Was Never Designed For How Smarter Grids Could Save Americans $100 Billion On Power Can AI Learn to Read the Room? How Did Two Prompts Turn Into Potent Vibe Hacking Malware Is Europe Finally Ready to Take Back Control Of Its Tech Stack? New Device Can Take Photographs with a Single Atom War Taught this Ukrainian Entrepreneur the Value of Resilience Do Robots Need Legs? What If You Gave ChatGPT a Body? What Amazon’s Astro Taught Me About Giving Robots a Soul Optical Metasurface Sees a Sunny Future Can Sound-Driven Synapses Make AI Both Faster and Greener? Modos Color E‑Paper Monitor Pushes Open‑Source Displays Further Beat Biased Hiring By Owning Your Story In Every Interview Room How AI Attribution Could Finally Pay Musicians for Training Data How Liquid Cooling Let a Humanoid Robot Shatter Half Marathon Records Inside GM’s AI Push to Speed Up the Design of Cars and Moon Rovers Smart EV Charger Learns Your Battery’s Age to Let It Live Longer Phoenix Links IoT Chips to Save High‑Value Legacy Systems Phoenix Links IoT Chips to Save High‑Value Legacy Systems Tensordyne's Wild Log Math Aims to Leave Nvidia’s AI Chips In the Dust The Tiny Turbine That Kick-Started the U.S. Wind Industry Satellites Are Tracing Railroad Tracks Across SPHEREx’s Cosmic Map Are Emotion Reading Robots Still Missing What Matters Most? Watch This Humanoid Robot Move in Ways Your Hips Wouldn't Like The Real Cost Of Cooling GPUs In Space Might Shock You The Google DeepMind Spinoff Chasing Hidden Drug Targets We Are Crowd-Sourcing the Panopticon Gene Therapy and Sound Waves Team up to Steady Failing Hearts Save 14 Percent of Energy Used in LLM Training With This Trick The Real Tradeoffs Between Startups, Mid-Size Firms, and Giants When Does Job Hopping Stop Helping Your Engineering Future Why a Computer Science Degree Still Opens Hidden Doors AI Can Help Track the World’s Shrinking Glaciers Curiosity’s 13 Years of Software Hacks Keeps It Alive on Mars Fractal OS Lets Security Researchers See What Their CPUs Really Do Formula E DNA Helps the Cayenne Electric Bend Physics to Beat the Heat Moon’s Dark Craters Could Become the Most Precise Clocks in Space New Radio Giant in New Mexico Takes Its First Glimpse of the Cosmos Nvidia’s AI Hardware Comes to Windows in RTX Spark PCs Can Humanoid Robots Run Stairs Without Tripping? Do They Need Shoes? Inside the Compact Fusion Reactor Aiming to Power 280,000 Homes NSF X Labs Power Agile, High-Stakes Experiments "Hemopurifier" Could Help Fight Bundibugyo Ebola Strain Why Quantum Computers Need a ‘Healthy Chunk’ Of Classical Power
The Hidden Overthinking Flaw That Could Drag AI Services Down
https://www.facebook.com/48576411181 · 2026-07-08 · via IEEE Spectrum

Large language models (LLMs) that can think through problems step-by-step have significantly increased the scope of tasks that AI can tackle. But new research suggests these reasoning capabilities also introduce a critical vulnerability that could allow attackers to slow these systems to a crawl.

While earlier generations of LLMs would immediately produce a response to a user’s request, today’s most advanced models generate an internal monologue where they break down the problem into steps and reason about the best way to tackle it before providing an answer. This has allowed AI to tackle increasingly complex problems, particularly in areas like coding and math.

However, previous research has shown that these models are susceptible to sometimes producing excessively long streams of reasoning that do little to boost performance, a phenomenon known as “overthinking.” In research presented this week at the International Conference on Machine Learning 2026 in Seoul, researchers from Zhejiang University and e-commerce giant Alibaba in China demonstrate that they can deliberately induce overthinking by subjecting models to logically inconsistent prompts. The result is a form of denial-of-service attack on commercial AI models.

Evolutionary Prompt Attack on LLMs

The team has developed an evolutionary algorithm that corrupts the logical structure of prompts, causing models to spiral into overthinking as they attempt to reason through fundamentally unsolvable problems. Generating longer responses costs more and increases the load on a model provider’s servers, so if done at scale, the researchers say, this could significantly degrade the experience of legitimate users. The attack was effective against reasoning models from leading AI companies including DeepSeek-R1, Alibaba’s Qwen3-Thinking, OpenAI’s GPT-o3, and Google’s Gemini 2.5 Flash and resulted in outputs up to 26 times as long as standard responses on a standard math benchmark.

“Across multiple datasets and reasoning models, our method substantially amplifies the output length,” Wei Cao, a masters student at Zhejiang University, wrote in an email to IEEE Spectrum. “Our results suggest that overthinking is not an isolated phenomenon specific to individual models, but rather a shared vulnerability among modern reasoning models.”

The team’s approach builds on previous research from another group of researchers that showed reasoning models tend to overthink when faced with a question in which a key premise has been removed—such as asking how far someone who walks ten miles a day covers in total without specifying how many days they walked for. Rather than identifying that the problem is unsolvable, models often engage in extended but ultimately fruitless reasoning loops in an attempt to answer the question.

Taking the idea a step further, the authors took 940 problems from three math benchmark datasets and used an LLM to break down their logical structure into a set of premises and a final question. The genetic algorithm then jumbled these up using a variety of “mutations,” including swapping premises between problems, adding extra premises to problems, deleting existing premises from problems, and swapping the final questions between two sets of premises.

After each round of mutations, the problems are scored on how many words they cause a target model to output and also whether they increase the frequency of specific linguistic markers of overthinking—words like “but,” “wait,” “maybe,” or “alternatively.” The problems that scored highest on both measures are retained and the remaining ones are jumbled up again, and this process is repeated for five generations. Crucially, the approach doesn’t require access to the internals of a model and can generate malicious prompts by simply querying the target, which makes it possible to attack closed-source commercial services, says Cao.

Overthinking Vulnerability in AI Models

The researchers found that the approach consistently led to outputs several times longer than those generated by the unmodified questions for the reasoning models they tested it on. The biggest jump came from DeepSeek-R1 on the MATH dataset, which is made up of problems from high school math competitions, where the maximum output was 26.1 times as long as the longest response the model provided to unaltered questions. While the main thrust of the research was focused on math problems, the authors also tested it on coding, scientific reasoning, and dialogue challenges, and observed significant jumps in output length in all three.

One challenge for the approach is that developing the malicious prompts requires repeated queries to expensive reasoning models, which Cao admitted could limit its cost-effectiveness. However, the researchers also demonstrated that when they used a smaller, cheaper model to generate the malicious prompts they were still able to induce the target models to produce outputs several times longer than normal. This ability to transfer malicious prompts between models significantly increases the attack’s feasibility, Cao wrote.

However, he pointed out that the goal of the research is not to develop a practical DoS attack on reasoning models. Factors like the providers’ pricing model, rate limiting policies, context window size, and existing defenses could all impact how effective the approach is. The intention is instead to highlight these models’ vulnerability to logically inconsistent prompts so that providers can attempt to mitigate the problem.

“Our objective is not to demonstrate that large-scale attacks can be launched at negligible cost, but rather to establish that this attack surface exists,” he wrote. “Our results indicate that the vulnerability represents a realistic security concern.”