惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

H
Hackread – Cybersecurity News, Data Breaches, AI and More
U
Unit 42
Vercel News
Vercel News
Martin Fowler
Martin Fowler
云风的 BLOG
云风的 BLOG
爱范儿
爱范儿
MongoDB | Blog
MongoDB | Blog
J
Java Code Geeks
F
Fortinet All Blogs
MyScale Blog
MyScale Blog
C
Check Point Blog
N
Netflix TechBlog - Medium
Microsoft Azure Blog
Microsoft Azure Blog
aimingoo的专栏
aimingoo的专栏
博客园_首页
WordPress大学
WordPress大学
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
IT之家
IT之家
Last Week in AI
Last Week in AI
罗磊的独立博客
大猫的无限游戏
大猫的无限游戏
Jina AI
Jina AI
V
Visual Studio Blog
小众软件
小众软件

Hacker News - Newest: "AI"

AI can't read an investor deck AI as an attorney? Student uses ChatGPT, Gemini to sue UW over alleged racial discrimination Hacking MCP Servers in AI Systems – The Rug Pull: Tool Changes After Approval GitHub - MeepCastana/KubeezCut: Free Web based video editor Can AI judge journalism? A Thiel-backed startup says yes, even if it risks chilling whistleblowers Coming soon: 10 Things That Matter in AI Right Now DARPA built an AI to fact-check enemy weapons claims What explains heterogeneity in AI adoption? When AI Meets Muscle: Context-Aware Electrical Stimulation Promises a New Way to Guide Human Movements - Department of Computer Science AI Changed How We Build. It Did Not Change What Matters. Linux rules on using AI-generated code - Copilot is OK, but humans must take 'full responsibility for the… Meta spins up AI version of Mark Zuckerberg to engage with employees Code Mode: Let Your AI Write Programs, Not Just Call Tools | TanStack Blog GitHub - Delavalom/graft: Go framework for building AI agents. Type-safe tools, multi-provider (OpenAI, Anthropic, Gemini, Bedrock), zero vendor SDKs. India's TCS tops estimates, says new AI models did not dent services demand Gen Z's fading AI hype Strong feeling: we are in a folded AI reality GitHub - machinarii/total-recall-catalog: A reference catalog of latest knowledge retrieval, memory & RAG systems GitHub - mensfeld/code-on-incus: Give each AI agent its own isolated machine with root, Docker, and systemd. Active defense detects and stops threats automatically.. Quantization, LoRA, and the 8% Problem: Benchmarking Local LLMs for Production AI Iran war: We spoke to the man making Lego-style AI videos that experts say are powerful propaganda Powell, Bessent discussed Anthropic's Mythos AI cyber threat with major U.S. banks GitHub - immartian/bellamem: Persistent belief-graph memory for AI agents. Retrieves decisive context by importance — not recency, not RAG, not /compact. recursive-mode: The Repo-Native Operating System for AI Engineering After the attack on Sam Altman's home, will AI CEO's go on the offensive? The biggest advance in AI since the LLM Opus 4.6 vs GPT 5.4 One Prompt Unity World Generation Test “AI polls” are fake polls Client Challenge Can AI be a 'child of God'? Inside Anthropic's meeting with Christian leaders
Fighting the AI scraperbot scourge
By Jonathan CorbetFebruary 14, 2025 · 2026-05-27 · via Hacker News - Newest: "AI"
Please consider subscribing to LWN

Subscriptions are the lifeblood of LWN.net. If you appreciate this content and would like to see more of it, your subscription will help to ensure that LWN continues to thrive. Please visit this page to join up and keep LWN on the net.

There are many challenges involved with running a web site like LWN. Some of them, such as finding the courage to write for people who know more about the subject matter than we do, simply come with the territory we have chosen. But others show up as an unwelcome surprise; the ongoing task of fending off bots determined to scrape the entire Internet to (seemingly) feed into the insatiable meat grinder of AI training is certainly one of those. Readers have, at times, expressed curiosity about that fight and how we are handling it; read on for a description of a modern-day plague.

Training the models for the generative AI systems that, we are authoritatively informed, are going to transform our lives for the better requires vast amounts of data. The most prominent companies working in this area have made it clear that they feel an unalienable entitlement to whatever data they can get their virtual hands on. But that is just the companies that are being at least slightly public about what they are doing. With no specific examples to point to, I nonetheless feel quite certain that, for every company working in the spotlight, there are many others with model-building programs that they are telling nobody about. Strangely enough, these operations do not seem to talk to each other or share the data they pillage from sites across the net.

The LWN content-management system contains over 750,000 items (articles, comments, security alerts, etc) dating back to the adoption of the "new" site code in 2002. We still have, in our archives, everything we did in the over four years we operated prior to the change as well. In addition, the mailing-list archives contain many hundreds of thousands of emails. All told, if you are overcome by an irresistible urge to download everything on the site, you are going to have to generate a vast amount of traffic to obtain it all. If you somehow feel the need to do this download repeatedly, just in case something changed since yesterday, your traffic will be multiplied accordingly. Factor in some unknown number of others doing the same thing, and it can add up to an overwhelming amount of traffic.

LWN is not served by some massive set of machines just waiting to keep the scraperbots happy. The site is, we think, reasonably efficiently written, and is generally responsive. But when traffic spikes get large enough, the effects will be felt by our readers; that is when we start to get rather grumpier than usual. And it is not just us; this problem has been felt by maintainers of resources all across our community and beyond.

In discussions with others and through our own efforts, we have looked at a number of ways of dealing with this problem. Some of them are more effective than others.

For example, the first suggestion from many is to put the offending scrapers into robots.txt, telling them politely to go away. This approach offers little help, though. While the scraperbots will hungrily pull down any content on the site they can find, most of them religiously avoid ever looking at robots.txt. The people who run these systems are absolutely uninterested in our opinion about how they should be accessing our site. To make this point even more clear, most of these robots go out of their way to avoid identifying themselves as such; they try as hard as possible to look like just another reader with a web browser.

Throttling is another frequently suggested solution. The LWN site has implemented basic IP-based throttling for years; even in the pre-AI days, it would often happen that somebody tried to act on a desire to download the entire site, preferably in less than five minutes. There are also systems like commix that will attempt to exploit every command-injection vulnerability its developers can think of, at a rate of thousands per second. Throttling is necessary to deal with such actors but, for reasons that we will get into momentarily, throttling is relatively ineffective against the current crop of bots.

Others suggest tarpits, such as Nepenthes, that will lead AI bots into a twisty little maze of garbage pages, all alike. Solutions like this bring an additional risk of entrapping legitimate search-engine scrapers that (normally) follow the rules. While LWN has not tried such a solution, we believe that this, too, would be ineffective. Among other things, these bots do not seem to care whether they are getting garbage or not, and serving garbage to bots still consumes server resources. If we are going to burn kilowatts and warm the planet, we would like the effort to be serving a better goal than that.

But there is a deeper reason why both throttling and tarpits do not help: the scraperbots have been written with these defenses in mind. They spread their HTTP activity across a set of IP addresses so that none reach the throttling threshold. In some cases, those addresses are all clearly coming from the same subnet; a certain amount of peace has been obtained by treating the entire Internet as a set of class-C subnetworks and applying a throttling threshold to each. Some operators can be slowed to a reasonable pace in this way. (Interestingly, scrapers almost never use IPv6).

But, increasingly, the scraperbot traffic does not fit that pattern. Instead, traffic will come from literally millions of IP addresses, where no specific address is responsible for more than two or three hits over the course of a week. Watching the traffic on the site, one can easily see scraping efforts that are fetching a sorted list of URLs in an obvious sequence, but the same IP address will not appear twice in that sequence. The specific addresses involved come from all over the globe, with no evident pattern.

In other words, this scraping is being done by botnets, quite likely bought in underground markets and consisting of compromised machines. There really is not any other explanation that fits the observed patterns. Once upon a time, compromised systems were put to work mining cryptocurrency; now, it seems, there is more money to be had in repeatedly scraping the same web pages. When one of these botnets goes nuts, the result is indistinguishable from a distributed denial-of-service (DDOS) attack — it is a distributed denial-of-service attack. Should anybody be in doubt about the moral integrity of the people running these systems, a look at the techniques they use should make the situation abundantly clear.

That leads to the last suggestion that often is heard: use a commercial content-delivery network (CDN). These networks are working to add scraperbot protections to the DDOS protections they already have. It may come to that, but it is not a solution we favor. Exposing our traffic (and readers) to another middleman seems undesirable. Many of the techniques that they use to fend off scraperbots — such as requiring the user and/or browser to answer a JavaScript-based challenge — run counter to how we want the site to work.

So, for the time being, we are relying on a combination of throttling and some server-configuration work to clear out a couple of performance bottlenecks. Those efforts have had the effect of stabilizing the load and, for now, eliminating the site delays that we had been experiencing. None of this stops the activity in question, which is frustrating for multiple reasons, but it does prevent it from interfering with the legitimate operation of the site. It seems certain, though, that this situation will only get worse over time. Everybody wants their own special model, and governments show no interest in impeding them in any way. It is a net-wide problem, and it is increasingly unsustainable.

LWN was born in the era when the freedom to put a site onto the Internet was a joy to experience. That freedom has since been beaten back in many ways, but still exists for the most part. If, though, we reach a point where the only way to operate a site of any complexity is to hide it behind one of a tiny number of large CDN providers (each of which probably has AI initiatives of its own), the net will be a sad place indeed. The humans will have been driven off (admittedly, some may see that as a good thing) and all that will be left is AI systems incestuously scraping pages from each other.

Index entries for this article
SecurityWeb