惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Jina AI
Jina AI
Recent Announcements
Recent Announcements
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
月光博客
月光博客
A
About on SuperTechFans
Vercel News
Vercel News
博客园 - 【当耐特】
爱范儿
爱范儿
Blog — PlanetScale
Blog — PlanetScale
阮一峰的网络日志
阮一峰的网络日志
V
V2EX
D
Docker
博客园 - 叶小钗
The Cloudflare Blog
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
H
Help Net Security
I
InfoQ
博客园 - 三生石上(FineUI控件)
博客园 - Franky
Microsoft Azure Blog
Microsoft Azure Blog
The GitHub Blog
The GitHub Blog
大猫的无限游戏
大猫的无限游戏
MongoDB | Blog
MongoDB | Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报

Forbes - Consumer Tech

This Unhackable Quantum Navigation System Is The Size Of A Loaf Of Bread Apple At 50 — A Leadership Shift And An AR Future We Are Under-Investing In Robotics ... 90% Of Humanoid Robots Are Made In China Ditch The Apple White: Beats Expands Colorful Cable Line-Up With New 10-Foot Option Satechi’s New ChargeView 140W Desktop GaN Charger With Real-Time Display The Hasselblad In Your Pocket: Oppo’s Find X9 Ultra Challenges The Galaxy S26 Ultra There's No Such Thing As Brain Honey How AI Agents Could Rebuild Fashion’s Visual Production Layer Sennheiser’s New Closed-Back Headphones Are Made For The Studio QClaw Goes Global. The Agent Built Itself In 5 Days Apple’s Tim Cook Exit Hides A $4 Trillion Agentic AI Power Move EZQuest Reveals A New Line Of Pro Series USB-C Hubs For MacBook Neo Samsung Galaxy Z TriFold 2 Already In The Works, Report Claims Apple Revealed New Siri Release Date For iPhone, Latest Report Claims How Arcani’s HARK Is Designed For Modern Battlefield Acoustics The Newest Trend In Tech Embraces Femininity And Fun Samsung’s 75R95H Ushers In A New World Of LCD TVs New Apple iPhone Fold Design Pushes Smartphone Rivals To Go Wider And Taller iPhone 18 Pro Report: Four New Colors Leak As Apple Cancels Popular Shade Nothing’s Design-Led Strategy: Carl Pei Reveals The Tech Brand’s Philosophy iOS 26.5 Release Date: When To Expect Your iPhone Messaging Upgrade Google Pixel And Highsnobiety Build A Talent Pipeline For Fashion Android Circuit: Samsung Raises Galaxy Prices, Oppo Pad Mini Teased, Microsoft Closing Outlook App Apple Loop: iPhone Fold Launch Dates, iPad Air Upgrade, iPhone 18 Pro Specs Comcast $117.5 Million Breach Settlement — Are You Eligible? Amazfit Cheetah 2 Pro Takes Aim At The Garmin Audience Disney’s Launches ‘Infinity Vision’ Certification For Premium Theaters SoundPeats Reveals New Air6 HS Semi-Open Wireless Earbuds Amazon’s $11.57 Billion Leap Into Space: A Challenge To Starlink Meta Quest 3 Hit With $100 Price Increase
Why Major News Sites Are Blocking The Internet Archive’s ...
Anisha Sircar · 2026-04-14 · via Forbes - Consumer Tech
Alcatel Minitel communication terminal, France, 1983.

UNITED KINGDOM - JULY 22: The Minitel system represents an independent French internet. Users can access up to 22,000 databases and services, for which they have to pay an access charge. (Photo by SSPL/Getty Images)

SSPL via Getty Images

For nearly three decades, the Internet Archive’s Wayback Machine has served as a go-to for anyone looking to access its vast treasure trove of archived internet pages.

Its mission of crawling and preserving the public web has made it an indispensable resource for journalists, historians, researchers, courts and beyond. As of October 2025, the Wayback Machine contained more than one trillion archived web pages.

Now, that mission is facing a threat that could make it harder to access digital history.

According to an analysis by the AI-detection startup Originality AI, 23 major news sites currently block ia_archiverbot, the web crawler the Internet Archive commonly uses for the Wayback project. However, that understates the full scope: in total, 241 news sites from nine countries explicitly disallow at least one of the four Internet Archive crawling bots.

Most of those sites — 87% — are owned by USA Today Co., the largest newspaper conglomerate in the United States, formerly known as Gannett. The company operates more than 200 media outlets, making its decision to block the Archive particularly jarring. The New York Times has gone further, actively “hard blocking” the Internet Archive’s crawlers — measures that go beyond the web's standard robots.txt conventions. At the end of 2025, the Times also added one of those crawlers — archive.org_bot — to its robots.txt file. Reddit, meanwhile, announced in August 2025 that it would block the Internet Archive.

The Guardian takes a more surgical approach. The publisher decided to limit the Internet Archive’s access to published articles, minimizing the chance that AI companies might scrape its content via the nonprofit’s repository. Its regional homepages and topic pages continue to appear in the Wayback Machine — just not its journalism.

Why the blockade?

The publishers aren’t targeting the Internet Archive specifically — the culprit, they say, is bots and artificial intelligence.

“This effort is not about specifically blocking the Internet Archive,” USA Today Co. spokesperson Lark-Marie Anton told Wired, describing it instead as part of the company’s broader effort to block all scraping bots.

Robert Hahn, the Guardian’s director of business affairs and licensing, said the outlet has been in conversation with the Archive over concerns about potential misuse of crawled content by AI companies.

Some publishers have been direct.

“The issue is that Times content on the Internet Archive is being used by AI companies in violation of copyright law to directly compete with us,” New York Times spokesperson Graham James told Wired, declining to clarify whether this referred to documented violations or a hypothetical concern. Reddit cited the same rationale — it didn’t want the Wayback Machine to become a backdoor for AI companies to access content Reddit is now licensing.

The anxiety may not be entirely unfounded, particularly given evidence that the Wayback Machine has been used to train large language models.

An analysis of Google’s C4 dataset by the Washington Post in 2023 showed that the Internet Archive was among the websites in the training data used to build Google’s T5 model and Meta’s Llama models.

The problem is escalating, too: in May 2023, the Internet Archive went temporarily offline after an AI company caused a server overload, sending tens of thousands of requests per second to extract text data from the nonprofit’s public domain archives.

There is also a legal angle. Publishers including the Times have filed copyright lawsuits against AI companies like OpenAI and Perplexity, and the ongoing conversation about AI training data is playing out across more than 100 active copyright cases in the United States.

Pushback

Individual reporters are pushing back, and they’re organized.

Fight for the Future, the Electronic Frontier Foundation and Public Knowledge coordinated a letter thanking the Internet Archive for its preservation of news and history at a moment when major outlets are reconsidering their relationship with the Archive.

The coalition collected more than 100 signatures from journalists and presented the letter to the Internet Archive. Signatories range from television anchor Rachel Maddow to independent reporters like Kat Tenbarge and Taylor Lorenz.

“In previous generations, journalists would turn to the physical archives of a local newspaper or of a local public library to access historical reporting and follow the threads of the present back into history,” the letter reads.

“With many newspapers closed, and no clear path for local public libraries to preserve digital-only reporting, the work of safeguarding journalism’s record increasingly falls to the Internet Archive.”

Access to History

There is an interesting irony undergirding this developing story.

A recent USA Today investigation illustrated the tool’s indispensable role: journalists used the Wayback Machine to analyze how U.S. Immigration and Customs Enforcement altered detention statistics and changed policies — reporting that depended heavily on archived web data. Meanwhile, USA Today Co. blocks the Archive’s crawler from preserving its own work.

"They're able to pull together their story research because the Wayback Machine exists,” Wayback Machine director Mark Graham told Wired. “At the same time, they're blocking access."

That’s not entirely new: In 2016, the Archive documented the New York Times revising a Bernie Sanders article — the kind of editorial accountability that becomes complicated when outlets control their own historical records.

Graham sees the publishers’ restrictions as part of a trend toward closing off what was once an open web.

“It’s not just making it harder for the Wayback Machine to preserve this material and make it available for future generations,” he reportedly said. “But it’s having a profound impact on anyone’s ability to access quality journalistic material.”

Graham described the Archive as “collateral damage” in a copyright war it didn't start.

No comparable public alternative to the Wayback Machine exists so far — and if major outlets continue restricting access, the ability to verify claims, track editorial changes and research historical context will deplete, creating information asymmetries where only large organizations can effectively control records of what they once said.

The Internet Archive said it remains in conversation with blocked outlets. However, resolution looks uncertain as AI copyright battles intensify across the industry — and the historical record continues to shrink.