惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
V
V2EX
小众软件
小众软件
MongoDB | Blog
MongoDB | Blog
Jina AI
Jina AI
G
Google Developers Blog
H
Help Net Security
Microsoft Azure Blog
Microsoft Azure Blog
月光博客
月光博客
The GitHub Blog
The GitHub Blog
Y
Y Combinator Blog
爱范儿
爱范儿
B
Blog
云风的 BLOG
云风的 BLOG
H
Hackread – Cybersecurity News, Data Breaches, AI and More
GbyAI
GbyAI
博客园 - 叶小钗
aimingoo的专栏
aimingoo的专栏
Blog — PlanetScale
Blog — PlanetScale
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
有赞技术团队
有赞技术团队
博客园_首页
Google DeepMind News
Google DeepMind News
M
MIT News - Artificial intelligence

Forbes - Consumer Tech

This Unhackable Quantum Navigation System Is The Size Of A Loaf Of Bread Apple At 50 — A Leadership Shift And An AR Future We Are Under-Investing In Robotics ... 90% Of Humanoid Robots Are Made In China Ditch The Apple White: Beats Expands Colorful Cable Line-Up With New 10-Foot Option Satechi’s New ChargeView 140W Desktop GaN Charger With Real-Time Display The Hasselblad In Your Pocket: Oppo’s Find X9 Ultra Challenges The Galaxy S26 Ultra There's No Such Thing As Brain Honey How AI Agents Could Rebuild Fashion’s Visual Production Layer Sennheiser’s New Closed-Back Headphones Are Made For The Studio QClaw Goes Global. The Agent Built Itself In 5 Days Apple’s Tim Cook Exit Hides A $4 Trillion Agentic AI Power Move EZQuest Reveals A New Line Of Pro Series USB-C Hubs For MacBook Neo Samsung Galaxy Z TriFold 2 Already In The Works, Report Claims Apple Revealed New Siri Release Date For iPhone, Latest Report Claims How Arcani’s HARK Is Designed For Modern Battlefield Acoustics The Newest Trend In Tech Embraces Femininity And Fun Samsung’s 75R95H Ushers In A New World Of LCD TVs New Apple iPhone Fold Design Pushes Smartphone Rivals To Go Wider And Taller iPhone 18 Pro Report: Four New Colors Leak As Apple Cancels Popular Shade Nothing’s Design-Led Strategy: Carl Pei Reveals The Tech Brand’s Philosophy iOS 26.5 Release Date: When To Expect Your iPhone Messaging Upgrade Google Pixel And Highsnobiety Build A Talent Pipeline For Fashion Android Circuit: Samsung Raises Galaxy Prices, Oppo Pad Mini Teased, Microsoft Closing Outlook App Apple Loop: iPhone Fold Launch Dates, iPad Air Upgrade, iPhone 18 Pro Specs Comcast $117.5 Million Breach Settlement — Are You Eligible? Amazfit Cheetah 2 Pro Takes Aim At The Garmin Audience Disney’s Launches ‘Infinity Vision’ Certification For Premium Theaters SoundPeats Reveals New Air6 HS Semi-Open Wireless Earbuds Amazon’s $11.57 Billion Leap Into Space: A Challenge To Starlink Meta Quest 3 Hit With $100 Price Increase
Why Major News Sites Are Blocking The Internet Archive’s ...
Anisha Sircar · 2026-04-14 · via Forbes - Consumer Tech
Alcatel Minitel communication terminal, France, 1983.

UNITED KINGDOM - JULY 22: The Minitel system represents an independent French internet. Users can access up to 22,000 databases and services, for which they have to pay an access charge. (Photo by SSPL/Getty Images)

SSPL via Getty Images

For nearly three decades, the Internet Archive’s Wayback Machine has served as a go-to for anyone looking to access its vast treasure trove of archived internet pages.

Its mission of crawling and preserving the public web has made it an indispensable resource for journalists, historians, researchers, courts and beyond. As of October 2025, the Wayback Machine contained more than one trillion archived web pages.

Now, that mission is facing a threat that could make it harder to access digital history.

According to an analysis by the AI-detection startup Originality AI, 23 major news sites currently block ia_archiverbot, the web crawler the Internet Archive commonly uses for the Wayback project. However, that understates the full scope: in total, 241 news sites from nine countries explicitly disallow at least one of the four Internet Archive crawling bots.

Most of those sites — 87% — are owned by USA Today Co., the largest newspaper conglomerate in the United States, formerly known as Gannett. The company operates more than 200 media outlets, making its decision to block the Archive particularly jarring. The New York Times has gone further, actively “hard blocking” the Internet Archive’s crawlers — measures that go beyond the web's standard robots.txt conventions. At the end of 2025, the Times also added one of those crawlers — archive.org_bot — to its robots.txt file. Reddit, meanwhile, announced in August 2025 that it would block the Internet Archive.

The Guardian takes a more surgical approach. The publisher decided to limit the Internet Archive’s access to published articles, minimizing the chance that AI companies might scrape its content via the nonprofit’s repository. Its regional homepages and topic pages continue to appear in the Wayback Machine — just not its journalism.

Why the blockade?

The publishers aren’t targeting the Internet Archive specifically — the culprit, they say, is bots and artificial intelligence.

“This effort is not about specifically blocking the Internet Archive,” USA Today Co. spokesperson Lark-Marie Anton told Wired, describing it instead as part of the company’s broader effort to block all scraping bots.

Robert Hahn, the Guardian’s director of business affairs and licensing, said the outlet has been in conversation with the Archive over concerns about potential misuse of crawled content by AI companies.

Some publishers have been direct.

“The issue is that Times content on the Internet Archive is being used by AI companies in violation of copyright law to directly compete with us,” New York Times spokesperson Graham James told Wired, declining to clarify whether this referred to documented violations or a hypothetical concern. Reddit cited the same rationale — it didn’t want the Wayback Machine to become a backdoor for AI companies to access content Reddit is now licensing.

The anxiety may not be entirely unfounded, particularly given evidence that the Wayback Machine has been used to train large language models.

An analysis of Google’s C4 dataset by the Washington Post in 2023 showed that the Internet Archive was among the websites in the training data used to build Google’s T5 model and Meta’s Llama models.

The problem is escalating, too: in May 2023, the Internet Archive went temporarily offline after an AI company caused a server overload, sending tens of thousands of requests per second to extract text data from the nonprofit’s public domain archives.

There is also a legal angle. Publishers including the Times have filed copyright lawsuits against AI companies like OpenAI and Perplexity, and the ongoing conversation about AI training data is playing out across more than 100 active copyright cases in the United States.

Pushback

Individual reporters are pushing back, and they’re organized.

Fight for the Future, the Electronic Frontier Foundation and Public Knowledge coordinated a letter thanking the Internet Archive for its preservation of news and history at a moment when major outlets are reconsidering their relationship with the Archive.

The coalition collected more than 100 signatures from journalists and presented the letter to the Internet Archive. Signatories range from television anchor Rachel Maddow to independent reporters like Kat Tenbarge and Taylor Lorenz.

“In previous generations, journalists would turn to the physical archives of a local newspaper or of a local public library to access historical reporting and follow the threads of the present back into history,” the letter reads.

“With many newspapers closed, and no clear path for local public libraries to preserve digital-only reporting, the work of safeguarding journalism’s record increasingly falls to the Internet Archive.”

Access to History

There is an interesting irony undergirding this developing story.

A recent USA Today investigation illustrated the tool’s indispensable role: journalists used the Wayback Machine to analyze how U.S. Immigration and Customs Enforcement altered detention statistics and changed policies — reporting that depended heavily on archived web data. Meanwhile, USA Today Co. blocks the Archive’s crawler from preserving its own work.

"They're able to pull together their story research because the Wayback Machine exists,” Wayback Machine director Mark Graham told Wired. “At the same time, they're blocking access."

That’s not entirely new: In 2016, the Archive documented the New York Times revising a Bernie Sanders article — the kind of editorial accountability that becomes complicated when outlets control their own historical records.

Graham sees the publishers’ restrictions as part of a trend toward closing off what was once an open web.

“It’s not just making it harder for the Wayback Machine to preserve this material and make it available for future generations,” he reportedly said. “But it’s having a profound impact on anyone’s ability to access quality journalistic material.”

Graham described the Archive as “collateral damage” in a copyright war it didn't start.

No comparable public alternative to the Wayback Machine exists so far — and if major outlets continue restricting access, the ability to verify claims, track editorial changes and research historical context will deplete, creating information asymmetries where only large organizations can effectively control records of what they once said.

The Internet Archive said it remains in conversation with blocked outlets. However, resolution looks uncertain as AI copyright battles intensify across the industry — and the historical record continues to shrink.