























UNITED KINGDOM - JULY 22: The Minitel system represents an independent French internet. Users can access up to 22,000 databases and services, for which they have to pay an access charge. (Photo by SSPL/Getty Images)
SSPL via Getty Images
For nearly three decades, the Internet Archive’s Wayback Machine has served as a go-to for anyone looking to access its vast treasure trove of archived internet pages.
Its mission of crawling and preserving the public web has made it an indispensable resource for journalists, historians, researchers, courts and beyond. As of October 2025, the Wayback Machine contained more than one trillion archived web pages.
Now, that mission is facing a threat that could make it harder to access digital history.
According to an analysis by the AI-detection startup Originality AI, 23 major news sites currently block ia_archiverbot, the web crawler the Internet Archive commonly uses for the Wayback project. However, that understates the full scope: in total, 241 news sites from nine countries explicitly disallow at least one of the four Internet Archive crawling bots.
Most of those sites — 87% — are owned by USA Today Co., the largest newspaper conglomerate in the United States, formerly known as Gannett. The company operates more than 200 media outlets, making its decision to block the Archive particularly jarring. The New York Times has gone further, actively “hard blocking” the Internet Archive’s crawlers — measures that go beyond the web's standard robots.txt conventions. At the end of 2025, the Times also added one of those crawlers — archive.org_bot — to its robots.txt file. Reddit, meanwhile, announced in August 2025 that it would block the Internet Archive.
The Guardian takes a more surgical approach. The publisher decided to limit the Internet Archive’s access to published articles, minimizing the chance that AI companies might scrape its content via the nonprofit’s repository. Its regional homepages and topic pages continue to appear in the Wayback Machine — just not its journalism.
The publishers aren’t targeting the Internet Archive specifically — the culprit, they say, is bots and artificial intelligence.
“This effort is not about specifically blocking the Internet Archive,” USA Today Co. spokesperson Lark-Marie Anton told Wired, describing it instead as part of the company’s broader effort to block all scraping bots.
Robert Hahn, the Guardian’s director of business affairs and licensing, said the outlet has been in conversation with the Archive over concerns about potential misuse of crawled content by AI companies.
Some publishers have been direct.
“The issue is that Times content on the Internet Archive is being used by AI companies in violation of copyright law to directly compete with us,” New York Times spokesperson Graham James told Wired, declining to clarify whether this referred to documented violations or a hypothetical concern. Reddit cited the same rationale — it didn’t want the Wayback Machine to become a backdoor for AI companies to access content Reddit is now licensing.
The anxiety may not be entirely unfounded, particularly given evidence that the Wayback Machine has been used to train large language models.
An analysis of Google’s C4 dataset by the Washington Post in 2023 showed that the Internet Archive was among the websites in the training data used to build Google’s T5 model and Meta’s Llama models.
The problem is escalating, too: in May 2023, the Internet Archive went temporarily offline after an AI company caused a server overload, sending tens of thousands of requests per second to extract text data from the nonprofit’s public domain archives.
There is also a legal angle. Publishers including the Times have filed copyright lawsuits against AI companies like OpenAI and Perplexity, and the ongoing conversation about AI training data is playing out across more than 100 active copyright cases in the United States.
Individual reporters are pushing back, and they’re organized.
Fight for the Future, the Electronic Frontier Foundation and Public Knowledge coordinated a letter thanking the Internet Archive for its preservation of news and history at a moment when major outlets are reconsidering their relationship with the Archive.
The coalition collected more than 100 signatures from journalists and presented the letter to the Internet Archive. Signatories range from television anchor Rachel Maddow to independent reporters like Kat Tenbarge and Taylor Lorenz.
“In previous generations, journalists would turn to the physical archives of a local newspaper or of a local public library to access historical reporting and follow the threads of the present back into history,” the letter reads.
“With many newspapers closed, and no clear path for local public libraries to preserve digital-only reporting, the work of safeguarding journalism’s record increasingly falls to the Internet Archive.”
There is an interesting irony undergirding this developing story.
A recent USA Today investigation illustrated the tool’s indispensable role: journalists used the Wayback Machine to analyze how U.S. Immigration and Customs Enforcement altered detention statistics and changed policies — reporting that depended heavily on archived web data. Meanwhile, USA Today Co. blocks the Archive’s crawler from preserving its own work.
"They're able to pull together their story research because the Wayback Machine exists,” Wayback Machine director Mark Graham told Wired. “At the same time, they're blocking access."
That’s not entirely new: In 2016, the Archive documented the New York Times revising a Bernie Sanders article — the kind of editorial accountability that becomes complicated when outlets control their own historical records.
Graham sees the publishers’ restrictions as part of a trend toward closing off what was once an open web.
“It’s not just making it harder for the Wayback Machine to preserve this material and make it available for future generations,” he reportedly said. “But it’s having a profound impact on anyone’s ability to access quality journalistic material.”
Graham described the Archive as “collateral damage” in a copyright war it didn't start.
No comparable public alternative to the Wayback Machine exists so far — and if major outlets continue restricting access, the ability to verify claims, track editorial changes and research historical context will deplete, creating information asymmetries where only large organizations can effectively control records of what they once said.
The Internet Archive said it remains in conversation with blocked outlets. However, resolution looks uncertain as AI copyright battles intensify across the industry — and the historical record continues to shrink.
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。