





















Common Crawl's impact on research has grown substantially since its beginning. Our crawls have become a vital resource for researchers in various fields, from natural language processing to red teaming.
Our data has so far been cited in over 7,000 academic publications, highlighting its value to the research community. We recently published a blog post on this, and plan to further investigate the connections in this network.

We're excited to announce that Common Crawl’s statistics are now available on Hugging Face! The Common Crawl Statistics dataset includes metrics such as the number of URLs, domains, bytes, and content types crawled over specific periods. This dataset is important for users who need comprehensive, structured insights into the composition and trends within the web data collected by Common Crawl. For more details on the statistics see our recent blog post.
We now provide a refreshed crawl every month. To date, we've delivered over 100 total crawl archives. In August alone, we crawled more than 2.3 billion web pages, amounting to over 320 TiB of uncompressed content. The total size of our corpus now exceeds 8 PiB, with WARC data alone exceeding 7 PiB—a growth of 10.87% in the past year. In addition to this, we now also generate Web Graphs on a monthly basis (as opposed to once every three months), which has significantly improved the quality of our recent crawls by giving us fresher ranking data.

We're actively influencing and shaping policy discussions for a free and open Internet. Earlier this year we hosted a conference in New York titled "AI & the Right to Learn on an Open Internet", co-hosted with Professor Jeff Jarvis at the Craig Newmark Graduate School of Journalism at CUNY.
We have been conducting experiments to evaluate the prevalence of the emerging ML and AI opt-out protocols, and we have been engaging in discussions with our users and collaborators on the best way to put in place, apply, and standardize these protocols.
We're also taking part in a workshop hosted by the Internet Architecture Board in Washington DC in September. The workshop will explore practical opt-out mechanisms for AI data collection, focusing on how content creators can control the use of their online content in training LLMs and other AI systems.
Based on user feedback, we're pursuing the following initiatives:
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。