惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

aimingoo的专栏
aimingoo的专栏
宝玉的分享
宝玉的分享
博客园 - 【当耐特】
博客园 - 司徒正美
L
LangChain Blog
有赞技术团队
有赞技术团队
大猫的无限游戏
大猫的无限游戏
Stack Overflow Blog
Stack Overflow Blog
Engineering at Meta
Engineering at Meta
U
Unit 42
Microsoft Azure Blog
Microsoft Azure Blog
I
InfoQ
博客园 - 叶小钗
H
Hackread – Cybersecurity News, Data Breaches, AI and More
J
Java Code Geeks
月光博客
月光博客
量子位
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
博客园_首页
Last Week in AI
Last Week in AI
人人都是产品经理
人人都是产品经理
Google DeepMind News
Google DeepMind News
云风的 BLOG
云风的 BLOG
D
DataBreaches.Net

Common Crawl

Common Crawl - Blog - Introducing the AI Visibility Audit Common Crawl - Blog - Host- and Domain-Level Web Graphs March, April, and May 2026 Common Crawl - Blog - May 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Crawl Archive Now Available in a Hugging Face Storage Bucket Common Crawl - Blog - You can now build directly on Common Crawl from the browser Common Crawl - Blog - Host- and Domain-Level Web Graphs February, March, and April 2026 Common Crawl - Blog - April 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Common Crawl Newsletter Common Crawl - Blog - Announcing a Change to Common Crawl Dataset Size Reporting Common Crawl - Blog - Host- and Domain-Level Web Graphs January, February, and March 2026 Common Crawl - Blog - March 2026 Crawl Archive Now Available Common Crawl - Blog - IPv6 Adoption Across the Top 100K Web Hosts Common Crawl - Blog - Web Graph Statistics Gets a Proper Upgrade Common Crawl - Blog - Measuring Web Accessibility from Crawl Archives Common Crawl - Blog - Announcing the Whirlwind Tour of Common Crawl's Datasets Using Java Common Crawl - Blog - Host- and Domain-Level Web Graphs December 2025 and January/February 2026 Common Crawl - Blog - Introducing the New Examples & Resources Browser Common Crawl - Blog - February 2026 Crawl Archive Now Available Common Crawl - Blog - AI Plumbers at FOSDEM’26 Common Crawl - Blog - CC-Citations: A Visualization of Research Papers Referencing Common Crawl Common Crawl - Blog - CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data Common Crawl - Blog - Host- and Domain-Level Web Graphs November/December 2025 and January 2026 Common Crawl - Blog - January 2026 Crawl Archive Now Available Common Crawl - Blog - Web Archives for Social Sciences Datathon, Bristol Common Crawl - Blog - How SEOs Are Using Common Crawl's Web Graph Data for AI Ranking Signals Common Crawl - Blog - GneissWeb Annotations Examples Common Crawl - Blog - Common Crawl at the Mozilla Festival 2025 Common Crawl - Blog - Host- and Domain-Level Web Graphs October, November, December 2025 Common Crawl - Blog - December 2025 Crawl Archive Now Available Common Crawl - Blog - A Sampling of 2025 Research Referencing Common Crawl
Common Crawl - Blog - August/September 2024 Newsletter
2024-09-10 · via Common Crawl

Table of Contents

  • Common Crawl Citations in Academic Research
  • Common Crawl Statistics on Hugging Face
  • Monthly Crawl Updates
  • Updates on our Policy Efforts
  • Roadmap and Future Plans

Common Crawl Citations in Academic Research

Common Crawl's impact on research has grown substantially since its beginning.  Our crawls have become a vital resource for researchers in various fields, from natural language processing to red teaming.

Our data has so far been cited in over 7,000 academic publications, highlighting its value to the research community.  We recently published a blog post on this, and plan to further investigate the connections in this network.

A chart showing the number of citations in Google Scholar mentioning Common Crawl, up to January 2024

Common Crawl Statistics on Hugging Face

We're excited to announce that Common Crawl’s statistics are now available on Hugging Face! The Common Crawl Statistics dataset includes metrics such as the number of URLs, domains, bytes, and content types crawled over specific periods. This dataset is important for users who need comprehensive, structured insights into the composition and trends within the web data collected by Common Crawl. For more details on the statistics see our recent blog post.


Monthly Crawl Updates

We now provide a refreshed crawl every month.  To date, we've delivered over 100 total crawl archives. In August alone, we crawled more than 2.3 billion web pages, amounting to over 320 TiB of uncompressed content.  The total size of our corpus now exceeds 8 PiB, with WARC data alone exceeding 7 PiB—a growth of 10.87% in the past year.  In addition to this, we now also generate Web Graphs on a monthly basis (as opposed to once every three months), which has significantly improved the quality of our recent crawls by giving us fresher ranking data.

A chart showing the cumulative size of WARC data in Common Crawl web archives up to August 2024

Charted data includes WARCs up to August 2024 (CC-MAIN-2024-33)

Updates on our Policy Efforts

We're actively influencing and shaping policy discussions for a free and open Internet.  Earlier this year we hosted a conference in New York titled "AI & the Right to Learn on an Open Internet", co-hosted with Professor Jeff Jarvis at the Craig Newmark Graduate School of Journalism at CUNY.

We have been conducting experiments to evaluate the prevalence of the emerging ML and AI opt-out protocols, and we have been engaging in discussions with our users and collaborators on the best way to put in place, apply, and standardize these protocols.

We're also taking part in a workshop hosted by the Internet Architecture Board in Washington DC in September.  The workshop will explore practical opt-out mechanisms for AI data collection, focusing on how content creators can control the use of their online content in training LLMs and other AI systems.

Roadmap and Future Plans

Based on user feedback, we're pursuing the following initiatives:

  • Significantly increasing our monthly crawl's size, by both depth and breadth, while both remaining a polite crawler and maintaining our data quality.
  • Reducing the carbon footprint of our crawl by modernizing and optimizing our current tools, while also encouraging companies to use our data instead of conducting their own crawls, further reducing the ecological impact of web crawling.
  • Start exploring content-based ranking methods based on the latest Natural Language Processing technologies, allowing us to filter out undesirable content better, improving data quality as we increase the size of our monthly crawls.
  • Develop and release open source and user-friendly tools to allow our users to better explore and understand our dataset.
  • Currently, English represents more than 40% of the data we crawl on a monthly basis.  As part of our efforts to increase the size of our crawl, we want to also make it more representative of the true multilingual nature of the open web, by significantly improving our language identification algorithms.  This in turn will allow us to increase the coverage of our crawls for underrepresented communities.