惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

G
Google Developers Blog
阮一峰的网络日志
阮一峰的网络日志
A
About on SuperTechFans
大猫的无限游戏
大猫的无限游戏
Engineering at Meta
Engineering at Meta
V
Visual Studio Blog
Martin Fowler
Martin Fowler
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
博客园 - 叶小钗
I
InfoQ
B
Blog RSS Feed
aimingoo的专栏
aimingoo的专栏
Y
Y Combinator Blog
Blog — PlanetScale
Blog — PlanetScale
IT之家
IT之家
P
Proofpoint News Feed
WordPress大学
WordPress大学
小众软件
小众软件
B
Blog
MongoDB | Blog
MongoDB | Blog
人人都是产品经理
人人都是产品经理
量子位
Hugging Face - Blog
Hugging Face - Blog
月光博客
月光博客

Common Crawl

Common Crawl - Blog - Introducing the AI Visibility Audit Common Crawl - Blog - Host- and Domain-Level Web Graphs March, April, and May 2026 Common Crawl - Blog - May 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Crawl Archive Now Available in a Hugging Face Storage Bucket Common Crawl - Blog - You can now build directly on Common Crawl from the browser Common Crawl - Blog - Host- and Domain-Level Web Graphs February, March, and April 2026 Common Crawl - Blog - April 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Common Crawl Newsletter Common Crawl - Blog - Announcing a Change to Common Crawl Dataset Size Reporting Common Crawl - Blog - Host- and Domain-Level Web Graphs January, February, and March 2026 Common Crawl - Blog - March 2026 Crawl Archive Now Available Common Crawl - Blog - IPv6 Adoption Across the Top 100K Web Hosts Common Crawl - Blog - Web Graph Statistics Gets a Proper Upgrade Common Crawl - Blog - Measuring Web Accessibility from Crawl Archives Common Crawl - Blog - Announcing the Whirlwind Tour of Common Crawl's Datasets Using Java Common Crawl - Blog - Host- and Domain-Level Web Graphs December 2025 and January/February 2026 Common Crawl - Blog - Introducing the New Examples & Resources Browser Common Crawl - Blog - February 2026 Crawl Archive Now Available Common Crawl - Blog - AI Plumbers at FOSDEM’26 Common Crawl - Blog - CC-Citations: A Visualization of Research Papers Referencing Common Crawl Common Crawl - Blog - CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data Common Crawl - Blog - Host- and Domain-Level Web Graphs November/December 2025 and January 2026 Common Crawl - Blog - January 2026 Crawl Archive Now Available Common Crawl - Blog - Web Archives for Social Sciences Datathon, Bristol Common Crawl - Blog - How SEOs Are Using Common Crawl's Web Graph Data for AI Ranking Signals Common Crawl - Blog - GneissWeb Annotations Examples Common Crawl - Blog - Common Crawl at the Mozilla Festival 2025 Common Crawl - Blog - Host- and Domain-Level Web Graphs October, November, December 2025 Common Crawl - Blog - December 2025 Crawl Archive Now Available Common Crawl - Blog - A Sampling of 2025 Research Referencing Common Crawl
Common Crawl - Blog - January/February 2025 Newsletter
2025-02-03 · via Common Crawl

Table of Contents

  • Annotation for Language Identification
  • cc-downloader Command Line Tool
  • Citations Updates
  • Common Crawl at SXSW 2025
  • Software Heritage Symposium at UNESCO
  • NeurIPS 2024 Social with Wikimedia

Annotation for Language Identification

In December we introduced an annotation campaign for Language Identification (LID or LangID) that we will conduct in collaboration with MLCommons. In this annotation campaign we will ask participants to do simple LangID annotations on Common Crawl data. We would like to get as many annotations as possible and cover as many languages as possible, in order to create the first web-based LangID dataset. Our ultimate goal with this project is to train a small language classifier that would help us make better decisions at crawl time ensuring that we crawl data for as many languages as possible, so that our dataset will hopefully better reflect the vast cultural and linguistic diversity of the web.

If you would like to contribute and participate in our annotation campaign, please visit MLCommons' Dynabench Platform.  For more details about our efforts to expand language coverage in Common Crawl, including LangID and our Web Languages project, see our related blog post.

cc-downloader Command Line Tool

We recently introduced cc-downloader, an experimental command-line tool for downloading Common Crawl data via HTTPS.  cc-downloader is intended to be a user-friendly and polite downloader.  For more details, please visit the cc-downloader GitHub repository and the related blog post.

Citations Updates

We have recently updated our Common Crawl Citations to include 2024 research paper citations.  Please see our updated Research Papers Citations graph for a look at Common Crawl citations in research papers through 2024.

Plot of Common Crawl citations (cumulative) in Google Scholar until January 2025

Source: cc-citations

Common Crawl at SXSW 2025

Common Crawl will be at SXSW in March.  If you will be in Austin that week we would love to meet up with you.  Please get in touch with us if you would like to arrange a coffee or meet-up.

Software Heritage Symposium at UNESCO

Left to right: Thom Vaughan, Pedro Ortiz Suarez at UNESCO Headquarters, Paris

On 29 January 2025, members of the Common Crawl Foundation attended the Software Heritage Symposium at UNESCO Headquarters in Paris. The event brought together experts from academia, industry, and policy to discuss key topics such as cybersecurity, AI transparency, open science, and cultural preservation. Speakers highlighted the role of open infrastructures in building a secure and inclusive digital future.

NeurIPS 2024 Social with Wikimedia

The Common Crawl Foundation attended NeurIPS 2024, connecting with organizations, hosting a social event on tech and social impact, and showcasing contributions to AI research and data access.  For more details on our social with Wikimedia and additional conference highlights, please see our related blog post.