惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

阮一峰的网络日志
阮一峰的网络日志
雷峰网
雷峰网
Last Week in AI
Last Week in AI
T
Tailwind CSS Blog
V
Visual Studio Blog
Jina AI
Jina AI
博客园 - 司徒正美
The Cloudflare Blog
Hugging Face - Blog
Hugging Face - Blog
博客园_首页
S
SegmentFault 最新的问题
博客园 - 三生石上(FineUI控件)
有赞技术团队
有赞技术团队
小众软件
小众软件
V
V2EX
Apple Machine Learning Research
Apple Machine Learning Research
美团技术团队
博客园 - 【当耐特】
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
IT之家
IT之家
WordPress大学
WordPress大学
爱范儿
爱范儿
月光博客
月光博客
大猫的无限游戏
大猫的无限游戏

Common Crawl

Common Crawl - Blog - Introducing the AI Visibility Audit Common Crawl - Blog - Host- and Domain-Level Web Graphs March, April, and May 2026 Common Crawl - Blog - May 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Crawl Archive Now Available in a Hugging Face Storage Bucket Common Crawl - Blog - You can now build directly on Common Crawl from the browser Common Crawl - Blog - Host- and Domain-Level Web Graphs February, March, and April 2026 Common Crawl - Blog - April 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Common Crawl Newsletter Common Crawl - Blog - Announcing a Change to Common Crawl Dataset Size Reporting Common Crawl - Blog - Host- and Domain-Level Web Graphs January, February, and March 2026 Common Crawl - Blog - March 2026 Crawl Archive Now Available Common Crawl - Blog - IPv6 Adoption Across the Top 100K Web Hosts Common Crawl - Blog - Web Graph Statistics Gets a Proper Upgrade Common Crawl - Blog - Measuring Web Accessibility from Crawl Archives Common Crawl - Blog - Announcing the Whirlwind Tour of Common Crawl's Datasets Using Java Common Crawl - Blog - Host- and Domain-Level Web Graphs December 2025 and January/February 2026 Common Crawl - Blog - Introducing the New Examples & Resources Browser Common Crawl - Blog - February 2026 Crawl Archive Now Available Common Crawl - Blog - AI Plumbers at FOSDEM’26 Common Crawl - Blog - CC-Citations: A Visualization of Research Papers Referencing Common Crawl Common Crawl - Blog - CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data Common Crawl - Blog - Host- and Domain-Level Web Graphs November/December 2025 and January 2026 Common Crawl - Blog - January 2026 Crawl Archive Now Available Common Crawl - Blog - Web Archives for Social Sciences Datathon, Bristol Common Crawl - Blog - How SEOs Are Using Common Crawl's Web Graph Data for AI Ranking Signals Common Crawl - Blog - GneissWeb Annotations Examples Common Crawl - Blog - Common Crawl at the Mozilla Festival 2025 Common Crawl - Blog - Host- and Domain-Level Web Graphs October, November, December 2025 Common Crawl - Blog - December 2025 Crawl Archive Now Available Common Crawl - Blog - A Sampling of 2025 Research Referencing Common Crawl
Common Crawl - Blog - Expanding the Language and Cultural...
2024-12-11 · via Common Crawl

At Common Crawl our mission has always been to make Open Web Data easily accessible for our users, so that they can benefit from high quality crawl data that was previously only available to large search engine corporations. However, from our own statistics, we know that our data has always been biased towards English content making our dataset difficult to use for individuals and organizations from smaller linguistic communities.

We have always wanted to make Common Crawl as representative as possible of the Open Web, so in recent months we have been working on some projects that we hope will allow us to expand the language and cultural coverage of our crawls, making it more representative of the actual linguistic and cultural diversity found on the web.

These projects will require input from the community, as our team is small and we speak but a handful of languages, and as we believe that the languages and the content written in them belong in the end to their respective linguistic communities.

The first initiative that we’re introducing today is the Web Languages Project. With this, we are asking speakers of Languages Other Than English (LOTE), to contribute URLs of websites that they know and that contain content written in their language. We will then inject these URLs into our seed crawl, which we hope will allow us to discover more web content written in these languages. We will of course respect Robots Exclusion Protocol directives, ensuring that all this new linguistic content that we will discover is crawled as politely as we have always crawled. If you want to contribute to this project please visit our GitHub Repository for more instructions.

The second initiative that we’re introducing is an annotation campaign for Language Identification (LID or LangID) that we will conduct in collaboration with MLCommons. In this annotation campaign we will ask participants to do simple LangID annotations on Common Crawl data. We would like to get as many annotations as possible and cover as many languages as possible, in order to create the first web-based LangID dataset. Our ultimate goal with this project is to train a small language classifier that would help us make better decisions at crawl time ensuring that we crawl data for as many languages as possible, so that our dataset will hopefully better reflect the vast cultural and linguistic diversity of the web. If you want to contribute and participate in our annotation campaign, please visit MLCommon’s Dynabench Platform, where you can already start annotating data today.

Dynabench interface showing highlighting of multilingual text

Interface in Dynabench

Finally, if you want to join the conversation about this project please join our Discord, there you will be able to share your feedback and engage with other contributors and community members.