惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

D
Docker
B
Blog RSS Feed
Microsoft Security Blog
Microsoft Security Blog
Y
Y Combinator Blog
N
Netflix TechBlog - Medium
M
MIT News - Artificial intelligence
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
B
Blog
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
C
Check Point Blog
The GitHub Blog
The GitHub Blog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
P
Proofpoint News Feed
Martin Fowler
Martin Fowler
大猫的无限游戏
大猫的无限游戏
GbyAI
GbyAI
博客园_首页
A
About on SuperTechFans
Blog — PlanetScale
Blog — PlanetScale
人人都是产品经理
人人都是产品经理
T
Tailwind CSS Blog
aimingoo的专栏
aimingoo的专栏
T
The Blog of Author Tim Ferriss
The Cloudflare Blog

Common Crawl

Common Crawl - Blog - Introducing the AI Visibility Audit Common Crawl - Blog - Host- and Domain-Level Web Graphs March, April, and May 2026 Common Crawl - Blog - May 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Crawl Archive Now Available in a Hugging Face Storage Bucket Common Crawl - Blog - You can now build directly on Common Crawl from the browser Common Crawl - Blog - Host- and Domain-Level Web Graphs February, March, and April 2026 Common Crawl - Blog - April 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Common Crawl Newsletter Common Crawl - Blog - Announcing a Change to Common Crawl Dataset Size Reporting Common Crawl - Blog - Host- and Domain-Level Web Graphs January, February, and March 2026 Common Crawl - Blog - March 2026 Crawl Archive Now Available Common Crawl - Blog - IPv6 Adoption Across the Top 100K Web Hosts Common Crawl - Blog - Web Graph Statistics Gets a Proper Upgrade Common Crawl - Blog - Measuring Web Accessibility from Crawl Archives Common Crawl - Blog - Announcing the Whirlwind Tour of Common Crawl's Datasets Using Java Common Crawl - Blog - Host- and Domain-Level Web Graphs December 2025 and January/February 2026 Common Crawl - Blog - Introducing the New Examples & Resources Browser Common Crawl - Blog - February 2026 Crawl Archive Now Available Common Crawl - Blog - AI Plumbers at FOSDEM’26 Common Crawl - Blog - CC-Citations: A Visualization of Research Papers Referencing Common Crawl Common Crawl - Blog - CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data Common Crawl - Blog - Host- and Domain-Level Web Graphs November/December 2025 and January 2026 Common Crawl - Blog - January 2026 Crawl Archive Now Available Common Crawl - Blog - Web Archives for Social Sciences Datathon, Bristol Common Crawl - Blog - How SEOs Are Using Common Crawl's Web Graph Data for AI Ranking Signals Common Crawl - Blog - GneissWeb Annotations Examples Common Crawl - Blog - Common Crawl at the Mozilla Festival 2025 Common Crawl - Blog - Host- and Domain-Level Web Graphs October, November, December 2025 Common Crawl - Blog - December 2025 Crawl Archive Now Available Common Crawl - Blog - A Sampling of 2025 Research Referencing Common Crawl
Common Crawl - Blog - The First WMDQS-Masakhane LangID Ha...
2025-07-08 · via Common Crawl

Since the end of 2024, the Common Crawl Foundation has committed to expanding the language coverage of its crawls in order to facilitate the creation of web and language technologies for underrepresented languages. In this effort to improve coverage, we have already started two initiatives: the Web Languages Project, where the community can contribute URLs in underrepresented languages for our seed crawl, and the LangID Project where users can add language identification (LangID or LID) to Common Crawl data, in order for us to improve the models we use to annotate our crawls, and to discover new data.

In line with these two initiatives, we also announced the 1st Workshop on Multilingual Data Quality Signals that we are organizing in collaboration with our colleagues at MLCommons, EleutherAI, and Johns Hopkins' HLTCOE. This workshop which will be collocated with COLM 2025, in Montréal, Canada, will also host a shared task on language identification where we expect to collect more annotations for our LangID, and then develop new LangID solutions with participants that are robust, lightweight, and open source, and that can be later maintained by us and our collaborators.

In the context of this shared task, Common Crawl and MLCommons hosted a hackathon on June 26th, to collect LangID annotations for African languages, in collaboration with our friends and colleagues at Masakhane and the Data Science for Social Impact research group at the University of Pretoria. We’re pleased to share that the hackathon was a great success, generating approximately 5,000 document annotations and bringing our total to over 17,000 across more than 70 languages. We attribute much of this momentum to the hackathon’s impact.

This hackathon allowed us to set a solid foundation for our dataset and our shared task, and constitutes a large community towards developing web and language technologies for African languages. As such, we would like to express our deepest gratitude to Masakahne, the Data Science for Social Impact research group, and all of the community members who participated in the hackathon and those who have continued to contribute with annotations and feedback. We would also like to thank Idris Abdulmumin and Vukosi Marivate in particular, who made this hackathon possible.

The WMDQS LangID shared task remains open to contributions, and for those who would like to participate directly, especially with new LangID models and solutions, registration is open until the 21st of July.