惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园 - Franky
雷峰网
雷峰网
The Cloudflare Blog
WordPress大学
WordPress大学
博客园 - 聂微东
人人都是产品经理
人人都是产品经理
IT之家
IT之家
V
V2EX
博客园 - 司徒正美
小众软件
小众软件
博客园_首页
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
酷 壳 – CoolShell
酷 壳 – CoolShell
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Hugging Face - Blog
Hugging Face - Blog
T
Tailwind CSS Blog
Last Week in AI
Last Week in AI
Jina AI
Jina AI
博客园 - 叶小钗
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
阮一峰的网络日志
阮一峰的网络日志
爱范儿
爱范儿

Common Crawl

Common Crawl - Blog - Introducing the AI Visibility Audit Common Crawl - Blog - Host- and Domain-Level Web Graphs March, April, and May 2026 Common Crawl - Blog - May 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Crawl Archive Now Available in a Hugging Face Storage Bucket Common Crawl - Blog - You can now build directly on Common Crawl from the browser Common Crawl - Blog - Host- and Domain-Level Web Graphs February, March, and April 2026 Common Crawl - Blog - April 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Common Crawl Newsletter Common Crawl - Blog - Announcing a Change to Common Crawl Dataset Size Reporting Common Crawl - Blog - Host- and Domain-Level Web Graphs January, February, and March 2026 Common Crawl - Blog - March 2026 Crawl Archive Now Available Common Crawl - Blog - IPv6 Adoption Across the Top 100K Web Hosts Common Crawl - Blog - Web Graph Statistics Gets a Proper Upgrade Common Crawl - Blog - Measuring Web Accessibility from Crawl Archives Common Crawl - Blog - Announcing the Whirlwind Tour of Common Crawl's Datasets Using Java Common Crawl - Blog - Host- and Domain-Level Web Graphs December 2025 and January/February 2026 Common Crawl - Blog - Introducing the New Examples & Resources Browser Common Crawl - Blog - February 2026 Crawl Archive Now Available Common Crawl - Blog - AI Plumbers at FOSDEM’26 Common Crawl - Blog - CC-Citations: A Visualization of Research Papers Referencing Common Crawl Common Crawl - Blog - CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data Common Crawl - Blog - Host- and Domain-Level Web Graphs November/December 2025 and January 2026 Common Crawl - Blog - January 2026 Crawl Archive Now Available Common Crawl - Blog - Web Archives for Social Sciences Datathon, Bristol Common Crawl - Blog - How SEOs Are Using Common Crawl's Web Graph Data for AI Ranking Signals Common Crawl - Blog - GneissWeb Annotations Examples Common Crawl - Blog - Common Crawl at the Mozilla Festival 2025 Common Crawl - Blog - Host- and Domain-Level Web Graphs October, November, December 2025 Common Crawl - Blog - December 2025 Crawl Archive Now Available Common Crawl - Blog - A Sampling of 2025 Research Referencing Common Crawl
Common Crawl - Blog - Common Crawl Foundation at ACL 2025
2025-08-13 · via Common Crawl

From 27 July to 1 August 2025, Laurie, Pedro, and Malte from Common Crawl’s engineering team attended the 63rd Annual Meeting of the Association of Computational Linguistics in Vienna, Austria.  ACL is one of the biggest and most prestigious conferences in the field of natural language processing (NLP), with over 6,000 attendees and more than 3,000 accepted papers!

The programme featured keynote talks, oral presentations, poster sessions and social events, plus tutorials on the Sunday before the conference and two days of workshops directly afterwards. It was a great opportunity to learn about the latest work in NLP as well as to develop partnerships with the research community.

Left to right: Malte Ostendorff, Laurie Burchell, and Pedro Ortiz Suarez at ACL 2025 in Vienna

Left to right: Malte Ostendorff, Laurie Burchell, and Pedro Ortiz Suarez at ACL 2025 in Vienna

Research With and By Common Crawl

Many of the papers featured at ACL 2025 made use of Common Crawl’s data products, either indirectly through their use of large language models (LLMs) trained on our data, or directly as a key part of their research. To give just one example, in "Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset" (Su et al., 2025), the authors use our data as the basis for a high-quality English LLM training dataset, leveraging smart data filtering and synthetic rephrasing to improve downstream task performance. We also had a lot of very positive informal feedback from attendees: many told us about the value of Common Crawl’s open data and how it was a key part of making their research happen.

We were also very pleased to have three papers by Common Crawl team members presented at ACL!

Laurie Burchell with a poster on HPLT v2

Laurie Burchell with a poster on HPLT v2

"An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)" (Burchell et al., 2025): presenting HPLT v2, a large-scale collection of high-quality multilingual monolingual and parallel corpora derived from Common Crawl and Internet Archive data.

Pedro Ortiz Suarez with a poster on mOSCAR

Pedro Ortiz Suarez with a poster on mOSCAR

"mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpus" (Futeral et al., 2025): introducing the mOSCAR dataset: the first large-scale multilingual and multimodal document corpus crawled from the web.

Malte, Pedro, and their DFKI colleagues with their award for “best paper runner up” at the 4th Table Representation Learning Workshop

Left to right: Malte Ostendorff, Pedro Ortiz Suarez, Ekaterina Borisova, Georg Rehm, Nils Feldhus, and Raia Abu Ahmad, with their award for “best paper runner up” at the 4th Table Representation Learning Workshop

"Table Understanding and (Multimodal) LLMs: A Cross-Domain Case Study on Scientific vs. Non-Scientific Data" (Borisova et al., 2025): investigating the effectiveness of both text-based and multimodal LLMs on table understanding tasks through a cross-domain and cross-modality evaluation. This paper was the runner up for the best paper award at the Fourth Table Representation Learning Workshop! 🏆

Next Steps

We had a great time at ACL 2025, strengthening our connections within the NLP community and exploring the latest work in the field. We look forward to attending more conferences - come find us at the workshop we’re co-organising at COLM 2025, the First Workshop on Data Quality Signals!