惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
Jina AI
Jina AI
月光博客
月光博客
F
Fortinet All Blogs
Stack Overflow Blog
Stack Overflow Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
V
Visual Studio Blog
小众软件
小众软件
博客园 - 三生石上(FineUI控件)
博客园 - 司徒正美
P
Proofpoint News Feed
酷 壳 – CoolShell
酷 壳 – CoolShell
M
MIT News - Artificial intelligence
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
B
Blog RSS Feed
Apple Machine Learning Research
Apple Machine Learning Research
S
SegmentFault 最新的问题
博客园_首页
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
J
Java Code Geeks
L
LangChain Blog
博客园 - 聂微东
G
Google Developers Blog
博客园 - Franky

Common Crawl

Common Crawl - Blog - Host- and Domain-Level Web Graphs March, April, and May 2026 Common Crawl - Blog - May 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Crawl Archive Now Available in a Hugging Face Storage Bucket Common Crawl - Blog - You can now build directly on Common Crawl from the browser Common Crawl - Blog - Host- and Domain-Level Web Graphs February, March, and April 2026 Common Crawl - Blog - April 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Common Crawl Newsletter Common Crawl - Blog - Announcing a Change to Common Crawl Dataset Size Reporting Common Crawl - Blog - Host- and Domain-Level Web Graphs January, February, and March 2026 Common Crawl - Blog - March 2026 Crawl Archive Now Available Common Crawl - Blog - IPv6 Adoption Across the Top 100K Web Hosts Common Crawl - Blog - Web Graph Statistics Gets a Proper Upgrade Common Crawl - Blog - Measuring Web Accessibility from Crawl Archives Common Crawl - Blog - Announcing the Whirlwind Tour of Common Crawl's Datasets Using Java Common Crawl - Blog - Host- and Domain-Level Web Graphs December 2025 and January/February 2026 Common Crawl - Blog - Introducing the New Examples & Resources Browser Common Crawl - Blog - February 2026 Crawl Archive Now Available Common Crawl - Blog - AI Plumbers at FOSDEM’26 Common Crawl - Blog - CC-Citations: A Visualization of Research Papers Referencing Common Crawl Common Crawl - Blog - CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data Common Crawl - Blog - Host- and Domain-Level Web Graphs November/December 2025 and January 2026 Common Crawl - Blog - January 2026 Crawl Archive Now Available Common Crawl - Blog - Web Archives for Social Sciences Datathon, Bristol Common Crawl - Blog - How SEOs Are Using Common Crawl's Web Graph Data for AI Ranking Signals Common Crawl - Blog - GneissWeb Annotations Examples Common Crawl - Blog - Common Crawl at the Mozilla Festival 2025 Common Crawl - Blog - Host- and Domain-Level Web Graphs October, November, December 2025 Common Crawl - Blog - December 2025 Crawl Archive Now Available Common Crawl - Blog - A Sampling of 2025 Research Referencing Common Crawl Common Crawl - Blog - Host- and Domain-Level Web Graphs September, October, and November 2025
Common Crawl - Blog - Common Crawl Foundation at IIPC-WAC...
2026-06-10 · via Common Crawl

Members of the Common Crawl Foundation team (Laurie Burchell, Sebastian Nagel, and Pedro Ortiz Suarez) attended the 2026 IIPC Web Archiving Conference (WAC) and General Assembly (GA) held at KBR, the Royal Library of Belgium in Brussels.

Royal Library of Belgium, picture by EmDee CC BY-SA 4.0

Themes for this year's conference consisted of Access & Research Use, Tools & Infrastructure, Collection Development, Legal & Ethical Issues, Policies & Standards, and Environmental Impact. In accordance with that last theme, participants were asked to use the stairs over the elevators.

Common Crawl Contributions

Common Crawl was well represented across the event, with both presentations and references in the work of other participants.

A photograph of Sebastian Nagel, Laurie Burchell, and Pedro Ortiz Suarez.

Sebastian Nagel, Laurie Burchell, and Pedro Ortiz Suarez.

Work on improved language identification for crawl data was presented by Laurie.  Pedro presented statistics on data-access methods used to download Common Crawl data and showcased cc-downloader adoption over the last year, with Sebastian presenting the research on Web crawling policies and opt-outs at the General Assembly done in collaboration with CCF.  Presentations are expected to be posted to the IIPC YouTube channel in the near future.

A photograph of Sebastian Nagel giving a presentation on CCBot

Sebastian Nagel presenting CCBot

Common Crawl’s data products were often referenced, both during talks and behind the scenes, notable mentions include the “Responsible Strategies” session by Abbie Grotke, "End of Term Web Archive: Harmonizing WARC contributions from multiple crawling partner" presented by Mark Phillips, and "Crawl, cloud, carbon: measuring and reducing emissions for web archivists" by Simon Ponsford.

We look forward to more discussions with our friends (new and old) from the IIPC in the near future.

A photograph of Laurie Burchell presenting CommonLID at Howest

Laurie Burchell presenting CommonLID at Howest

A photograph of the closing panel: Web Archiving for Accountability, Shown from left to right: Emily Tripp, Marvin Milatz, Friedhelm Weinberg, Basile Simon

Closing panel: Web Archiving for Accountability, shown from left to right: Emily Tripp, Marvin Milatz, Friedhelm Weinberg, Basile Simon