惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

L
LangChain Blog
有赞技术团队
有赞技术团队
博客园_首页
IT之家
IT之家
爱范儿
爱范儿
量子位
小众软件
小众软件
Jina AI
Jina AI
WordPress大学
WordPress大学
酷 壳 – CoolShell
酷 壳 – CoolShell
博客园 - 聂微东
The Cloudflare Blog
博客园 - 司徒正美
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
V
V2EX
大猫的无限游戏
大猫的无限游戏
月光博客
月光博客
雷峰网
雷峰网
V
Visual Studio Blog
博客园 - Franky
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
美团技术团队
Last Week in AI
Last Week in AI
S
SegmentFault 最新的问题

Common Crawl

Common Crawl - Blog - Introducing the AI Visibility Audit Common Crawl - Blog - Host- and Domain-Level Web Graphs March, April, and May 2026 Common Crawl - Blog - May 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Crawl Archive Now Available in a Hugging Face Storage Bucket Common Crawl - Blog - You can now build directly on Common Crawl from the browser Common Crawl - Blog - Host- and Domain-Level Web Graphs February, March, and April 2026 Common Crawl - Blog - April 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Common Crawl Newsletter Common Crawl - Blog - Announcing a Change to Common Crawl Dataset Size Reporting Common Crawl - Blog - Host- and Domain-Level Web Graphs January, February, and March 2026 Common Crawl - Blog - March 2026 Crawl Archive Now Available Common Crawl - Blog - IPv6 Adoption Across the Top 100K Web Hosts Common Crawl - Blog - Web Graph Statistics Gets a Proper Upgrade Common Crawl - Blog - Measuring Web Accessibility from Crawl Archives Common Crawl - Blog - Announcing the Whirlwind Tour of Common Crawl's Datasets Using Java Common Crawl - Blog - Host- and Domain-Level Web Graphs December 2025 and January/February 2026 Common Crawl - Blog - Introducing the New Examples & Resources Browser Common Crawl - Blog - February 2026 Crawl Archive Now Available Common Crawl - Blog - AI Plumbers at FOSDEM’26 Common Crawl - Blog - CC-Citations: A Visualization of Research Papers Referencing Common Crawl Common Crawl - Blog - CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data Common Crawl - Blog - Host- and Domain-Level Web Graphs November/December 2025 and January 2026 Common Crawl - Blog - January 2026 Crawl Archive Now Available Common Crawl - Blog - Web Archives for Social Sciences Datathon, Bristol Common Crawl - Blog - How SEOs Are Using Common Crawl's Web Graph Data for AI Ranking Signals Common Crawl - Blog - GneissWeb Annotations Examples Common Crawl - Blog - Common Crawl at the Mozilla Festival 2025 Common Crawl - Blog - Host- and Domain-Level Web Graphs October, November, December 2025 Common Crawl - Blog - December 2025 Crawl Archive Now Available Common Crawl - Blog - A Sampling of 2025 Research Referencing Common Crawl
Common Crawl - Blog - Reflections on Recent Talks at the ...
2024-11-04 · via Common Crawl

At the turn of October into November, our colleagues Thom Vaughan and Pedro Ortiz Suarez had the opportunity to present at two events, sharing insights into the work of the Common Crawl Foundation and the impact of open web data on research and industry.

Turing Institute NLP Special Interest Group

Thom Vaughan, Pedro Ortiz Suarez, Common Crawl Foundation, beside an Enigma machine on loan to the Turing Institute from GCHQ. Photo credit: Robert Blackwell

The first presentation took place at the Turing Institute, as part of the NLP Special Interest Group. Addressing a knowledgeable audience of NLP researchers and practitioners, they discussed how Common Crawl's web-scale data has become a crucial resource for applications in many different areas. Thom and Pedro outlined the dataset's role in training language models and enabling diverse linguistic research, and addressed key challenges associated with curating large-scale web data and the ethical considerations that are inherent in its use. The session concluded with some constructive discussion, which reflected a growing interest in using open data responsibly.

Co-hosted Talk at UCL with Valyu

Thom Vaughan, Pedro Ortiz Suarez, Common Crawl Foundation. Photo credit: Valyu

The second event was held at University College London, co-hosted with Valyu. The talk was on the transformative potential of open datasets for research and innovation. Thom and Pedro showcased examples of how Common Crawl is used in various academic and industrial projects, showing examples of the dataset's contribution to advancements in data science and machine learning. The discussion also focused on strategies to enhance data accessibility and the crucial role of collaboration in promoting a healthy open-data ecosystem. Representatives from Valyu's team Hirsh Pithadia and Harvey Yorke talked about the implications of measured rises of restrictions of data in web archives.

Hirsh Pithadia, Valyu.

Summary

Both events underscored the relevance and importance of accessible web data for driving forward scientific and technological progress. We are grateful to the Turing Institute and UCL, along with Valyu, for facilitating these discussions and for their commitment to advancing the open data landscape. As we continue our work, we look forward to further engaging with the community and supporting new and impactful applications of our datasets.

We’d like to thank Robert Blackwell and Anthony Rhys Hills at the Turing Institute for the opportunity to present at the NLP Special Interest Group, our friends at Valyu for their insightful talk, and Professor Philip Treleaven from UCL for the warm introduction.