惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

云风的 BLOG
云风的 BLOG
The GitHub Blog
The GitHub Blog
Y
Y Combinator Blog
博客园 - 三生石上(FineUI控件)
T
The Blog of Author Tim Ferriss
宝玉的分享
宝玉的分享
Hugging Face - Blog
Hugging Face - Blog
WordPress大学
WordPress大学
V
Visual Studio Blog
小众软件
小众软件
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
MongoDB | Blog
MongoDB | Blog
V
V2EX
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
博客园 - 【当耐特】
Microsoft Azure Blog
Microsoft Azure Blog
The Cloudflare Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Engineering at Meta
Engineering at Meta
L
LangChain Blog
Martin Fowler
Martin Fowler
GbyAI
GbyAI
博客园 - 司徒正美

Common Crawl

Common Crawl - Blog - Introducing the AI Visibility Audit Common Crawl - Blog - Host- and Domain-Level Web Graphs March, April, and May 2026 Common Crawl - Blog - May 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Crawl Archive Now Available in a Hugging Face Storage Bucket Common Crawl - Blog - You can now build directly on Common Crawl from the browser Common Crawl - Blog - Host- and Domain-Level Web Graphs February, March, and April 2026 Common Crawl - Blog - April 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Common Crawl Newsletter Common Crawl - Blog - Announcing a Change to Common Crawl Dataset Size Reporting Common Crawl - Blog - Host- and Domain-Level Web Graphs January, February, and March 2026 Common Crawl - Blog - March 2026 Crawl Archive Now Available Common Crawl - Blog - IPv6 Adoption Across the Top 100K Web Hosts Common Crawl - Blog - Web Graph Statistics Gets a Proper Upgrade Common Crawl - Blog - Measuring Web Accessibility from Crawl Archives Common Crawl - Blog - Announcing the Whirlwind Tour of Common Crawl's Datasets Using Java Common Crawl - Blog - Host- and Domain-Level Web Graphs December 2025 and January/February 2026 Common Crawl - Blog - Introducing the New Examples & Resources Browser Common Crawl - Blog - February 2026 Crawl Archive Now Available Common Crawl - Blog - AI Plumbers at FOSDEM’26 Common Crawl - Blog - CC-Citations: A Visualization of Research Papers Referencing Common Crawl Common Crawl - Blog - CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data Common Crawl - Blog - Host- and Domain-Level Web Graphs November/December 2025 and January 2026 Common Crawl - Blog - January 2026 Crawl Archive Now Available Common Crawl - Blog - Web Archives for Social Sciences Datathon, Bristol Common Crawl - Blog - How SEOs Are Using Common Crawl's Web Graph Data for AI Ranking Signals Common Crawl - Blog - GneissWeb Annotations Examples Common Crawl - Blog - Common Crawl at the Mozilla Festival 2025 Common Crawl - Blog - Host- and Domain-Level Web Graphs October, November, December 2025 Common Crawl - Blog - December 2025 Crawl Archive Now Available Common Crawl - Blog - A Sampling of 2025 Research Referencing Common Crawl
Common Crawl - Blog - October/November 2024 Newsletter
2024-11-25 · via Common Crawl

Table of Contents

  • Web Languages Project
  • NeurIPS Social with Common Crawl and Wikimedia
  • Event Updates
  • Open Job Positions

Web Languages Project

We have launched the Web Languages project, a volunteer effort with the goal of improving our crawling by making a human-curated list of important non-English websites.

Common Crawl recognizes many languages in its datasets, and we can see that we don't have enough data in languages like Hindi (which has 500+ million speakers!), smaller countries’ languages like Hungarian, and regional languages like Catalan. We are interested in languages from all over the world. By contributing, you can help improve the coverage of underrepresented languages, making a meaningful impact on their visibility and accessibility.
For more details about the project please see our Web Languages GitHub repo, and join our Discord for further discussion and questions.

NeurIPS Social with Common Crawl and Wikimedia

Common Crawl and Wikimedia will host an in-person social event at NeurIPS, Nonprofits Bridging Tech and Social Impact.  If you will be at NeurIPS this December in Vancouver, join us to explore the intersections between nonprofit organizations and the tech community. This session, held on December 11 at 7:30pm, will feature representatives from the Wikimedia Foundation and Common Crawl Foundation, offering an opportunity to connect with nonprofits committed to using technology for social missions.

The event will begin with presentations from both organizations, highlighting their goals, projects and research (e.g., Wikipedia, Common Crawl datasets), and challenges facing the open commons community. Following the presentations, the session will transition into roundtable discussions focused on current initiatives and an open Q&A.

Event Updates

We’ve been busy attending numerous events this Fall.  In late September, we had the privilege of participating in a groundbreaking workshop on AI-CONTROL hosted by the Internet Architecture Board (IAB) in Washington DC.  This event brought together experts from crawling companies, web publishers, AI companies, and “bot defense” companies to discuss the intersection of artificial intelligence and Internet protocols.  For more details, see our blog post.

We attended the IETF 121 meeting in Dublin, where there was further discussion on the initial results from the recent AI CONTROL workshop.  Here are some notes from the chairs Mark Nottingham and Suresh Krishnan.

In October, we had the honor of briefing the White House Office of Science and Technology Policy (OSTP) on the role of The Common Crawl Foundation as critical infrastructure in the artificial intelligence ecosystem and how we can support U.S. federal efforts in advancing responsible AI use and research.  More details about the event and its follow-ups, see our blog post.

In November, we had the opportunity to present at two events, sharing insights into the work of the Common Crawl Foundation and the impact of open web data on research and industry. The first presentation took place at the Turing Institute, as part of the NLP Special Interest Group. The second event was held at University College London, co-hosted with Valyu.  For more on these discussions, see our blog post.

Common Crawl presentation at UCL, November 2024

Open Job Positions

We now have a Jobs page on our website.  Learn about our open roles and how to get in touch with us if you are interested in joining our collaborative team, where your contributions will help shape the future of web data.