惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

G
Google Developers Blog
WordPress大学
WordPress大学
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
小众软件
小众软件
人人都是产品经理
人人都是产品经理
美团技术团队
Blog — PlanetScale
Blog — PlanetScale
S
SegmentFault 最新的问题
博客园 - 【当耐特】
V
V2EX
Microsoft Azure Blog
Microsoft Azure Blog
博客园 - 叶小钗
Google DeepMind News
Google DeepMind News
量子位
罗磊的独立博客
月光博客
月光博客
N
Netflix TechBlog - Medium
大猫的无限游戏
大猫的无限游戏
博客园_首页
P
Proofpoint News Feed
Jina AI
Jina AI
云风的 BLOG
云风的 BLOG
博客园 - 司徒正美
腾讯CDC

Common Crawl

Common Crawl - Blog - Introducing the AI Visibility Audit Common Crawl - Blog - Host- and Domain-Level Web Graphs March, April, and May 2026 Common Crawl - Blog - May 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Crawl Archive Now Available in a Hugging Face Storage Bucket Common Crawl - Blog - You can now build directly on Common Crawl from the browser Common Crawl - Blog - Host- and Domain-Level Web Graphs February, March, and April 2026 Common Crawl - Blog - April 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Common Crawl Newsletter Common Crawl - Blog - Announcing a Change to Common Crawl Dataset Size Reporting Common Crawl - Blog - Host- and Domain-Level Web Graphs January, February, and March 2026 Common Crawl - Blog - March 2026 Crawl Archive Now Available Common Crawl - Blog - IPv6 Adoption Across the Top 100K Web Hosts Common Crawl - Blog - Web Graph Statistics Gets a Proper Upgrade Common Crawl - Blog - Measuring Web Accessibility from Crawl Archives Common Crawl - Blog - Announcing the Whirlwind Tour of Common Crawl's Datasets Using Java Common Crawl - Blog - Host- and Domain-Level Web Graphs December 2025 and January/February 2026 Common Crawl - Blog - Introducing the New Examples & Resources Browser Common Crawl - Blog - February 2026 Crawl Archive Now Available Common Crawl - Blog - AI Plumbers at FOSDEM’26 Common Crawl - Blog - CC-Citations: A Visualization of Research Papers Referencing Common Crawl Common Crawl - Blog - CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data Common Crawl - Blog - Host- and Domain-Level Web Graphs November/December 2025 and January 2026 Common Crawl - Blog - January 2026 Crawl Archive Now Available Common Crawl - Blog - Web Archives for Social Sciences Datathon, Bristol Common Crawl - Blog - How SEOs Are Using Common Crawl's Web Graph Data for AI Ranking Signals Common Crawl - Blog - GneissWeb Annotations Examples Common Crawl - Blog - Common Crawl at the Mozilla Festival 2025 Common Crawl - Blog - Host- and Domain-Level Web Graphs October, November, December 2025 Common Crawl - Blog - December 2025 Crawl Archive Now Available Common Crawl - Blog - A Sampling of 2025 Research Referencing Common Crawl
Common Crawl - Blog - July/August 2025 Newsletter
2025-08-26 · via Common Crawl

Table of Contents

Stanford HAI Seminar in October

Summer Event Highlights

The First WMDQS-Masakhane LangID Hackathon

SEO to AIO, Search 1.0 to 2.0

The Common Crawl engineering team’s weekly meeting

The Common Crawl engineering team’s weekly meeting

Stanford HAI Seminar in October

Common Crawl Foundation is thrilled to present at an upcoming Stanford Institute for Human-Centered Artificial Intelligence (HAI) Seminar entitled Preserving Humanity's Knowledge and Making it Accessible: Addressing Challenges of Public Web Data.

Learn about Common Crawl's insights from a recent data product and informed solutions for the future of public web data. The seminar is Wednesday, October 22 from noon to 1:15 pm.  For registration (in person and virtual) and more details please see the event listing. The Common Crawl team (including several of our engineers!) will be around for a few hours after the talk for followup chats.

Summer Event Highlights

The Common Crawl team attended the 63rd Annual Meeting of the Association of Computational Linguistics (ACL) in Vienna, presenting recent published work and strengthening links with the research community.  More details about the event and links to papers with and about Common Crawl can be found in our recent blog post.

In July we had the happy opportunity to attend IETF 123, held at the Meliã Castilla in Madrid. As ever, the event was packed full of discussions, new draft proposals, and connections from the Internet protocol community. More details in our blog post.

And, back in June the Common Crawl Foundation team was in New York City for the United Nations Open Source Week, and several industry side-events.  Over the course of the week we engaged with developers, researchers, and policymakers on all things related to Open Source and AI.  For highlights from the week, see our blog post.

The First WMDQS-Masakhane LangID Hackathon

In June 2025 the Common Crawl Foundation, MLCommons, and EleutherAI had the pleasure of hosting a virtual hackathon in partnership with Masakhane in order to collect language identification annotations for African languages.  For more about the hackathon as well as the Shared Task on Improving Language Identification for Web Text (to be held at COLM in October) see our blog post.

SEO to AIO, Search 1.0 to 2.0

Publishers and brands are shifting from SEO to AIO. Many SEOs unknowingly block their sites from AI search by restricting CCBot in robots.txt. As Search 2.0 transforms discovery, ensuring content can train AI models becomes as crucial as traditional SEO. Read our recent in-depth post on this topic here.