惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

GbyAI
GbyAI
The GitHub Blog
The GitHub Blog
小众软件
小众软件
美团技术团队
博客园 - 司徒正美
G
Google Developers Blog
Blog — PlanetScale
Blog — PlanetScale
Hugging Face - Blog
Hugging Face - Blog
博客园_首页
大猫的无限游戏
大猫的无限游戏
罗磊的独立博客
Recent Announcements
Recent Announcements
酷 壳 – CoolShell
酷 壳 – CoolShell
D
Docker
J
Java Code Geeks
Last Week in AI
Last Week in AI
V
Visual Studio Blog
Microsoft Azure Blog
Microsoft Azure Blog
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
P
Proofpoint News Feed
V
V2EX
C
Check Point Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
MyScale Blog
MyScale Blog

Common Crawl

Common Crawl - Blog - Introducing the AI Visibility Audit Common Crawl - Blog - Host- and Domain-Level Web Graphs March, April, and May 2026 Common Crawl - Blog - May 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Crawl Archive Now Available in a Hugging Face Storage Bucket Common Crawl - Blog - You can now build directly on Common Crawl from the browser Common Crawl - Blog - Host- and Domain-Level Web Graphs February, March, and April 2026 Common Crawl - Blog - April 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Common Crawl Newsletter Common Crawl - Blog - Announcing a Change to Common Crawl Dataset Size Reporting Common Crawl - Blog - Host- and Domain-Level Web Graphs January, February, and March 2026 Common Crawl - Blog - March 2026 Crawl Archive Now Available Common Crawl - Blog - IPv6 Adoption Across the Top 100K Web Hosts Common Crawl - Blog - Web Graph Statistics Gets a Proper Upgrade Common Crawl - Blog - Measuring Web Accessibility from Crawl Archives Common Crawl - Blog - Announcing the Whirlwind Tour of Common Crawl's Datasets Using Java Common Crawl - Blog - Host- and Domain-Level Web Graphs December 2025 and January/February 2026 Common Crawl - Blog - Introducing the New Examples & Resources Browser Common Crawl - Blog - February 2026 Crawl Archive Now Available Common Crawl - Blog - AI Plumbers at FOSDEM’26 Common Crawl - Blog - CC-Citations: A Visualization of Research Papers Referencing Common Crawl Common Crawl - Blog - CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data Common Crawl - Blog - Host- and Domain-Level Web Graphs November/December 2025 and January 2026 Common Crawl - Blog - January 2026 Crawl Archive Now Available Common Crawl - Blog - Web Archives for Social Sciences Datathon, Bristol Common Crawl - Blog - How SEOs Are Using Common Crawl's Web Graph Data for AI Ranking Signals Common Crawl - Blog - GneissWeb Annotations Examples Common Crawl - Blog - Common Crawl at the Mozilla Festival 2025 Common Crawl - Blog - Host- and Domain-Level Web Graphs October, November, December 2025 Common Crawl - Blog - December 2025 Crawl Archive Now Available Common Crawl - Blog - A Sampling of 2025 Research Referencing Common Crawl
Common Crawl - Blog - March/April 2025 Newsletter
2025-04-14 · via Common Crawl

Table of Contents

Event Updates

AI Action Summit + ROOST Launch

Submission to UK Copyright and AI Consultation

Common Crawl AI Agent by Ready AI

Language Updates

Event Updates

We have been busy participating in events this Winter and Spring.  In February, we presented at HPLT Winter School, which had a focus this year on Pre-training Data Quality and Multilingual LLM Evaluation.  Also in February, we attended the AI Action Summit (see separate post below).

Valyu x Common Crawl x UCL: AI Agents, Crawling and the Future of the Web was a co-hosted event in London this February to discuss AI-driven retrieval, web crawling for AI agents, AI preference signaling, opt-in/opt-out models, and related topics.

In March we attended SXSW in Austin, and hosted a networking social, attended others’ socials and events, and met up with many partners and friends of Common Crawl.

Left-to-right: Benjamin Diggles, Chris Tolles, Erik Bethel, Aidan Clifford, speaking at the Digital Chamber Summit in Washington DC in 2025

Left-to-right: Benjamin Diggles, Chris Tolles, Erik Bethel, Aidan Clifford, speaking at the Digital Chamber Summit in Washington DC in 2025

Also in March, we participated in a panel discussion on AI and blockchain with partner Constellation Network at the DC Blockchain Summit.  Watch the complete panel discussion here, and learn more about Constellation Network’s launch at the summit of Digital Evidence product in our blog post.

In April, we attended the IIPC Web Archiving Conference.  Stayed tuned for a full report in a separate blog post coming soon!

AI Action Summit + ROOST Launch

In February, Common Crawl attended the AI Action Summit in Paris, which saw the launch of several projects, standards, and partnerships.  A coalition of major technology companies and foundations announced the launch of ROOST: Robust Online Open Safety Tools (https://roost.tools). ROOST makes critical data and tools for online safety openly accessible to benefit everyone; a mission which closely aligns with ours at Common Crawl. To learn more about the ROOST launch, please see our blog post.

Submission to UK Copyright and AI Consultation

The frontispiece of Ted Nelson's Computer Lib/Dream Machines (1974). Original image by John R. Neill for L. Frank Baum's Tik-tok of Oz (1914).

The frontispiece of Ted Nelson's Computer Lib/Dream Machines (1974). Original image by John R. Neill for L. Frank Baum's Tik-tok of Oz (1914).

Common Crawl made a submission to the UK Copyright and AI Consultation supporting a legal exception for text and data mining (TDM) while respecting creators’ rights.  Read the full submission in our blog post.

Common Crawl AI Agent by Ready AI

Two logos side-by-side, ReadyAI and Common Crawl

Announcing the launch of an experimental AI Agent, developed by our friends at ReadyAI

We recently announced the launch of an experimental AI Agent, developed by our friends at ReadyAI. The agent offers a conversational interface designed to help users explore Common Crawl’s data, use cases, and community initiatives. Learn more about the agent in our blog post, and try it out here.

Language Updates

At the end of last year, we introduced two new language initiatives, LangID and web-languages.  LangID, our annotation campaign for language identification in collaboration with MLCommons, now has over 600 contributions.  Learn more and contribute to the LangID task here.  Our web-languages project, in which we are asking speakers of Languages Other Than English (LOTE), to contribute URLs of websites that they know and that contain content written in their language, has had 21 pull requests merged since our last newsletter.  Our web-languages GitHub repo has more details on contributing to the project.

A chart showing modified language files in our web-languages repository

A chart showing modified language files in our web-languages repository