惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Apple Machine Learning Research
Apple Machine Learning Research
Jina AI
Jina AI
博客园_首页
WordPress大学
WordPress大学
罗磊的独立博客
小众软件
小众软件
Last Week in AI
Last Week in AI
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
Hugging Face - Blog
Hugging Face - Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
爱范儿
爱范儿
The Cloudflare Blog
GbyAI
GbyAI
C
Check Point Blog
腾讯CDC
MyScale Blog
MyScale Blog
有赞技术团队
有赞技术团队
博客园 - 聂微东
IT之家
IT之家
雷峰网
雷峰网
H
Help Net Security
博客园 - 叶小钗
美团技术团队
D
DataBreaches.Net

Common Crawl

Common Crawl - Blog - Introducing the AI Visibility Audit Common Crawl - Blog - Host- and Domain-Level Web Graphs March, April, and May 2026 Common Crawl - Blog - May 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Crawl Archive Now Available in a Hugging Face Storage Bucket Common Crawl - Blog - You can now build directly on Common Crawl from the browser Common Crawl - Blog - Host- and Domain-Level Web Graphs February, March, and April 2026 Common Crawl - Blog - April 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Common Crawl Newsletter Common Crawl - Blog - Announcing a Change to Common Crawl Dataset Size Reporting Common Crawl - Blog - Host- and Domain-Level Web Graphs January, February, and March 2026 Common Crawl - Blog - March 2026 Crawl Archive Now Available Common Crawl - Blog - IPv6 Adoption Across the Top 100K Web Hosts Common Crawl - Blog - Web Graph Statistics Gets a Proper Upgrade Common Crawl - Blog - Measuring Web Accessibility from Crawl Archives Common Crawl - Blog - Announcing the Whirlwind Tour of Common Crawl's Datasets Using Java Common Crawl - Blog - Host- and Domain-Level Web Graphs December 2025 and January/February 2026 Common Crawl - Blog - Introducing the New Examples & Resources Browser Common Crawl - Blog - February 2026 Crawl Archive Now Available Common Crawl - Blog - AI Plumbers at FOSDEM’26 Common Crawl - Blog - CC-Citations: A Visualization of Research Papers Referencing Common Crawl Common Crawl - Blog - CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data Common Crawl - Blog - Host- and Domain-Level Web Graphs November/December 2025 and January 2026 Common Crawl - Blog - January 2026 Crawl Archive Now Available Common Crawl - Blog - Web Archives for Social Sciences Datathon, Bristol Common Crawl - Blog - How SEOs Are Using Common Crawl's Web Graph Data for AI Ranking Signals Common Crawl - Blog - GneissWeb Annotations Examples Common Crawl - Blog - Common Crawl at the Mozilla Festival 2025 Common Crawl - Blog - Host- and Domain-Level Web Graphs October, November, December 2025 Common Crawl - Blog - December 2025 Crawl Archive Now Available Common Crawl - Blog - A Sampling of 2025 Research Referencing Common Crawl
Common Crawl - Blog - Common Crawl at the United Nations ...
2025-06-30 · via Common Crawl

Pedro Ortiz Suarez and Sebastian Nagel at the United Nations in New York, attending the UN Open Source Week, NY.

Left-to-right: Pedro Ortiz Suarez and Sebastian Nagel at the United Nations in New York, attending the UN Open Source Week, NY.

From the 16th to the 20th of June, the Common Crawl Foundation team was in New York City for the United Nations Open Source Week, and select industry side-events.  Over the course of the week we engaged with developers, researchers, and policymakers on all things related to Open Source and AI. We presented at IBM’s Thomas J. Watson Research Center, and co-hosted the “AI Unconference” event at IBM One Madison: a gathering designed for open discussions of what we see as some of the most important issues facing the industry today: transparency, safety, diversity, and the importance of ethical data pipelines.

UN Open Source Maintain-a-thon

The CCF team attended the United Nations for the “Maintain-a-thon”, as part of UN Open Source Week, NY.

The CCF team attended the United Nations for the “Maintain-a-thon”, as part of UN Open Source Week, NY.

Our team attended the United Nations for the Open Source Maintain-a-thon on Tuesday.  Attendees from numerous global organisations split into groups and produced “Today I Learned” takeaways, “Tomorrow I Will” actions, and “Gee, I Wish” ideas for application in various areas of the (AI) industry. This culminated in a collective playbook for maintainability which will be released at a later date via the United Nations website.

IBM Thomas J. Watson Research Center

Inside the IBM Thomas J. Watson Research Center, Yorktown Heights, NY

Inside the IBM Thomas J. Watson Research Center, Yorktown Heights, NY.

The team travelled to Yorktown Heights to IBM’s Thomas J. Watson Research Center, where our distinguished engineer Sebastian Nagel gave a series of presentations on Common Crawl’s activities, goals, and partnerships.  Our team then met with dozens of representatives from departments across IBM to discuss mutual goals and identify areas where collaboration might benefit the industry at large.

Sebastian Nagel (Distinguished Engineer, Common Crawl) presenting at the IBM Thomas J. Watson Research Center, Yorktown Heights, NY.

Sebastian Nagel (Distinguished Engineer, Common Crawl) presenting at the IBM Thomas J. Watson Research Center, Yorktown Heights, NY.

Side-events at LinkedIn, Meta, and PwC

Common Crawl Foundation team members attended LinkedIn’s event “AI and the Future of Work: The ICT Sector in Transition” and their Empire State Building offices in midtown. This was a chance for the team to meet with more industry professionals and policymakers.

Pedro Ortiz Suarez (Senior Research Scientist, Common Crawl) and Laurie Burchell (Senior Research Engineer, Common Crawl) also attended two further side-events: the first of which took place at Meta’s NYC offices on Friday, where Mary Williamson of Meta presented on the Open Language Data Initiative, which Laurie is co-organising. The second was held at PwC, where Pedro gave a brief general presentation about Common Crawl and data-driven open source software. Pedro and Laurie also met and discussed with additional industry experts and policymakers. These engagements contributed to ongoing discussions around language data, openness, and cross-sector collaboration.

AI Unconference, IBM One Madison

Our main event was the AI Unconference, part of the official UN Open Source Week side-events, which Common Crawl co-hosted with our friends at IBM, the AI Alliance and BrightQuery. One attendee described it as ‘the most impactful AI event of the year’.

The event brought together over 100 attendees from around the world, including leading technologists and industry pioneers.  Highlights included talks from Rich Skrenta (Executive Director, Common Crawl), Jose Plehn-Dujowich (CEO, BrightQuery), Andrea Greco (Research Business Partnerships, IBM), Dean Wampler (Chief Technical Representative to the AI Alliance, IBM), and Thom Vaughan (Principal Technologist, Common Crawl).

Rich Skrenta opened the event with a welcome from Common Crawl, followed by an introduction to the AI Alliance by Andrea Greco, an introduction to BrightQuery from Jose Plehn-Dujowich, an introduction to the AI Alliance’s Open Trusted Data Initiative by Dean Wampler, and a detailed presentation by Thom Vaughan on Common Crawl’s mission. Roberto di Cosmo (Director, Software Heritage) also gave a presentation on their efforts in ethical data collection operations.

Andrea Greco introducing the AI Alliance at the AI Unconference at IBM One Madison.

Andrea Greco introducing the AI Alliance at the AI Unconference at IBM One Madison.

Dean Wampler introducing the Open Trusted Data Initiative at the AI Unconference at IBM One Madison.

Dean Wampler introducing the Open Trusted Data Initiative at the AI Unconference at IBM One Madison.

Thom Vaughan presenting on Common Crawl’s mission and accomplishments at the AI Unconference at IBM One Madison. Yes, we had over an exabyte downloaded from our S3 bucket in 2024!

Thom Vaughan presenting on Common Crawl’s mission and accomplishments at the AI Unconference at IBM One Madison. Yes, we had over an exabyte downloaded from our S3 bucket in 2024!

This was followed by a dynamic and well-received panel discussion, featuring Jose Plehn-Dujowich, Dean Wampler, Lilith Bat-Leah (DMLR Working Group Co-chair, MLCommons), Dave Buckley (Senior Policy Manager, OpenMined), Greg Lindahl (CTO, Common Crawl), and Roberto di Cosmo. Thom Vaughan served as moderator.

The issues around transparency and accountability in AI discussed by the panel are critical to the industry. A repeated term was “full chain of transparency” (thanks, Dave Buckley!) which it was broadly agreed is desperately needed across the industry. Another key theme was that attribution and provenance in training data should be systematic; embedded into data practices by design, rather than as an afterthought. Several panellists also highlighted the developers’ responsibility to uphold ethical standards, with repeated reference to Croissant, the community-developed metadata standard from MLCommons, as a promising tool for responsible data documentation.

Left-to-right: Thom Vaughan, Dave Buckley, Greg Lindahl, Jose Plehn-Dujowich, Lilith Bat-Leah, and Dean Wampler on the panel discussion at the AI Unconference at IBM One Madison.

Left-to-right: Thom Vaughan, Dave Buckley, Greg Lindahl, Jose Plehn-Dujowich, Lilith Bat-Leah, and Dean Wampler on the panel discussion at the AI Unconference at IBM One Madison.

As Dean Wampler noted during the discussions, “Constraints liberate, liberties constrain”, a saying at IBM that resonated with the panel. In the context of AI, the idea points to how well-designed boundaries like clear data documentation standards, transparent governance structures, and ethical constraints can enable greater innovation, trust, and collaboration, rather than limit progress.

Breakout sessions followed the panel, discussing the ethics of large scale data collection, preserving authenticity in user preference signals, governance and transparency in AI training data usage, collaborative standards across the AI ecosystem, and building trust in public data pipelines.

We would like to thank our friends at the AI Alliance, BrightQuery, and IBM for co-hosting this special event with us. Thanks in particular to Tim Bonnemann, Community Lead at IBM, for his tireless efforts and thoughtful coordination throughout the event.

It was a full and productive week in New York.  We had meaningful conversations, made valued connections, and saw real interest in the work we’re doing at Common Crawl.

Our thanks to everyone who took part.  We’re looking forward to what comes next.