惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

L
LangChain Blog
N
Netflix TechBlog - Medium
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
V
V2EX
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Blog — PlanetScale
Blog — PlanetScale
Microsoft Security Blog
Microsoft Security Blog
D
Docker
WordPress大学
WordPress大学
罗磊的独立博客
J
Java Code Geeks
博客园 - 【当耐特】
博客园 - 司徒正美
雷峰网
雷峰网
H
Help Net Security
酷 壳 – CoolShell
酷 壳 – CoolShell
Last Week in AI
Last Week in AI
宝玉的分享
宝玉的分享
Martin Fowler
Martin Fowler
T
Tailwind CSS Blog
Google DeepMind News
Google DeepMind News
M
MIT News - Artificial intelligence
Recent Announcements
Recent Announcements
B
Blog

Common Crawl

Common Crawl - Blog - Introducing the AI Visibility Audit Common Crawl - Blog - Host- and Domain-Level Web Graphs March, April, and May 2026 Common Crawl - Blog - May 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Crawl Archive Now Available in a Hugging Face Storage Bucket Common Crawl - Blog - You can now build directly on Common Crawl from the browser Common Crawl - Blog - Host- and Domain-Level Web Graphs February, March, and April 2026 Common Crawl - Blog - April 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Common Crawl Newsletter Common Crawl - Blog - Announcing a Change to Common Crawl Dataset Size Reporting Common Crawl - Blog - Host- and Domain-Level Web Graphs January, February, and March 2026 Common Crawl - Blog - March 2026 Crawl Archive Now Available Common Crawl - Blog - IPv6 Adoption Across the Top 100K Web Hosts Common Crawl - Blog - Web Graph Statistics Gets a Proper Upgrade Common Crawl - Blog - Measuring Web Accessibility from Crawl Archives Common Crawl - Blog - Announcing the Whirlwind Tour of Common Crawl's Datasets Using Java Common Crawl - Blog - Host- and Domain-Level Web Graphs December 2025 and January/February 2026 Common Crawl - Blog - Introducing the New Examples & Resources Browser Common Crawl - Blog - February 2026 Crawl Archive Now Available Common Crawl - Blog - AI Plumbers at FOSDEM’26 Common Crawl - Blog - CC-Citations: A Visualization of Research Papers Referencing Common Crawl Common Crawl - Blog - CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data Common Crawl - Blog - Host- and Domain-Level Web Graphs November/December 2025 and January 2026 Common Crawl - Blog - January 2026 Crawl Archive Now Available Common Crawl - Blog - Web Archives for Social Sciences Datathon, Bristol Common Crawl - Blog - How SEOs Are Using Common Crawl's Web Graph Data for AI Ranking Signals Common Crawl - Blog - GneissWeb Annotations Examples Common Crawl - Blog - Common Crawl at the Mozilla Festival 2025 Common Crawl - Blog - Host- and Domain-Level Web Graphs October, November, December 2025 Common Crawl - Blog - December 2025 Crawl Archive Now Available Common Crawl - Blog - A Sampling of 2025 Research Referencing Common Crawl
Common Crawl - Blog - White House Briefing on Open Data’s...
2024-10-08 · via Common Crawl

We recently had the honor of briefing the White House Office of Science and Technology Policy (OSTP) on the role of The Common Crawl Foundation as critical infrastructure in the artificial intelligence ecosystem and how we can support U.S. federal efforts in advancing responsible AI use and research.

We were invited by Travis Hoppe, Assistant Director for AI Research and Development for the OSTP, who hosted the meeting.  Rich Skrenta, Executive Director of the Common Crawl Foundation led the briefing, accompanied by Hugh Marbury and Chris Tolles from our advisory board.  Other attendees both in person and online included representatives from the OSTP, the U.S. Department of Commerce, and the White House Office of Management and Budget.

Hugh Marbury, Rich Skrenta, and Chris Tolles photographed at the White House, Washington DC

Left to right: Hugh Marbury, Rich Skrenta, Chris Tolles. Photo credit: Travis Hoppe

Before the briefing, we attended a roundtable discussion titled "Democratizing Government Data with Gen AI" organized by the Kapor Foundation, the Omidyar Network, and the nonprofit Center for Open Data Enterprise (CODE).  There, we connected with featured presenter Oliver Wise, the Chief Data Officer at the U.S. Department of Commerce, who facilitated the chain of introductions leading to our briefing.  

Presenting our work and mission to leaders dedicated to public service at the historic Old Executive Office Building was a significant opportunity, and the follow-up from this briefing has been highly productive and encouraging.  As discussions turned to the role that Common Crawl can play as a responsible actor in the open data space, we were a signatory to an announcement from the White House on September 12, 2024, regarding voluntary private sector commitments to responsibly source their datasets and safeguard them from image-based sexual abuse.

Additionally, we have had follow-up meetings graciously facilitated by the OSTP with other interested parties inside and outside the federal government.  We were especially excited to meet with executive management leading the National Science Foundation’s National Artificial Intelligence Research Resource (NAIRR) pilot project to discuss the vision for a shared national research infrastructure for responsible discovery and innovation in AI.

We are looking forward to more involvement with various agencies within the federal government to explore how the Common Crawl Foundation can help promote artificial intelligence for the benefit of our nation and the world.  We would like to thank Travis Hoppe for kicking-off these exciting collaborations.