惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

T
Threat Research - Cisco Blogs
NISL@THU
NISL@THU
A
Arctic Wolf
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
S
Schneier on Security
T
Tenable Blog
I
Intezer
S
Securelist
Scott Helme
Scott Helme
V
Visual Studio Blog
Simon Willison's Weblog
Simon Willison's Weblog
Google DeepMind News
Google DeepMind News
T
The Blog of Author Tim Ferriss
D
Darknet – Hacking Tools, Hacker News & Cyber Security
AWS News Blog
AWS News Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
MongoDB | Blog
MongoDB | Blog
L
LangChain Blog
F
Fortinet All Blogs
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
C
Cisco Blogs
L
Lohrmann on Cybersecurity
M
MIT News - Artificial intelligence
G
GRAHAM CLULEY
博客园 - 司徒正美
aimingoo的专栏
aimingoo的专栏
雷峰网
雷峰网
MyScale Blog
MyScale Blog
D
DataBreaches.Net
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Recent Announcements
Recent Announcements
C
CXSECURITY Database RSS Feed - CXSecurity.com
量子位
博客园 - 三生石上(FineUI控件)
P
Proofpoint News Feed
Blog — PlanetScale
Blog — PlanetScale
云风的 BLOG
云风的 BLOG
大猫的无限游戏
大猫的无限游戏
GbyAI
GbyAI
Cisco Talos Blog
Cisco Talos Blog
Security Latest
Security Latest
Project Zero
Project Zero
K
Kaspersky official blog
罗磊的独立博客
Know Your Adversary
Know Your Adversary
T
The Exploit Database - CXSecurity.com
P
Privacy International News Feed
P
Privacy & Cybersecurity Law Blog
C
Cyber Attacks, Cyber Crime and Cyber Security
Apple Machine Learning Research
Apple Machine Learning Research

Common Crawl

Common Crawl - Blog - Introducing the AI Visibility Audit Common Crawl - Blog - Host- and Domain-Level Web Graphs March, April, and May 2026 Common Crawl - Blog - May 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Crawl Archive Now Available in a Hugging Face Storage Bucket Common Crawl - Blog - You can now build directly on Common Crawl from the browser Common Crawl - Blog - Host- and Domain-Level Web Graphs February, March, and April 2026 Common Crawl - Blog - April 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Common Crawl Newsletter Common Crawl - Blog - Announcing a Change to Common Crawl Dataset Size Reporting Common Crawl - Blog - Host- and Domain-Level Web Graphs January, February, and March 2026 Common Crawl - Blog - March 2026 Crawl Archive Now Available Common Crawl - Blog - IPv6 Adoption Across the Top 100K Web Hosts Common Crawl - Blog - Web Graph Statistics Gets a Proper Upgrade Common Crawl - Blog - Measuring Web Accessibility from Crawl Archives Common Crawl - Blog - Announcing the Whirlwind Tour of Common Crawl's Datasets Using Java Common Crawl - Blog - Host- and Domain-Level Web Graphs December 2025 and January/February 2026 Common Crawl - Blog - Introducing the New Examples & Resources Browser Common Crawl - Blog - February 2026 Crawl Archive Now Available Common Crawl - Blog - AI Plumbers at FOSDEM’26 Common Crawl - Blog - CC-Citations: A Visualization of Research Papers Referencing Common Crawl Common Crawl - Blog - CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data Common Crawl - Blog - Host- and Domain-Level Web Graphs November/December 2025 and January 2026 Common Crawl - Blog - January 2026 Crawl Archive Now Available Common Crawl - Blog - Web Archives for Social Sciences Datathon, Bristol Common Crawl - Blog - How SEOs Are Using Common Crawl's Web Graph Data for AI Ranking Signals Common Crawl - Blog - GneissWeb Annotations Examples Common Crawl - Blog - Common Crawl at the Mozilla Festival 2025 Common Crawl - Blog - Host- and Domain-Level Web Graphs October, November, December 2025 Common Crawl - Blog - December 2025 Crawl Archive Now Available Common Crawl - Blog - A Sampling of 2025 Research Referencing Common Crawl Common Crawl - Blog - Host- and Domain-Level Web Graphs September, October, and November 2025 Common Crawl - Blog - November 2025 Crawl Archive Now Available Common Crawl - Blog - Common Crawl Celebrates World Digital Preservation Day Common Crawl - Blog - Setting the Record Straight: Common Crawl’s Commitment to Transparency, Fair Use, and the Public Good Common Crawl - Blog - October/November 2025 Newsletter Common Crawl - Blog - Common Crawl Foundation at Stanford HAI Common Crawl - Blog - Host- and Domain-Level Web Graphs August, September, and October 2025 Common Crawl - Blog - October 2025 Crawl Archive Now Available Common Crawl - Blog - Common Crawl Foundation at COLM 2025 Common Crawl - Blog - Announcing GneissWeb Annotations Common Crawl - Blog - Web Languages Needing Review by Native Speakers Common Crawl - Blog - Host- and Domain-Level Web Graphs July, August, and September 2025 Common Crawl - Blog - From SEO to AIO: Why Your Content Needs to Exist in AI Training Data Common Crawl - Blog - September 2025 Crawl Archive Now Available Common Crawl - Blog - Common Crawl Foundation Opt-Out Registry Common Crawl - Blog - Trip Report: AI_dev (Linux Foundation) August 2025 Common Crawl - Blog - July/August 2025 Newsletter Common Crawl - Blog - Host- and Domain-Level Web Graphs June, July, and August 2025 Common Crawl - Blog - August 2025 Crawl Archive Now Available Common Crawl - Blog - Common Crawl Foundation at ACL 2025 Common Crawl - Blog - AI Optimization Is Here: Are You Ready for Search 2.0? Common Crawl - Blog - IETF 123 Report Common Crawl - Blog - Host- and Domain-Level Web Graphs May, June, and July 2025 Common Crawl - Blog - July 2025 Crawl Archive Now Available Common Crawl - Blog - WMDQS Shared Task on Language Identification Common Crawl - Blog - The First WMDQS-Masakhane LangID Hackathon Common Crawl - Blog - Host- and Domain-Level Web Graphs April, May, and June 2025 Common Crawl - Blog - Common Crawl at the United Nations Open Source Week, June 2025 Common Crawl - Blog - June 2025 Crawl Archive Now Available Common Crawl - Blog - May/June 2025 Newsletter Common Crawl - Blog - Announcing the Whirlwind Tour of Common Crawl's Datasets using Python Common Crawl - Blog - Host- and Domain-Level Web Graphs March, April, and May 2025 Common Crawl - Blog - May 2025 Crawl Archive Now Available Common Crawl - Blog - Announcing the First Workshop on Multilingual Data Quality Signals Common Crawl - Blog - Host- and Domain-Level Web Graphs February, March, and April 2025 Common Crawl - Blog - April 2025 Crawl Archive Now Available Common Crawl - Blog - Introducing the Host Index Common Crawl - Blog - IIPC General Assembly & Web Archiving Conference 2025 Common Crawl - Blog - March/April 2025 Newsletter Common Crawl - Blog - Providing Authenticity & Data Provenance for Common Crawl Using Blockchain: Our Work with Constellation Network Common Crawl - Blog - Host- and Domain-Level Web Graphs January, February, and March 2025 Common Crawl - Blog - March 2025 Crawl Archive Now Available Common Crawl - Blog - Introducing Common Crawl AI Agent by ReadyAI Common Crawl - Blog - Submission to the UK’s Copyright and AI Consultation Common Crawl - Blog - Host- and Domain-Level Web Graphs December 2024 and January/February 2025 Common Crawl - Blog - February 2025 Crawl Archive Now Available Common Crawl - Blog - Opening the Gates to Online Safety Common Crawl - Blog - January/February 2025 Newsletter Common Crawl - Blog - Host- and Domain-Level Web Graphs November/December 2024 and January 2025 Common Crawl - Blog - January 2025 Crawl Archive Now Available Common Crawl - Blog - Introducing cc-downloader Common Crawl - Blog - Host- and Domain-Level Web Graphs October, November, and December 2024 Common Crawl - Blog - December 2024 Crawl Archive Now Available Common Crawl - Blog - Common Crawl Foundation at NeurIPS 2024: Expanding Horizons and Building Connections Common Crawl - Blog - Expanding the Language and Cultural Coverage of Common Crawl Common Crawl - Blog - October/November 2024 Newsletter Common Crawl - Blog - Host- and Domain-Level Web Graphs September, October, November 2024 Common Crawl - Blog - November 2024 Crawl Archive Now Available Common Crawl - Blog - Reflections on Recent Talks at the Turing Institute and UCL Common Crawl - Blog - Introducing the Common Crawl Errata Page for Data Transparency Common Crawl - Blog - Host- and Domain-Level Web Graphs August, September, and October 2024 Common Crawl - Blog - October 2024 Crawl Archive Now Available Common Crawl - Blog - White House Briefing on Open Data’s Role in Technology Common Crawl - Blog - IAB Workshop on AI-CONTROL Common Crawl - Blog - Host- and Domain-Level Web Graphs July, August, and September 2024 Common Crawl - Blog - September 2024 Crawl Archive Now Available Common Crawl - Blog - August/September 2024 Newsletter Common Crawl - Blog - Host- and Domain-Level Web Graphs June, July, and August 2024 Common Crawl - Blog - August 2024 Crawl Archive Now Available Common Crawl - Blog - The Increase of Common Crawl Citations in Academic Research
Common Crawl - Blog - Common Crawl Foundation at Stanford HAI: A Shared Legacy of Data and Innovation
2025-09-08 · via Common Crawl

The Stanford Human-Centered AI Institute (HAI), co-founded by Dr. Fei-Fei Li, is a prominent institution that aims to improve the human condition through AI. Li's foundational work with ImageNet, a massive visual database, revolutionized computer vision by demonstrating the importance of large-scale, high-quality data. By providing a common benchmark for researchers, ImageNet helped catalyze the deep learning revolution, proving that a vast amount of curated data could make algorithms more accurate.

This very principle (the power of open, accessible data to drive innovation) is also at the core of the Common Crawl Foundation, a non-profit founded by Gil Elbaz. Both Li and Elbaz share a connection to the California Institute of Technology (Caltech); Li earned her PhD there, while Elbaz is an alumnus with a double major in Engineering & Applied Science and Economics.

Elbaz founded Applied Semantics which was the only company acquired by Google (2003) before their IPO. During his tenure at Google, Elbaz continued to work on the Applied Semantics technology and AdSense. AdSense helped establish Google’s position as a leader in online advertising and has been responsible for a substantial amount of revenue since its launch in 2005.  In 2007, Gil Elbaz founded Common Crawl with the mission to democratize access to web information, providing a petabyte-scale web crawl that is free for public use.

The shared academic foundation and commitment to open data among these two figures highlights a central theme: the collective, open source approach to data is essential for meaningful progress in artificial intelligence and machine learning, as well as in strengthening the ties between language, community, and culture.

This philosophy brings the Common Crawl Foundation to Stanford HAI for their seminar, "Preserving Humanity's Knowledge and Making it Accessible: Addressing Challenges of Public Web Data". The seminar, taking place on October 22, 2025, will focus on crucial topics such as privacy, safety, and security. By presenting insights from a new data product, Common Crawl will advocate for greater transparency and informed solutions for the future of public web data, continuing to build on the legacy of open data that pioneers like Dr. Fei-Fei Li and Gil Elbaz have championed.

Please join us in person or virtually at this seminal event.