惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

H
Help Net Security
C
Cybersecurity and Infrastructure Security Agency CISA
S
Secure Thoughts
Application and Cybersecurity Blog
Application and Cybersecurity Blog
Hacker News: Ask HN
Hacker News: Ask HN
Attack and Defense Labs
Attack and Defense Labs
N
News and Events Feed by Topic
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
A
Arctic Wolf
www.infosecurity-magazine.com
www.infosecurity-magazine.com
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
The GitHub Blog
The GitHub Blog
W
WeLiveSecurity
Simon Willison's Weblog
Simon Willison's Weblog
WordPress大学
WordPress大学
Y
Y Combinator Blog
Recent Commits to openclaw:main
Recent Commits to openclaw:main
Forbes - Security
Forbes - Security
NISL@THU
NISL@THU
博客园 - 聂微东
G
Google Developers Blog
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Schneier on Security
Schneier on Security
V
Vulnerabilities – Threatpost
V
V2EX
I
Intezer
S
Schneier on Security
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
雷峰网
雷峰网
L
LangChain Blog
小众软件
小众软件
Hugging Face - Blog
Hugging Face - Blog
T
The Blog of Author Tim Ferriss
L
LINUX DO - 热门话题
P
Proofpoint News Feed
Microsoft Security Blog
Microsoft Security Blog
Project Zero
Project Zero
V
Visual Studio Blog
Engineering at Meta
Engineering at Meta
爱范儿
爱范儿
H
Hacker News: Front Page
B
Blog
T
Threatpost
Spread Privacy
Spread Privacy
H
Heimdal Security Blog
F
Fortinet All Blogs
TaoSecurity Blog
TaoSecurity Blog
T
Threat Research - Cisco Blogs
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
The Cloudflare Blog

Common Crawl

Common Crawl - Blog - Introducing the AI Visibility Audit Common Crawl - Blog - Host- and Domain-Level Web Graphs March, April, and May 2026 Common Crawl - Blog - May 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Crawl Archive Now Available in a Hugging Face Storage Bucket Common Crawl - Blog - You can now build directly on Common Crawl from the browser Common Crawl - Blog - Host- and Domain-Level Web Graphs February, March, and April 2026 Common Crawl - Blog - April 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Common Crawl Newsletter Common Crawl - Blog - Announcing a Change to Common Crawl Dataset Size Reporting Common Crawl - Blog - Host- and Domain-Level Web Graphs January, February, and March 2026 Common Crawl - Blog - March 2026 Crawl Archive Now Available Common Crawl - Blog - IPv6 Adoption Across the Top 100K Web Hosts Common Crawl - Blog - Web Graph Statistics Gets a Proper Upgrade Common Crawl - Blog - Measuring Web Accessibility from Crawl Archives Common Crawl - Blog - Announcing the Whirlwind Tour of Common Crawl's Datasets Using Java Common Crawl - Blog - Host- and Domain-Level Web Graphs December 2025 and January/February 2026 Common Crawl - Blog - Introducing the New Examples & Resources Browser Common Crawl - Blog - February 2026 Crawl Archive Now Available Common Crawl - Blog - AI Plumbers at FOSDEM’26 Common Crawl - Blog - CC-Citations: A Visualization of Research Papers Referencing Common Crawl Common Crawl - Blog - CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data Common Crawl - Blog - Host- and Domain-Level Web Graphs November/December 2025 and January 2026 Common Crawl - Blog - January 2026 Crawl Archive Now Available Common Crawl - Blog - Web Archives for Social Sciences Datathon, Bristol Common Crawl - Blog - How SEOs Are Using Common Crawl's Web Graph Data for AI Ranking Signals Common Crawl - Blog - GneissWeb Annotations Examples Common Crawl - Blog - Common Crawl at the Mozilla Festival 2025 Common Crawl - Blog - Host- and Domain-Level Web Graphs October, November, December 2025 Common Crawl - Blog - December 2025 Crawl Archive Now Available Common Crawl - Blog - A Sampling of 2025 Research Referencing Common Crawl Common Crawl - Blog - Host- and Domain-Level Web Graphs September, October, and November 2025 Common Crawl - Blog - November 2025 Crawl Archive Now Available Common Crawl - Blog - Common Crawl Celebrates World Digital Preservation Day Common Crawl - Blog - Setting the Record Straight: Common Crawl’s Commitment to Transparency, Fair Use, and the Public Good Common Crawl - Blog - October/November 2025 Newsletter Common Crawl - Blog - Common Crawl Foundation at Stanford HAI Common Crawl - Blog - Host- and Domain-Level Web Graphs August, September, and October 2025 Common Crawl - Blog - October 2025 Crawl Archive Now Available Common Crawl - Blog - Common Crawl Foundation at COLM 2025 Common Crawl - Blog - Announcing GneissWeb Annotations Common Crawl - Blog - Web Languages Needing Review by Native Speakers Common Crawl - Blog - Host- and Domain-Level Web Graphs July, August, and September 2025 Common Crawl - Blog - From SEO to AIO: Why Your Content Needs to Exist in AI Training Data Common Crawl - Blog - September 2025 Crawl Archive Now Available Common Crawl - Blog - Common Crawl Foundation Opt-Out Registry Common Crawl - Blog - Trip Report: AI_dev (Linux Foundation) August 2025 Common Crawl - Blog - Common Crawl Foundation at Stanford HAI: A Shared Legacy of Data and Innovation Common Crawl - Blog - July/August 2025 Newsletter Common Crawl - Blog - Host- and Domain-Level Web Graphs June, July, and August 2025 Common Crawl - Blog - August 2025 Crawl Archive Now Available Common Crawl - Blog - Common Crawl Foundation at ACL 2025 Common Crawl - Blog - AI Optimization Is Here: Are You Ready for Search 2.0? Common Crawl - Blog - IETF 123 Report Common Crawl - Blog - Host- and Domain-Level Web Graphs May, June, and July 2025 Common Crawl - Blog - July 2025 Crawl Archive Now Available Common Crawl - Blog - WMDQS Shared Task on Language Identification Common Crawl - Blog - The First WMDQS-Masakhane LangID Hackathon Common Crawl - Blog - Host- and Domain-Level Web Graphs April, May, and June 2025 Common Crawl - Blog - Common Crawl at the United Nations Open Source Week, June 2025 Common Crawl - Blog - June 2025 Crawl Archive Now Available Common Crawl - Blog - May/June 2025 Newsletter Common Crawl - Blog - Announcing the Whirlwind Tour of Common Crawl's Datasets using Python Common Crawl - Blog - Host- and Domain-Level Web Graphs March, April, and May 2025 Common Crawl - Blog - May 2025 Crawl Archive Now Available Common Crawl - Blog - Announcing the First Workshop on Multilingual Data Quality Signals Common Crawl - Blog - Host- and Domain-Level Web Graphs February, March, and April 2025 Common Crawl - Blog - April 2025 Crawl Archive Now Available Common Crawl - Blog - Introducing the Host Index Common Crawl - Blog - IIPC General Assembly & Web Archiving Conference 2025 Common Crawl - Blog - March/April 2025 Newsletter Common Crawl - Blog - Providing Authenticity & Data Provenance for Common Crawl Using Blockchain: Our Work with Constellation Network Common Crawl - Blog - Host- and Domain-Level Web Graphs January, February, and March 2025 Common Crawl - Blog - March 2025 Crawl Archive Now Available Common Crawl - Blog - Introducing Common Crawl AI Agent by ReadyAI Common Crawl - Blog - Submission to the UK’s Copyright and AI Consultation Common Crawl - Blog - Host- and Domain-Level Web Graphs December 2024 and January/February 2025 Common Crawl - Blog - February 2025 Crawl Archive Now Available Common Crawl - Blog - Opening the Gates to Online Safety Common Crawl - Blog - January/February 2025 Newsletter Common Crawl - Blog - Host- and Domain-Level Web Graphs November/December 2024 and January 2025 Common Crawl - Blog - January 2025 Crawl Archive Now Available Common Crawl - Blog - Introducing cc-downloader Common Crawl - Blog - Host- and Domain-Level Web Graphs October, November, and December 2024 Common Crawl - Blog - December 2024 Crawl Archive Now Available Common Crawl - Blog - Common Crawl Foundation at NeurIPS 2024: Expanding Horizons and Building Connections Common Crawl - Blog - Expanding the Language and Cultural Coverage of Common Crawl Common Crawl - Blog - October/November 2024 Newsletter Common Crawl - Blog - Host- and Domain-Level Web Graphs September, October, November 2024 Common Crawl - Blog - November 2024 Crawl Archive Now Available Common Crawl - Blog - Reflections on Recent Talks at the Turing Institute and UCL Common Crawl - Blog - Introducing the Common Crawl Errata Page for Data Transparency Common Crawl - Blog - Host- and Domain-Level Web Graphs August, September, and October 2024 Common Crawl - Blog - October 2024 Crawl Archive Now Available Common Crawl - Blog - White House Briefing on Open Data’s Role in Technology Common Crawl - Blog - Host- and Domain-Level Web Graphs July, August, and September 2024 Common Crawl - Blog - September 2024 Crawl Archive Now Available Common Crawl - Blog - August/September 2024 Newsletter Common Crawl - Blog - Host- and Domain-Level Web Graphs June, July, and August 2024 Common Crawl - Blog - August 2024 Crawl Archive Now Available Common Crawl - Blog - The Increase of Common Crawl Citations in Academic Research
Common Crawl - Blog - IAB Workshop on AI-CONTROL
2024-09-30 · via Common Crawl

Earlier this month, the Common Crawl Foundation had the privilege of participating in a groundbreaking workshop hosted by the Internet Architecture Board (IAB) in Washington DC. The workshop, titled "IAB Workshop on AI-CONTROL," brought together experts from crawling companies, web publishers, AI companies, and “bot defense” companies to discuss the intersection of artificial intelligence and Internet protocols. Thom Vaughan from Common Crawl served on the program committee.

Left to right: Thom Vaughan, Paul Ohm, Carl Gahnberg, Jari Arkko, Farzaneh Badii (photo permission graciously approved)

Key Topics of Discussion

While adhering to Chatham House rules limits the specifics we can share, we can highlight some of the general themes that were explored:

Opt-out and Opt-in Vocabulary

The workshop attendees discussed various approaches for allowing individuals and organizations to opt-out of AI data collection and processing. We agreed that it was important to develop a vocabulary that clearly expressed the preferences of rights holders and authors. This vocabulary should include Creative Commons-style opt-in choices. This vocabulary will be important for the eventual EU text and data mining (TDM) opt-out registry.

The Rise of Bot Defenses

Attendees discussed the recent popularity of using “robot defenses” to stop crawling, instead of robots.txt.

Providers of these defenses are sometimes treating archive crawlers (like Common Crawl’s CCBot) the same as bots crawling for particular AI companies. This is unfortunate and opaque: robots.txt is usually a public document, and is obeyed by a significant number of crawlers. We discussed some examples of US government websites inadvertently blocking the official 2024 End of Term Archive.

Stakeholder Concerns

A wide range of perspectives was shared, and we discussed the balancing act between innovation, privacy, and ethical considerations.

The Future

The discussions at this workshop will undoubtedly influence future recommendations and standards in the realm of Generative AI and Internet protocols. As we enter what many are calling the "dawn of generative AI," the guidance provided by organizations like the IAB will be instrumental in shaping a responsible and innovative future.

While we can't share specific details of the presentations or discussions, we can say that the level of expertise and the depth of conversation were extensive.  Our organization was well-represented, with our CTO Greg Lindahl contributing valuable insights to the discussions. A conference report will soon be published on the IETF datatracker.

Conclusion

It's clear that the intersection of AI and internet protocols will remain a critical area of focus.  Workshops like this one play a vital role in cultivating collaboration and developing thoughtful approaches to emerging challenges.

We look forward to seeing how the ideas exchanged at this workshop will shape future guidelines and best practices in the field. If you want to continue the discussion you're more than welcome to join us in our Discord Server, or our Google Group.

Note: This blog post adheres to Chatham House rules. No specific statements or opinions have been attributed to individual participants.

Left to right: Thom Vaughan, Washington Monument, Greg Lindahl