惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Google DeepMind News
Google DeepMind News
MongoDB | Blog
MongoDB | Blog
有赞技术团队
有赞技术团队
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
人人都是产品经理
人人都是产品经理
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
B
Blog RSS Feed
T
Tor Project blog
T
Threat Research - Cisco Blogs
Microsoft Azure Blog
Microsoft Azure Blog
M
MIT News - Artificial intelligence
V
Vulnerabilities – Threatpost
Project Zero
Project Zero
C
CXSECURITY Database RSS Feed - CXSecurity.com
The Register - Security
The Register - Security
Latest news
Latest news
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
The Hacker News
The Hacker News
Google DeepMind News
Google DeepMind News
L
LINUX DO - 最新话题
U
Unit 42
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
博客园 - 司徒正美
T
Tenable Blog
H
Hacker News: Front Page
B
Blog
宝玉的分享
宝玉的分享
C
Check Point Blog
美团技术团队
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
C
CERT Recently Published Vulnerability Notes
P
Proofpoint News Feed
The GitHub Blog
The GitHub Blog
G
GRAHAM CLULEY
Google Online Security Blog
Google Online Security Blog
Security Archives - TechRepublic
Security Archives - TechRepublic
P
Proofpoint News Feed
GbyAI
GbyAI
酷 壳 – CoolShell
酷 壳 – CoolShell
Hugging Face - Blog
Hugging Face - Blog
Y
Y Combinator Blog
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Hacker News - Newest:
Hacker News - Newest: "LLM"
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Scott Helme
Scott Helme
L
Lohrmann on Cybersecurity
量子位
A
About on SuperTechFans
V2EX - 技术
V2EX - 技术
T
The Exploit Database - CXSecurity.com

Common Crawl

Common Crawl - Blog - Introducing the AI Visibility Audit Common Crawl - Blog - Host- and Domain-Level Web Graphs March, April, and May 2026 Common Crawl - Blog - May 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Crawl Archive Now Available in a Hugging Face Storage Bucket Common Crawl - Blog - You can now build directly on Common Crawl from the browser Common Crawl - Blog - Host- and Domain-Level Web Graphs February, March, and April 2026 Common Crawl - Blog - April 2026 Crawl Archive Now Available Common Crawl - Blog - April 2026 Common Crawl Newsletter Common Crawl - Blog - Announcing a Change to Common Crawl Dataset Size Reporting Common Crawl - Blog - Host- and Domain-Level Web Graphs January, February, and March 2026 Common Crawl - Blog - March 2026 Crawl Archive Now Available Common Crawl - Blog - IPv6 Adoption Across the Top 100K Web Hosts Common Crawl - Blog - Web Graph Statistics Gets a Proper Upgrade Common Crawl - Blog - Measuring Web Accessibility from Crawl Archives Common Crawl - Blog - Announcing the Whirlwind Tour of Common Crawl's Datasets Using Java Common Crawl - Blog - Host- and Domain-Level Web Graphs December 2025 and January/February 2026 Common Crawl - Blog - Introducing the New Examples & Resources Browser Common Crawl - Blog - February 2026 Crawl Archive Now Available Common Crawl - Blog - AI Plumbers at FOSDEM’26 Common Crawl - Blog - CC-Citations: A Visualization of Research Papers Referencing Common Crawl Common Crawl - Blog - CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data Common Crawl - Blog - Host- and Domain-Level Web Graphs November/December 2025 and January 2026 Common Crawl - Blog - January 2026 Crawl Archive Now Available Common Crawl - Blog - Web Archives for Social Sciences Datathon, Bristol Common Crawl - Blog - How SEOs Are Using Common Crawl's Web Graph Data for AI Ranking Signals Common Crawl - Blog - GneissWeb Annotations Examples Common Crawl - Blog - Common Crawl at the Mozilla Festival 2025 Common Crawl - Blog - Host- and Domain-Level Web Graphs October, November, December 2025 Common Crawl - Blog - December 2025 Crawl Archive Now Available Common Crawl - Blog - A Sampling of 2025 Research Referencing Common Crawl Common Crawl - Blog - Host- and Domain-Level Web Graphs September, October, and November 2025 Common Crawl - Blog - November 2025 Crawl Archive Now Available Common Crawl - Blog - Common Crawl Celebrates World Digital Preservation Day Common Crawl - Blog - Setting the Record Straight: Common Crawl’s Commitment to Transparency, Fair Use, and the Public Good Common Crawl - Blog - October/November 2025 Newsletter Common Crawl - Blog - Common Crawl Foundation at Stanford HAI Common Crawl - Blog - Host- and Domain-Level Web Graphs August, September, and October 2025 Common Crawl - Blog - October 2025 Crawl Archive Now Available Common Crawl - Blog - Common Crawl Foundation at COLM 2025 Common Crawl - Blog - Announcing GneissWeb Annotations Common Crawl - Blog - Web Languages Needing Review by Native Speakers Common Crawl - Blog - Host- and Domain-Level Web Graphs July, August, and September 2025 Common Crawl - Blog - September 2025 Crawl Archive Now Available Common Crawl - Blog - Common Crawl Foundation Opt-Out Registry Common Crawl - Blog - Trip Report: AI_dev (Linux Foundation) August 2025 Common Crawl - Blog - Common Crawl Foundation at Stanford HAI: A Shared Legacy of Data and Innovation Common Crawl - Blog - July/August 2025 Newsletter Common Crawl - Blog - Host- and Domain-Level Web Graphs June, July, and August 2025 Common Crawl - Blog - August 2025 Crawl Archive Now Available Common Crawl - Blog - Common Crawl Foundation at ACL 2025 Common Crawl - Blog - AI Optimization Is Here: Are You Ready for Search 2.0? Common Crawl - Blog - IETF 123 Report Common Crawl - Blog - Host- and Domain-Level Web Graphs May, June, and July 2025 Common Crawl - Blog - July 2025 Crawl Archive Now Available Common Crawl - Blog - WMDQS Shared Task on Language Identification Common Crawl - Blog - The First WMDQS-Masakhane LangID Hackathon Common Crawl - Blog - Host- and Domain-Level Web Graphs April, May, and June 2025 Common Crawl - Blog - Common Crawl at the United Nations Open Source Week, June 2025 Common Crawl - Blog - June 2025 Crawl Archive Now Available Common Crawl - Blog - May/June 2025 Newsletter Common Crawl - Blog - Announcing the Whirlwind Tour of Common Crawl's Datasets using Python Common Crawl - Blog - Host- and Domain-Level Web Graphs March, April, and May 2025 Common Crawl - Blog - May 2025 Crawl Archive Now Available Common Crawl - Blog - Announcing the First Workshop on Multilingual Data Quality Signals Common Crawl - Blog - Host- and Domain-Level Web Graphs February, March, and April 2025 Common Crawl - Blog - April 2025 Crawl Archive Now Available Common Crawl - Blog - Introducing the Host Index Common Crawl - Blog - IIPC General Assembly & Web Archiving Conference 2025 Common Crawl - Blog - March/April 2025 Newsletter Common Crawl - Blog - Providing Authenticity & Data Provenance for Common Crawl Using Blockchain: Our Work with Constellation Network Common Crawl - Blog - Host- and Domain-Level Web Graphs January, February, and March 2025 Common Crawl - Blog - March 2025 Crawl Archive Now Available Common Crawl - Blog - Introducing Common Crawl AI Agent by ReadyAI Common Crawl - Blog - Submission to the UK’s Copyright and AI Consultation Common Crawl - Blog - Host- and Domain-Level Web Graphs December 2024 and January/February 2025 Common Crawl - Blog - February 2025 Crawl Archive Now Available Common Crawl - Blog - Opening the Gates to Online Safety Common Crawl - Blog - January/February 2025 Newsletter Common Crawl - Blog - Host- and Domain-Level Web Graphs November/December 2024 and January 2025 Common Crawl - Blog - January 2025 Crawl Archive Now Available Common Crawl - Blog - Introducing cc-downloader Common Crawl - Blog - Host- and Domain-Level Web Graphs October, November, and December 2024 Common Crawl - Blog - December 2024 Crawl Archive Now Available Common Crawl - Blog - Common Crawl Foundation at NeurIPS 2024: Expanding Horizons and Building Connections Common Crawl - Blog - Expanding the Language and Cultural Coverage of Common Crawl Common Crawl - Blog - October/November 2024 Newsletter Common Crawl - Blog - Host- and Domain-Level Web Graphs September, October, November 2024 Common Crawl - Blog - November 2024 Crawl Archive Now Available Common Crawl - Blog - Reflections on Recent Talks at the Turing Institute and UCL Common Crawl - Blog - Introducing the Common Crawl Errata Page for Data Transparency Common Crawl - Blog - Host- and Domain-Level Web Graphs August, September, and October 2024 Common Crawl - Blog - October 2024 Crawl Archive Now Available Common Crawl - Blog - White House Briefing on Open Data’s Role in Technology Common Crawl - Blog - IAB Workshop on AI-CONTROL Common Crawl - Blog - Host- and Domain-Level Web Graphs July, August, and September 2024 Common Crawl - Blog - September 2024 Crawl Archive Now Available Common Crawl - Blog - August/September 2024 Newsletter Common Crawl - Blog - Host- and Domain-Level Web Graphs June, July, and August 2024 Common Crawl - Blog - August 2024 Crawl Archive Now Available Common Crawl - Blog - The Increase of Common Crawl Citations in Academic Research
Common Crawl - Blog - From SEO to AIO: Why Your Content Needs to Exist in AI Training Data
2025-09-23 · via Common Crawl

I've been working in search for over two decades, from the Open Directory Project days through building search engines at Blekko to my current role leading web intelligence at Common Crawl. But it wasn't until a customer found my motorcycle repair side hustle through ChatGPT that the magnitude of what's happening really hit me.

A few years back, I restored an old BMW and started fixing bikes out of my garage in Redwood City. Being obsessive about local SEO, I made sure the shop ranked well locally. Then something unexpected happened: customers started showing up saying an AI assistant told them about me. They'd asked ChatGPT where to get their motorcycle fixed, and it sent them to my garage.

That's when it clicked: being visible in AI systems isn't just about future-proofing anymore. It's already driving real business today.

The Fundamental Shift in User Behavior

For decades, we optimized for short, typed queries. "Buy shoes." "Hotels Bangkok." "Pizza near me." Two or three words into a search box, ten blue links back. The SEO playbook was straightforward: get indexed, then fight for rankings.

But watch how people interact with ChatGPT, Claude, or Perplexity now. They don't type "hotels Bangkok." They say: "I'm planning a trip to Chiang Mai for November, want something boutique near the old city with a pool, not too touristy, under $200 a night, what are my options?"

That's a twenty-word question, often spoken aloud. Users no longer want ten links to evaluate; they want the single synthesized answer, the filtered recommendation, the complete plan. They expect the system to read, compare, decide, and deliver.

This shift from two typed words to twenty spoken ones represents the most disruptive change search has ever seen. And it's why discovery itself is being re-architected from the ground up.

When Invisibility Becomes Life-Threatening

The stakes couldn't be higher. Children's Hospital of Los Angeles is one of the top pediatric cancer centers in the United States. Yet when parents search within Gemini or ChatGPT for "where should I take my child with leukemia in LA?" they don't see CHLA.

Why? The hospital's site sits behind Cloudflare, whose default settings inject a robots.txt that blocks AI crawlers, including our CCBot at Common Crawl. In the world of AI-powered search, this premier hospital effectively doesn't exist.

This goes beyond lost web traffic. Families may be unable to find potentially life-saving care for their children, not because the hospital deliberately opted out of AI systems, but because a default setting at the hospital’s content delivery provider blocked AI crawlers without the hospital realising it.

Understanding the New Discovery Pipeline

To grasp why this matters, you need to understand how large language models actually work. LLMs aren't real-time systems. They're trained on static snapshots of the web, a process that takes weeks or months. Once training concludes, the model's knowledge is frozen. That's why ChatGPT will tell you "my knowledge cutoff is April 2023" or similar.

Retrieval-augmented generation (RAG) bridges this gap. When you ask a question, these systems can fetch fresh web pages and blend that current information with the model's frozen knowledge base. That's how Perplexity can tell you about yesterday's news despite its underlying model knowing nothing about it.

Here's the critical nuance: Technically, you don't need to be in the training data to appear in RAG results. If your page is crawlable and indexed by whatever live source the system uses, you could theoretically be pulled in.

But practically if your brand, product, or entity is absent from the training data, the model doesn't know you exist. It won't expand queries to include you. It won't recognize you as relevant. Your retrieval chances plummet.

Being in the training corpus doesn't guarantee retrieval, but being absent from it dramatically lowers your odds of ever being surfaced. Visibility at the crawl layer has become as strategically vital as backlinks once were.

The Opt-Out Crisis We're Witnessing

At Common Crawl, we're experiencing this shift firsthand. After the New York Times blocked AI crawlers, we saw a wave of copycats. Publishers of all sizes, some polite, others threatening legal action, demanded removal from our dataset.

We created an Opt-Out Registry. When a site requests exclusion or threatens us legally, we don't just remove them from Common Crawl. We flag them for the entire ecosystem: OpenAI, Meta, Amazon, researchers, everyone.

For these publishers, exclusion isn't temporary. It's essentially permanent.

Here's the irony: AI models don't actually need these sites to answer user questions. The knowledge surfaces anyway through reviews, forums, citations, and other user-generated content. When a brand opts out, the conversation about them continues. What disappears is their authoritative voice in that conversation.

The Language Divide

Another critical factor is language representation. The vast majority of training data is English. Smaller languages like Estonian, Catalan, even Thai, are massively underrepresented in these models.

For businesses operating in smaller language markets, this is existential. Publish only in Catalan, and your content may be invisible in AI-driven answers because the model lacks sufficient Catalan material to generalize from.

The practical strategy? Publish in English alongside your local language. English acts as the gateway into the model. You're not abandoning your local audience; you're ensuring legibility to the systems that increasingly serve as the first point of discovery.

When Infrastructure Becomes a Ranking Factor

The infrastructure constraints are staggering. Training GPT-4 cost between $78-100 million in computation costs alone, with total costs reaching hundreds of millions when including hardware. The energy footprint is enormous: some training runs burn more electricity than a large hydroelectric dam generates per minute.

These constraints directly shape what gets crawled and processed. Microsoft is reviving Three Mile Island for AI power. Google is investing in small modular nuclear reactors. When computation is this expensive, crawlers must prioritize what they consider high-value content.

Infrastructure is no longer invisible; it's effectively become a ranking factor. Not every site gets equal treatment. If you're deemed lower priority, you might not make it into the next training run.

The New Visibility Funnel

Think of the new discovery funnel this way:

Training Data → LLM → RAG → Your Site → Conversion

At the top, if you're absent from training data, you're missing critical baseline awareness. In the middle, if you're blocked at the crawl layer, you're excluded entirely. At the bottom, if you're invisible to retrieval systems, you don't get surfaced, summarized, or converted.

This is the new reality of online visibility.

What You Need to Do Now

First, audit your crawl accessibility immediately. Check whether your site is being blocked at the CDN level. If you're on Cloudflare, verify that AI crawlers are actually allowed. Don't assume; check your logs. If you don't see CCBot, GPTBot, or ClaudeBot, you may already be invisible.

Second, publish in English and your local language. English remains the gateway language into these models.

Third, syndicate and distribute your content. Don't rely on a single domain. Just as backlinks once created resilience, syndication today increases the odds that your content survives preprocessing and makes it into training datasets.

Fourth, monitor the evolving landscape. Defaults change. New players emerge. Not everyone follows the stated rules.

Finally, educate your stakeholders. Many executives still think SEO is about title tags and keyword density. They don't realize their site may already be invisible in the fastest-growing discovery systems on earth.

The Bottom Line

SEO has always been about visibility. That hasn't changed. What's changed is the mechanism.

The old world was about index and rank. The new world is about training and retrieval.

Training data has become the new link graph. The strategic asset isn't your PageRank anymore; it's your presence in the corpus.

If you're not in the crawl, you're not in the model, and if you're not in the model, you may not be in the market.

The choice is yours, but make it consciously. Because right now decisions about your AI visibility might be getting made by your CDN, your legal department, or your trade association without you even knowing it.

Stephen Burns is the Web Intelligence Lead at Common Crawl Foundation, the nonprofit that provides open web data to AI researchers and companies worldwide. He also works in enterprise SEO at U.S. Bank and operates a motorcycle repair shop that customers keep finding through ChatGPT.