惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Jina AI
Jina AI
S
SegmentFault 最新的问题
D
DataBreaches.Net
H
Help Net Security
有赞技术团队
有赞技术团队
M
MIT News - Artificial intelligence
Martin Fowler
Martin Fowler
IT之家
IT之家
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
腾讯CDC
罗磊的独立博客
Y
Y Combinator Blog
阮一峰的网络日志
阮一峰的网络日志
云风的 BLOG
云风的 BLOG
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
WordPress大学
WordPress大学
Microsoft Security Blog
Microsoft Security Blog
J
Java Code Geeks
Vercel News
Vercel News
Hugging Face - Blog
Hugging Face - Blog
aimingoo的专栏
aimingoo的专栏
Stack Overflow Blog
Stack Overflow Blog
Recent Announcements
Recent Announcements
博客园 - 三生石上(FineUI控件)

Enterprise – Silicon Republic

Recovery readiness a missing link in cyber resilience, finds report EU finally gets its hands on Anthropic’s Mythos New Irish dispute body to tackle illegal online content launched Rhysida leaks 5.7TB of sensitive Berlin state data in major hack ‘Keeping OT security up to date is more than patching systems’ HSE fined €645,000 for storing records in decrepit conditions Boston Scientific cyberattack: Cork staff asked to work remotely Report: Only sectors aiding AI adoption sure to grow from it Why connectivity and cybersecurity can't be treated separately French authorities investigating hack of national tax authority Why employee retention is at the heart of the cyber skills gap Levi Strauss corporate data stolen in cyberattack Levi Strauss corporate data stolen in cyberattack How might a system ‘leak secrets’ without being hacked? What is a side-channel attack in cybersecurity? OpenAI agents breach Modal client system after Hugging Face hack OpenAI agents breach Modal client system after Hugging Face hack Report: EMEA businesses not reaping benefits from their AI spend Report: EMEA businesses not reaping benefits from their AI spend Transport for London hackers jailed for five and a half years Transport for London hackers jailed for five and a half years Commission refers Ireland to CJEU for failing to enact cyber rules Commission refers Ireland to CJEU for failing to enact cyber rules Cloudflare to block AI crawlers from ad-supported webpages by default Google ordered to pay Klarna nearly $2bn in abuse-of-power row Google ordered to pay Klarna nearly $2bn in abuse-of-power row Upcoming iPhone 18 model leaked in Tata Electronics hack New iPhone 18 models reportedly leaked in Tata Electronics hack Data breaches going unreported – Irish compliance survey Data breaches going unreported, says Irish compliance survey
Cloudflare to block AI crawlers from ad-supported webpage...
Colin Ryan · 2026-07-02 · via Enterprise – Silicon Republic

Come 15 September, multipurpose crawlers used by the likes of Google, Microsoft and Apple will be blocked by default according to Cloudflare’s new rules.

IT and network services provider Cloudflare has announced new rules designed to give website owners more control over the types of web crawlers that will be allowed on or blocked from their sites – along with plans to block multipurpose crawlers by default on ad-supported pages.

Traditionally, search engines and websites maintained a sort of “symbiotic relationship”, as Cloudflare puts it, whereby web owners allowed search engines to crawl their sites and in return, search engines sent users back to their pages.

The company explained that this crawl-to-referral process, when balanced, would help sites generate the pageviews needed to sustain advertising, affiliate revenue and subscriptions.

However, the rise of AI crawlers and agents changed things, as AI chatbots scrape sites to synthesise answers and bypass original sources – often leading to imbalanced crawl-to-referral ratios. Cloudflare’s own research from last year noted ratios ranging from 118:1 up to nearly 50,000:1 – meaning an AI crawler could have scraped a site tens of thousands of times and only sent back a single user.

Nowadays, many of these crawlers are used for multiple purposes – including AI training and search indexing – which puts website owners in a difficult position, as turning off all automation and crawler access to their sites could diminish their chances of showing up on search results.

Cloudflare hopes to tackle this issue with its new rules, which include options for managing crawler access by establishing three categories of crawler purposes: Search, Agent and Training.

‘Search’ refers to crawlers that are used for search indexing, ‘Agent’ refers to automated behaviours used by the likes of chatbots and browser-use agents, and ‘Training refers to crawlers that scrape content for fine-tuning AI models.

With these three classifications, website owners will be able to selectively allow or block crawlers that are used for each of the three classifications – meaning that if a web owner wanted to allow Search crawlers but block Agent and Training crawlers, they will now be able to do so

As part of these new rules, Cloudflare will also block Training and Agent crawlers by default on pages that display ads.

The default block settings, which will apply to any new domain onboarded to Cloudflare from 15 September, won’t apply to crawlers used for search indexing, while multipurpose crawlers – specifically those used for both search and training purposes – will be allowed or blocked “according to all of their behaviours”.

As a result, multipurpose crawlers used by the likes of Google, Microsoft and Apple will be blocked by default come 15 September.

“We believe it should be simple for all website owners to manage access for these three AI-centred use cases,” read a blogpost by Cloudflare. “We believe that bot operators should separate their crawlers because that creates more transparency for website owners, allowing them to better understand why a given crawler is visiting them as well as to better manage the access they extend to that crawler.

“If a company runs automation that builds Search indexes, acts as an Agent, and collects data to Train their models, then we strongly encourage that company to separate the automation into three separate crawlers.”

In the lead-up to the September default deadline, Cloudflare customers can opt out of the default settings if they want to.

Cloudflare’s new rules are the latest in the company’s attempts to curb crawler misuse.

This time last year, the company introduced new crawler controls for website owners, including a ‘pay per crawl’ system designed to integrate with existing web infrastructure and leverage HTTP status codes and established authentication mechanisms to create a framework for paid content access.

The year before that, Cloudflare introduced a tool that allowed website owners to block all bots at once.

Don’t miss out on the knowledge you need to succeed. Sign up for the Daily Brief, Silicon Republic’s digest of need-to-know sci-tech news.