惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

月光博客
月光博客
小众软件
小众软件
爱范儿
爱范儿
Y
Y Combinator Blog
博客园 - Franky
美团技术团队
博客园 - 【当耐特】
The Cloudflare Blog
罗磊的独立博客
Hugging Face - Blog
Hugging Face - Blog
Jina AI
Jina AI
IT之家
IT之家
人人都是产品经理
人人都是产品经理
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
大猫的无限游戏
大猫的无限游戏
Apple Machine Learning Research
Apple Machine Learning Research
博客园 - 聂微东
WordPress大学
WordPress大学
V
Visual Studio Blog
博客园_首页
阮一峰的网络日志
阮一峰的网络日志
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
有赞技术团队
有赞技术团队

Matthias Ott

Hello Again, World This, Still Not for Everyone The Shape of Friction WeissKlang L1 – Punching Above Its Weight Continvoucly Morged Value Webspace Invaders To Affinity and Beyond The Mystery of Storytelling Amateurs! Echoes of Connection Linear() Is Not (That) Linear View Transitions: The Smooth Parts Adding AVIF and WebP Support to My Craft CMS Site Challenge Acoustic Room Treatment and Building Sound Panels, Part 1: Planning Play On Overshoot The HTML Output Element Listening Closely Compressed Fluid Typography The Lifeblood of the Web What Could Go Wrong? That’s My Rank Making Space CSS :is() :where() the Magic Happens Visual Regression Testing for External URLs With Playwright Jane Goodall’s Famous Last Words European Tech Alternatives 🇪🇺 Independent Type Foundry Advent Calendar – Day 24: NaN Independent Type Foundry Advent Calendar – Day 23: Typotheque
Good Riddance, GPTBot · Matthias Ott
Matthias Ott · 2023-08-10 · via Matthias Ott

Just like Google is constantly indexing the Web, OpenAI is now crawling the open Web to scrape content from websites for free to train their LLM (lucrative language model) “AI” products.

But, as I learned from a post by Ethan on Mastodon, you can disallow GPTBot to get its tiny robot hands on your writing by adding those two lines of code to your website’s robots.txt:

User-agent: GPTBot
Disallow: /

Good riddance, GPTBot! 👋

~

43 Webmentions

  1. @matthiasott I really wish there was a way to disallow all AI bots. I don't want to opt out of search indices but it doesn't seem like there's a good way to get out of Google's LLM models

  2. @janboddez Yes, I wrote that post on the smartphone in bed and didn’t find the time to update the robots.txt myself yet. But thanks for the reminder! I just added the two lines. ✅;

  3. @brunomiguel @matthiasott Maybe we should start taking content off the net and ship discs (well, flash drives nowadays) to friends again instead. 🤷 Imagine, getting an USB drive with random cool stuff every month, without knowing what's on it.

  4. @frederic this reminds me of the old days of getting infected with malware 😍 @matthiasott

  5. @frederic @matthiasott tech bros making the world a worst place, one stupid shit at a time

  6. @matthiasott Hmm yes they will definitely respect this. Hmmm openai is very trustworthy. Mmmmm

  7. @matthiasott I also did it about an hour ago. If my hosting allowed it, I would even block their IP ranges

  8. @brunomiguel @matthiasott Well, looks like that's only the tip of the iceberg: https://searchengineland.com/google-content-available-ai-training-publishers-opt-out-430475 Google says all online content should be available for AI training unless publishers opt out

17 Reposts

17 Likes

ⓘ Webmentions are a way to notify other websites when you link to them, and to receive notifications when others link to you. Learn more about Webmentions.

More Notes