惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

美团技术团队
Microsoft Azure Blog
Microsoft Azure Blog
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
B
Blog
Y
Y Combinator Blog
博客园_首页
有赞技术团队
有赞技术团队
博客园 - Franky
腾讯CDC
G
Google Developers Blog
Recent Announcements
Recent Announcements
博客园 - 【当耐特】
D
Docker
The GitHub Blog
The GitHub Blog
MyScale Blog
MyScale Blog
H
Help Net Security
Apple Machine Learning Research
Apple Machine Learning Research
A
About on SuperTechFans
D
DataBreaches.Net
T
The Blog of Author Tim Ferriss
V
V2EX
U
Unit 42
aimingoo的专栏
aimingoo的专栏
WordPress大学
WordPress大学

Matthias Ott

Hello Again, World This, Still Not for Everyone The Shape of Friction WeissKlang L1 – Punching Above Its Weight Continvoucly Morged Value Webspace Invaders To Affinity and Beyond The Mystery of Storytelling Amateurs! Echoes of Connection Linear() Is Not (That) Linear View Transitions: The Smooth Parts Adding AVIF and WebP Support to My Craft CMS Site Challenge Acoustic Room Treatment and Building Sound Panels, Part 1: Planning Play On Overshoot The HTML Output Element Listening Closely Compressed Fluid Typography The Lifeblood of the Web What Could Go Wrong? That’s My Rank Making Space CSS :is() :where() the Magic Happens Visual Regression Testing for External URLs With Playwright Jane Goodall’s Famous Last Words European Tech Alternatives 🇪🇺 Independent Type Foundry Advent Calendar – Day 24: NaN Independent Type Foundry Advent Calendar – Day 23: Typotheque
Good Riddance, GPTBot · Matthias Ott
Matthias Ott · 2023-08-10 · via Matthias Ott

Just like Google is constantly indexing the Web, OpenAI is now crawling the open Web to scrape content from websites for free to train their LLM (lucrative language model) “AI” products.

But, as I learned from a post by Ethan on Mastodon, you can disallow GPTBot to get its tiny robot hands on your writing by adding those two lines of code to your website’s robots.txt:

User-agent: GPTBot
Disallow: /

Good riddance, GPTBot! 👋

~

43 Webmentions

  1. @matthiasott I really wish there was a way to disallow all AI bots. I don't want to opt out of search indices but it doesn't seem like there's a good way to get out of Google's LLM models

  2. @janboddez Yes, I wrote that post on the smartphone in bed and didn’t find the time to update the robots.txt myself yet. But thanks for the reminder! I just added the two lines. ✅;

  3. @brunomiguel @matthiasott Maybe we should start taking content off the net and ship discs (well, flash drives nowadays) to friends again instead. 🤷 Imagine, getting an USB drive with random cool stuff every month, without knowing what's on it.

  4. @frederic this reminds me of the old days of getting infected with malware 😍 @matthiasott

  5. @frederic @matthiasott tech bros making the world a worst place, one stupid shit at a time

  6. @matthiasott Hmm yes they will definitely respect this. Hmmm openai is very trustworthy. Mmmmm

  7. @matthiasott I also did it about an hour ago. If my hosting allowed it, I would even block their IP ranges

  8. @brunomiguel @matthiasott Well, looks like that's only the tip of the iceberg: https://searchengineland.com/google-content-available-ai-training-publishers-opt-out-430475 Google says all online content should be available for AI training unless publishers opt out

17 Reposts

17 Likes

ⓘ Webmentions are a way to notify other websites when you link to them, and to receive notifications when others link to you. Learn more about Webmentions.

More Notes