Post #2686064
2025-05-12 13:01 UTC
@SimmerVigor@mastodon.social Oh nice! From the title and ToC that looks very much like the sort of thing I've been ยฝ mulling...๐๐ป
I also feel like we need to move to more of an intent/purpose-based model of identification/robots.txt blocking. The current system is way too onerous to maintain. Would be a lot better to be able to say e.g.:
* Search Engine bots: Yes
* LLM data collectors: No
etc.
Obv that relies on well-mannered bot operators but we do right now too.
Replies (1)
-
@gunchleoc@mastodon.scot 2025-05-12 17:48
@tdp_org@mastodon.social @SimmerVigor@mastodon.social Not being well-mannered is definitely a problem. One of my sites got hammered by an AI crawler a while ago, we found them via the user agent and they have actually published their source code. The tool has an explicit option to turn off respecting robots.txt. When I e-mailed them about it, they did not understand why that might be a problem ๐