Post #2446134
2025-01-14 22:58 UTC
Replies (40)
-
@mdione@en.osm.town 2025-01-16 12:09
@aaron@chirp.zadzmo.org FWIW, @jwz wrote his own: https://www.jwz.org/blog/2025/01/exterminate-all-rational-ai-scrapers/
-
@tezoatlipoca@mas.to 2025-01-16 17:33
@aaron@chirp.zadzmo.org this is awesome and I'm install this this weekend.
-
@mitchkiah@mastodon.xyz 2025-01-16 18:09
@aaron@chirp.zadzmo.org @abraxos@mastodon.social you might get some amusement from this one
-
@fluffy@plush.city 2025-01-17 03:06
@aaron@chirp.zadzmo.org I used to run something similar to trap spam email harvesters. Definitely worth doing but it benefits from having its own domain to burn.
-
@madonius@chaos.social 2025-01-17 11:08
@aaron@chirp.zadzmo.org Since AI-Companies have been demonstrated to ignore robots.txt one could put the tool behind a forbidden path and punish everyone who "chooses" to ignore it. Regular web crawlers should not punish your site, no?
-
@ghouston@mamot.fr 2025-01-17 11:21
@aaron@chirp.zadzmo.org For AI you say? I am also finding this content intriguing: forests and NOx emissions was recently named one of the propagating organization. Semantics All of them. Have you not refuse me a drink from his art agent in Europe. The Jewish National Fund was founded in 1958. He established himself as conspicuous as though it is imagined lisps benefices gruff prevention maintenance
-
@para_paramoney@mastodon.social 2025-01-17 12:23
@aaron@chirp.zadzmo.org @teknologisk@koop.social jeg forstår ikke helt det her, men måske noget der er værd at lege med
-
@itgrrl@infosec.exchange 2025-01-18 07:51
@aaron@chirp.zadzmo.org another one for you, @KathyReid@aus.social 👀
-
@beaker@freeradical.zone 2025-01-18 14:56
@aaron@chirp.zadzmo.org bravo, I was considering creating something like this myself after I heard about some AI crawlers ignoring robots.txt. I'd combine this with a robots.txt that keeps away well behaved search engine crawlers and only punishes the bad actors.
-
@shanesemler@metalhead.club 2025-01-18 15:20
@aaron@chirp.zadzmo.org I fail to see how this helps people or the state of the internet.
-
@viq@social.hackerspace.pl 2025-01-18 15:48
@aaron@chirp.zadzmo.org I think that's about fifth OSS project called that 😂
-
@cxj@phpc.social 2025-01-18 16:38
@aaron@chirp.zadzmo.org W00t!
-
@kc@social.coop 2025-01-18 21:25
@aaron@chirp.zadzmo.org thank you for this. I've thrown one up on my site ! No matter what I do to get rid of these bots they just come at my server, cost me money and time ! Not figured out the Markov Babble though, managed to get an error 😅
-
@OmegaPolice@hachyderm.io 2025-01-22 22:48
@aaron@chirp.zadzmo.org This is genius and hilarious. Thank you! Out of curiosity: Does nepenthes provide a robots.txt so that well-behaved crawlers are safe? Can we _potty-train_ crawlers? 😬
-
@Suiseiseki@freesoftwareextremist.com 2025-01-16 11:48
@aaron@chirp.zadzmo.org What is the license? I see a version of MIT expat in the project root, but there's no details as to which files that applies to (if any) or license headers in the source files.
-
@aka_dude@mk.phreedom.club 2025-01-16 09:12
@aaron@chirp.zadzmo.org people really do be hatin that ai guy
-
@exa@mastodon.online 2025-01-16 07:27
@aaron@chirp.zadzmo.org Next project: Jekyll or hugo theme that looks exactly like #nepenthes so that when the AI crawlers get configured to avoid these things, they also finally leave my site alone
-
@thelovebing@mastodon.nu 2025-01-16 07:17
@aaron@chirp.zadzmo.org if you run this on a sub domain, would the main domain (or whatever it’s called) disappear from search engines too? Also, it would be great with some kind of one click installation á la Wordpress for us non-tech folks. Is that doable?
-
@Okanogen@mastodon.social 2025-01-15 23:48
@aaron@chirp.zadzmo.org Trapping Moriarty in the holodeck?
-
@lazynezumi@mstdn.social 2025-01-15 21:03
@aaron@chirp.zadzmo.org that's great but now I need a new door!
-
@tekhedd@byteheaven.net 2025-01-15 18:48
@aaron@chirp.zadzmo.org I love tarpits! If you put your tarpit behind a "robots.txt" it's amazing how many hits you still get.
-
@lena@social.treehouse.systems 2025-01-15 18:28
@aaron@chirp.zadzmo.org i want to try deploying this on a separate subdomain, integrated with fail2ban then, you add a hidden (and aria-hidden) tag to it. not like i care THAT MUCH about being listed on search engines, but still
-
@UweHalfHand@norcal.social 2025-01-15 18:23
@aaron@chirp.zadzmo.org Nice! Not ready to deploy this, I don’t have a website up and running, but making a note of it… 👍
-
@scarpentier@hachyderm.io 2025-01-15 18:22
@aaron@chirp.zadzmo.org love this kind of offensive tech. Poison them all! LET'S GOOOO! 🔥
-
@nat@gts.blahaj.pl 2025-01-15 17:14
@aaron@chirp.zadzmo.org Do you think they'll find a way to detect that they're crawling markov slop? Do you think there's a way to make this completely undistinguishable for a bot from a regular website? Maybe returning common headers from eg. websites running wordpress and generating a html structure similar to what some popular themes would return? If they tried detecting something like this maybe it would give them false-positives on real stuff 🐱
-
@Mia@social.translunar.academy 2025-01-15 16:42
@aaron@chirp.zadzmo.org hmm i wonder if endless torture inside a tarpit would create spontaneously sentience
-
@tati@kind.social 2025-01-15 16:27
@aaron@chirp.zadzmo.org chaotic good
-
@zillion@freeradical.zone 2025-01-15 16:14
@aaron@chirp.zadzmo.org http://libraryofbabel.info/
-
@Beachbum@mastodon.sdf.org 2025-01-15 15:28
@aaron@chirp.zadzmo.org Pretty smart.
-
@martinvermeer@fediscience.org 2025-01-15 15:06
@aaron@chirp.zadzmo.org Can it be made to punish only crawlers disrespecting robots.txt?
-
@funes@infosec.exchange 2025-01-15 15:05
@aaron@chirp.zadzmo.org in addition to suggesting its use for informing IP ban lists, you might consider adding in a feature some day to opt in to sharing the list of IPs caught in the trap back to a service that can aggregate them from all the instances for use as an intel feed, kind of like the d-shield honeypot https://en.m.wikipedia.org/wiki/DShield
-
@solonovamax@tech.lgbt 2025-01-15 14:52
@aaron@chirp.zadzmo.org this looks interesting, I'll check it out later I was thinking of building smth like this but it also uses an llm to generate garbage data to poison their datasets it would need a bit more power, but it'd cache the pages, so it'd use a tad bit more cpu. but, I also plannes for it to be like the absolute tiniest llm I could find. it could also stream the pages very slowly to the bots, just to be annoying. but this uses a markov chain, so that might be sufficient? unsure.
-
@riskythinking@infosec.exchange 2025-01-15 13:32
@aaron@chirp.zadzmo.org I do a similar thing for bots trying to harvest email addresses which disobey the robots.txt file. I suspect the pages need to be more structured and longer to have a significant effect: otherwise it's just random noise that can average out.
-
@ToweroftheArchmage@chirp.enworld.org 2025-01-15 13:13
@aaron@chirp.zadzmo.org is this the sort of thing that would run like SETI at home, when your computer is idle?
-
@0x4261756D@mastodon.online 2025-01-15 13:04
@aaron@chirp.zadzmo.org Do LLM crawlers typically honour robots.txt? Otherwise this could be a way to avoid spare innocent crawlers.
-
@froge@social.glitched.systems 2025-01-15 12:40
@aaron@chirp.zadzmo.org this is probably useful for lots of stuff related to trapping bots/scrapers, this is super nice!!!
-
@Jgmeadows@mstdn.ca 2025-01-15 12:12
@aaron@chirp.zadzmo.org @bjb@fosstodon.org Very cool! I don’t currently have any server space set up but you have me thinking!
-
@gcvsa@mstdn.plus 2025-01-15 12:08
@aaron@chirp.zadzmo.org you are in a maze of twisty little passages all alike
-
@destiny@social.hailstorm.gay 2025-01-15 12:00
@aaron@chirp.zadzmo.org This... I like this... Seems like a fun way to test a web crawler, but that made me think about robots.txt. You could for instance disallow the way back machine from entering it, but allow for all others. Not sure in practice what that'd result in, but seems interesting.
-
@whvholst@eupolicy.social 2025-01-15 11:06
@aaron@chirp.zadzmo.org Now combine that with gzip shenanigans!