Elektrine lite

← Feed

@claudius@darmstadt.social

Post #1359220

2025-10-25 21:11 UTC

@Codeberg AI companies crawl our websites. We ask that they stop by using the industry standard robots.txt AI companies ignore those rules. We start blocking the companies themselves with conventional tools like IP rules. AI companies start working around those blocks. We invent ways to specifically make life harder for their crawlers (stuff like Anubis). AI companies put considerable resources into circumventing that, too. This industry seriously needs to implode. Fast.

Replies (8)

  • @claudius@darmstadt.social 2025-10-25 21:13

    As a next step, AI companies are now offering "their" browser (read: Chromium ever so slightly themed with some company bullshit built in) In part, this is certainly done to have yet another way to crawl the web, but this time user-directed and indistinguishable from actual human requests.

    Open ##1359221

  • @adrianco@mastodon.social 2025-10-25 21:16

    @claudius @Codeberg Have you tried using Firecrawl as a test for your blocking? It seems to be a popular site that centralizes crawling technology.

    Open ##1359235

  • @claudius @Codeberg I feel like we are working towards a point where you have to redesign the whole web to account for AI ignoring rules. New browsers, new protocols, etc.

    Open ##1359237

  • @tknarr@mstdn.social 2025-10-26 02:30

    @claudius @Codeberg At some point we're going to start paying a lawyer a few dollars to send the AI companies a registered return-receipt-requested letter saying "You are denied access to my web site. I have taken every step possible to prevent you from accessing it. If you continue to circumvent these measures and access my site anyway, you will be billed $1000/access. This fee will take effect 14 days after you receive this notice." Then start sending bills.

    Open ##1359244

  • @zebratale@mastodon.nz 2025-10-26 06:31

    @claudius @Codeberg IANAL but I'm wondering if adding explicit copyright text to every page they consume might at least provide some future ability to claim against them, particularly if it included specific fees for use without permission?

    Open ##1359254

  • @theogrin@chaosfem.tw 2025-10-26 06:40

    @claudius @Codeberg More folks need to begin adopting ... unorthodox solutions for those groups which have been so wonderful as to ignore robots.txt. Disguised petabyte ZIP bombs. Poisoned pages. Image folders chock full of Nightshade. The legal argument to be made and adopted here is that if the companies weren't willfully breaking the law, then they wouldn't have subjected themselves to those attacks. It certainly doesn't even fall under entrapment in most cases.

    Open ##1359255

  • @FlohEinstein@chaos.social 2025-10-26 06:43

    @claudius @Codeberg don't fight them circumventing it. Feed them garbage. I let the crawlers download gigabytes of randomness everyday, generated from Discworld books on my instance of Iocaine by @algernon : https://olyfjan.blomi.is

    Open ##1359256

  • @claudius @Codeberg "They" do not appear to be very good with Captcha and similar mechanisms... but it would appear a lot of people don't like them either!

    Open ##1359257