Elektrine lite

← Feed

@khobochka@mastodon.social

Post #2563017

2024-12-27 10:25 UTC

#LLMs are a fucking scourge. Perceiving their training infrastructure as anything but a horrific all-consuming parasite destroying the internet (and wasting real-life resources at a grand scale) is delusional. #ChatGPT isn't a fun toy or a useful tool, it's a _someone else's_ utility built with complete disregard for human creativity and craft, mixed with malicious intent masquerading as "progress", and should be treated as such. https://pod.geraspora.de/posts/17342163

Replies (31)

  • @Lily_and_frog@mastodon.art 2024-12-27 13:52

    @khobochka@mastodon.social That's fucking insane!

    Open ##2972662

  • @kurio@sunny.garden 2024-12-27 14:10

    @khobochka@mastodon.social This is terrible. I updated the crawler blocklist on the robots.txt of my web and they seem to havemultiplied the crawl attempts x10 times. :blobdisapproval:

    Open ##2972663

  • @gimulnautti@mastodon.green 2024-12-27 14:42

    @khobochka@mastodon.social We need an international co-operative system of making these parties pay for scraping. It includes legislative changes. At the same time it can become a real-time pricing market for ”rights to scrape” and for creators to get paid. Here’s my whitepaper for a solution. Absolutely no cryptocurrency involved. #ai #scraping #copyright #technology #whitepaper https://docs.google.com/document/d/18cz-ZX1copCYiC4C2ReY8GLJjuhG2IH0MEBGaoSJhP4/edit

    Open ##2972669

  • @gimulnautti@mastodon.green 2024-12-28 12:49

    @kkarhan@infosec.space @khobochka@mastodon.social I don't believe in denying business or market incentives works either. Your scheme has low risk but terrible chance of adoption. My scheme has moderate risk but at least a possibility of adoption. Silicon Valley has itself participated in not getting regulated, by boosting this narrative of "all attempts to govern us will end in rich collecting schemes". We fight amongst ourselves. While they use "freedom" to inflict their tyranny on our assets.

    Open ##2972671

  • @ide@masto.ai 2024-12-27 14:53

    @khobochka@mastodon.social Doing ultra-wide dumb crawls without proper domain rate limit or caching and the like just sounds like severe developer incompetence.

    Open ##2972675

  • @niko@furry.engineer 2024-12-28 04:05

    @kkarhan@infosec.space may i also add i personally redirect these bots to gz.niko.lgbt which returns a innocent looking 100MB HTTP response with content-encoding: gzip that decompresses to 100GB and the only way to find out is to actually decompress it so if they want anything they gotta go through it

    Open ##2972679

  • @drwho@hackers.town 2024-12-28 22:22

    @kkarhan@infosec.space @niko@furry.engineer I use both. Belt and suspenders. What I want to do is maintain one set of files for the stuff I have in shared housing. I keep the ones for my website up to date and plan to write a script that copies them into other webdirs. The nginx servers are all behind http basic auth, so they can't see them anyway.

    Open ##2972683

  • @drwho@hackers.town 2024-12-29 05:56

    @kkarhan@infosec.space @niko@furry.engineer Sloworis but for clients?

    Open ##2972684

  • @clusterfcku@mastodon.social 2024-12-27 16:45

    @khobochka@mastodon.social Change your UA. put them on notice -with a legal letter- that certain of their bots will be charged $ per visit. Wait a reasonable amount of time (eg 4 weeks). Start invoicing. Sue in small claims. Donate the money to FOSS search/archive/fedi etc

    Open ##2972686

  • @khobochka@mastodon.social ...that's consistent with some weird periodic slowdown behaviour I've started seeing on my little personal website over the last year or so. So far I've just waited for it to go away, as the ssh console is unresponsive during the problem, but next time it happens I guess I'll check the access log for signs of Abominable Intelligence.

    Open ##2972687

  • @bposi@mastodontech.de 2024-12-27 21:47

    @khobochka@mastodon.social They are like pirates - or more prianha's.. I really can't say on how bored I am about his fucking hype on LLM. Everywhere AI is mentioned. Even at work the people act as pirates of trying to circumvent rules on compliance to dig the gold they believe is there. And risking everything.. Someone needs to write a scanner that is checking the Logs for such misuse - and block each of such activity.

    Open ##2972689

  • @oddevan@mastodon.social 2024-12-28 00:29

    @khobochka@mastodon.social I remember back in the aughts there was talk that getting around an IP or UA block violated the anti-circumvention clause in the DMCA. Wonder if we could test that here.

    Open ##2972690

  • @khobochka@mastodon.social It's a weapon.

    Open ##2972691

  • @trisweb@m.trisweb.com 2024-12-28 01:44

    @khobochka@mastodon.social the power costs of this across the entire internet might be approaching the power of running the LLMs themselves.

    Open ##2972692

  • @jay_nakrani@mastodon.world 2024-12-28 02:13

    @khobochka@mastodon.social LLMs, by themselves, aren't the issue. The issue is human greed -- to be the few winners in this new technology S-curve. That is what's turning people into assholes, and causing all of these problems. Greed leads to cut throat competition, which in turns leads to people ignoring rules (or creative interpretations/advocacy to suite their greed/ambition).

    Open ##2972693

  • @unlofl@mstdn.social 2024-12-28 02:40

    @khobochka@mastodon.social the fact they switch ips and user agents is so scummy. I've been thinking data poisoning is the only real defense, it doesn't save cost but fuck em, I'll burn CPU generating and maintaining wrong content to ingest.

    Open ##2972695

  • @arisunz@gts.arielaw.ar 2024-12-28 04:43

    @khobochka@mastodon.social damn, didnt know the damn crawlers got aggressive to the point of changing their fucking user agent now... luckily i still have a couple1 aces2 up my sleeve but ugh

    Open ##2972696

  • @pinkprius@chaos.social 2024-12-28 09:40

    @khobochka@mastodon.social straight up evil

    Open ##2972697

  • @vampirdaddy@chaos.social 2024-12-28 10:00

    @khobochka@mastodon.social robots.txt ist just _asking nicely, prettyplease_ who should not wander into which territories. For real blocks one has to block such systems - by IP or other indicators. Depending on server it's comparatively easy to filter out user agents (on lighttpd you can easily filter by useragent). While at it: does one have ideas for nice , ressource-light adversial learning garbage that I could feed those suckers? (though the zip bomb is nice ide, too)

    Open ##2972698

  • @wtfrank@mastodon.social 2024-12-28 12:35

    @khobochka@mastodon.social there's clearly a free rider problem, where these LLMs benefit from other people's work to train their models. In my view, when you use other people's data to train your model that should give them a partial copyright over the model, as the model becomes a derivative work. It's one thing a search engine crawling as that will send traffic your way, but an LLM crawling is all taken, no give.

    Open ##2972699

  • @JustinDerrick@mstdn.ca 2024-12-28 12:52

    @khobochka@mastodon.social I rescued an old forum whose administrator decided they weren’t interested anymore. I was aghast at the traffic coming from the big AI companies. The first few weeks, I spent an hour, every day, blocking huge netblocks at the firewall. I wrote a script that summarized the heavy-hitters to the web server, investigated each one manually, then added it to the firewall’s blocklist. At the end of the first month, I was blocking nearly 1% of all IPv4 addresses.

    Open ##2972700

  • @Herover@helvede.net 2024-12-28 13:00

    @khobochka@mastodon.social and the crawlers aren't just stupid for repeatedly crawl the same identical pages over and over, they also do completely nonsensical stuff. At my work we noticed that GPTBot had found our search button and did what looks like 100.000 requests per day to /search/amp;amp;amp; repeated 100 times and then our support phone number or a url to a random image followed by more amp;'s.

    Open ##2972703

  • @Isurandil@mastodon.online 2024-12-28 15:09

    @khobochka@mastodon.social So, we need to find ways to penalize them without them noticing. Rate-limiting and UA filtering did not do the trick, so we can assume that trying to feed poisoned content to them will also be detected easily. Maybe the only way to protect against automated LLM bot groping is to put expensive functionality and special content behind login walls, going back to the internet driven by small communities, except that this time they are mostly closed to discovery from the outside.

    Open ##2972704

  • @mathling@mastodon.social 2024-12-28 15:50

    @khobochka@mastodon.social Can confirm. They are a huge fraction of my website traffic and they don't honour any constraints I put on them short of hard blocks

    Open ##2972705

  • @snowyfox@deadinsi.de 2024-12-28 23:44

    .

    Open ##2972706

  • @groxx@hachyderm.io 2024-12-29 02:04

    @khobochka@mastodon.social time for a widely shared IP ban list? The only way they're gonna listen is if we make it expensive for them.

    Open ##2972707

  • @ClickyMcTicker@hachyderm.io 2024-12-29 04:40

    @khobochka@mastodon.social Laws like this one should apply to scraping that ignores the standard robots.txt: https://www.law.cornell.edu/uscode/text/18/1030 An automated program willfully ignoring an explicit order to stop should result in prison time for the perpetrators. Add additional financial penalties for businesses whose employees do it, and slap percentages of annual revenue on any noncompliance for failing to turn over employees implicated in illegal scraping. The issue would dry up immediately.

    Open ##2972708

  • @khobochka@mastodon.social and it is not even useful, I once asked ChatGPT how to destroy the world and all it could offer were platitudes

    Open ##2972709

  • @Sonic2k@oldbytes.space 2024-12-29 12:32

    @khobochka@mastodon.social fully agree.

    Open ##2972710

  • @Diziet@tech.lgbt 2024-12-29 12:40

    @khobochka@mastodon.social I think that if either the victim site, or the slop slurper, are in the UK, this kind of behaviour is a criminal offence contrary to the Computer Misuse Act. The scale, and intentionality (eg ban evasion), means people should be going to prison.

    Open ##2972711

  • @stevel@hachyderm.io 2024-12-29 12:47

    @khobochka@mastodon.social they'll be examining the diff entries just to pick up any human conversation in the comments-searching for the last sentences of non-machine-generated text. I believe IMDb used to deal with bots by throttling- not enough to trigger "countermeasures" but enough to reduce their damage. every 10s pause on a GET request occupies a TCP socket on the client which would otherwise put load on other sites

    Open ##2972712