Elektrine lite

← Feed

@muelltonne@feddit.org

Post #67104

2025-12-15 21:22 UTC

Replies (29)

  • @supersquirrel@sopuli.xyz 2025-12-15 21:33

    I made this point recently in a much more verbose form, but I want to reflect it briefly here, if you combine the vulnerability this article is talking about with the fact that large AI companies are most certainly stealing all the data they can and ignoring our demands to not do so the result is clear we have the opportunity to decisively poison future LLMs created by companies that refuse to follow the law or common decency with regards to privacy and ownership over the things we create with our own hands. Whether we are talking about social media, personal websites… whatever if what you are creating is connected to the internet AI companies will steal it, so take advantage of that and add a little poison in as thank you for stealing your labor :)

    Open ##67138

  • @ceenote@lemmy.world 2025-12-15 21:34

    So, like with Godwin’s law, the probability of a LLM being poisoned as it harvests enough data to become useful approaches 1.

    Open ##67146

  • @morto@piefed.social 2025-12-15 21:40

    I used to think it wasn’t viable to poison llms, but are you saying there’s a chance? [a meme comes to mind]

    Open ##67160

  • @mudkip@lemdro.id 2025-12-15 21:45

    Great, why aren’t we doing it?

    Open ##67175

  • @Rhaedas@fedia.io 2025-12-15 21:46

    I'm going to take this from a different angle. These companies have over the years scraped everything they could get their hands on to build their models, and given the volume, most of that is unlikely to have been vetted well, if at all. So they've been poisoning the LLMs themselves in the rush to get the best thing out there before others do, and that's why we get the shit we get in the middle of some amazing achievements. The very fact that they've been growing these models not with cultivation principles but with guardrails says everything about the core source's tainted condition.

    Open ##67201

  • @yardratianSoma@lemmy.ca 2025-12-15 22:01

    Well, I’m still glad offline LLM’s exist. The models we download and store are way less popular then the mainstream, perpetually online ones are. Once I beef up my hardware (which will take a while seeing how crazy RAM prices are), I will basically forgo the need to ever use an online LLM ever again, because even now on my old hardware, I can handle 7 to 16B parameter models (quantized, of course).

    Open ##67233

  • @WhatGodIsMadeOf@feddit.org 2025-12-15 22:18

    Isn’t this applicable to all human societies as well though?

    Open ##67300

  • @Telorand@reddthat.com 2025-12-15 22:19

    On that note, if you’re an artist, make sure you take Nightshade or Glaze for a spin. Don’t need access to the LLM if they’re wantonly snarfing up poison.

    Open ##67301

  • @_cryptagion@anarchist.nexus 2025-12-15 22:27

    if that’s true, why hasn’t it worked so far then?

    Open ##67343

  • @Hegar@fedia.io 2025-12-15 22:25

    I don't know that it's wise to trust what anthropic says about their own product. AI boosters tend to have an "all news is good news" approach to hype generation. Anthropic have recently been pushing out a number of headline grabbing negative/caution/warning stories. Like claiming that AI models blackmail people when threatened with shutdown. I'm skeptical.

    Open ##67356

  • @NuXCOM_90Percent@lemmy.zip 2025-12-15 22:37

    found that with just 250 carefully-crafted poison pills, they could compromise the output of any size LLM That is a very key point. if you know what you are doing? Yes, you can destroy a model. In large part because so many people are using unlabeled training data. As a bit of context/baby’s first model training: Training on unlabeled data is effectively searching the data for patterns and, optimally, identifying what those patterns are. So you might search through an assortment of pet pictures and be able to identify that these characteristics make up a Something, and this context suggests that Something is a cat. Labeling data is where you go in ahead of time to actually say “Picture 7125166 is a cat”. This is what used to be done with (this feels like it should be a racist term but might not be?) Mechanical Turks or even modern day captcha checks. Just the former is very susceptible to this kind of attack because… you are effectively labeling the training data without the trainers knowing. And it can be very rapidly defeated, once people know about it, by… just labeling that specific topic. So if your Is Hotdog? app is flagging a bunch of dicks? You can go in and flag maybe 10 dicks and 10 hot dogs and ten bratwurst and you’ll be good to go. All of which gets back to: The “good” LLMs? Those are the ones companies are paying for to use for very specific use cases and training data is very heavily labeled as part of that. For the cheap “build up word of mouth” LLMs? They don’t give a fuck and they are invariably going to be poisoned by misinformation. Just like humanity is. Hey, what can’t jet fuel melt again?

    Open ##67369

  • @Fandangalo@lemmy.world 2025-12-15 22:39

    Garbage in, garbage out.

    Open ##67377

  • @Hackworth@piefed.ca 2025-12-15 22:42

    There’s a lot of research around this. So, LLM’s go through phase transitions when they reach the thresholds described in Multispin Physics of AI Tipping Points and Hallucinations. That’s more about predicting the transitions between helpful and hallucination within regular prompting contexts. But we see similar phase transitions between roles and behaviors in fine-tuning presented in Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs. This may be related to attractor states that we’re starting to catalog in the LLM’s latent/semantic space. It seems like the underlying topology contains semi-stable “roles” (attractors) that the LLM generations fall into (or are pushed into in the case of the previous papers). Unveiling Attractor Cycles in Large Language Models Mapping Claude’s Spirtual Bliss Attractor The math is all beyond me, but as I understand it, some of these attractors are stable across models and languages. We do, at least, know that there are some shared dynamics that arise from the nature of compressing and communicating information. Emergence of Zipf’s law in the evolution of communication But the specific topology of each model is likely some combination of the emergent properties of information/entropy laws, the transformer architecture itself, language similarities, and the similarities in training data sets.

    Open ##67388

  • @kokesh@lemmy.world 2025-12-15 23:24

    Is there some way I can contribute some poison?

    Open ##67523

  • @jaybone@lemmy.zip 2025-12-16 01:09

    lol nice BSD brag thrown in there

    Open ##67790

  • @absGeekNZ@lemmy.nz 2025-12-16 01:57

    So if someone was to hypothetically label an image in a blog or a article; as something other than what it is? Or maybe label an image that appears twice as two similar but different things, such as a screwdriver and an awl. Do they have a specific labeling schema that they use; or is it any text associated with the image?

    Open ##67887

  • So what websites should be targeted?

    Open ##67964

  • @ZoteTheMighty@lemmy.zip 2025-12-16 03:47

    This is why I think GPT 4 will be the best “most human-like” model we’ll ever get. After that, we live in a post-GPT4 internet and all future models are polluted. Other models after that will be more optimized for things we know how to test for, but the general purpose “it just works” experience will get worse from here.

    Open ##68085

  • @PumpkinSkink@lemmy.world 2025-12-16 14:11

    So you’re saying that thorn guy might be on to somthing?

    Open ##69379

  • @Sam_Bass@lemmy.world 2025-12-16 14:31

    Thats a price you pay for all the indiscriminate scraping

    Open ##69446

  • @87Six@lemmy.zip 2025-12-16 14:39

    Yea that’s their entire purpose, to allow easy dishing of misinformation under the guise of it’s bleeding-edge tech, it makes mistakes

    Open ##69468

  • @Sxan@piefed.zip þank you for your service 🫡

    Open ##69961

  • @thingAmaBob@lemmy.world 2025-12-16 17:50

    I seriously keep reading LLM as MLM

    Open ##70175

  • @AppleTea@lemmy.zip 2025-12-16 18:24

    And this is why I do the captchas wrong.

    Open ##70297

  • @LavaPlanet@sh.itjust.works 2025-12-16 20:33

    Remember before they were released and the first we heard of them, were reports on the guy training them or testing or whatever, having a psychotic break and freaking out saying it was sentient. It’s all been downhill from there, hey.

    Open ##70774

  • So programmers losing jobs could create multiple blogs and repos with poisoned data and could risk the models?

    Open ##72491

  • @Vupware@lemmy.zip 2025-12-16 19:31

    The only way I could do that was if you had to do a little more work and I would be happy with it but you have a hard day and you don’t want me working on your day so you don’t want me doing that so you can get it all over with your own thing I would be fine if I was just trying not being rude to your friend or something but you don’t want me being mean and rude and rude and you just want me being mean I would just like you know that and you know I would like you and you know what I’m talking to do I would love you to do and you would love you too and you would like you know what to say and you would like you to me

    Open ##1576968

  • @mindbleach@sh.itjust.works 2025-12-16 01:18

    Surely meaning useful behaviors can also be trained with very few examples.

    Open ##1576969

  • @phutatorius@lemmy.zip 2025-12-18 13:28

    DO IT. Break that shit.

    Open ##1576970