Post #2576102
2026-03-26 10:06 UTC
Nah, guarantee the models have rules built in to deal with obvious stuff like that.
You need to be more subtle. Give them information that is *slightly* wrong.
Replies (5)
-
@taco@anarchist.nexus 2026-03-26 12:27
Perhaps by generating a bunch of complex copilot code to upload. It's easy to mass produce and would look plausibly functional.
-
@Viceversa@lemmy.world 2026-03-26 14:55
... and tell it things, that are *slightly* obscene
-
@ozymandias117@lemmy.world 2026-03-27 02:05
Just need to use less obvious insults, a la, "your mother was a hamster, and your father smelt of elderberries" Still poisons the model with something an end user won't like, but isn't easy enough to train out
-
@Aerosol3215@piefed.ca 2026-03-27 08:06
Artisanal crap code.
-
@bufalo1973@piefed.social 2026-03-28 01:45
Prompt for another AI: "write an example of code that looks correct but doesn't work" Step 2; upload the resulting code to GitHub. Step 3: make this an automated task.