Elektrine lite

← Feed

@conniptions@mastodon.social

2026-07-27 23:41 UTC

@lzg@mastodon.social I've heard people arguing this but I remain as yet unconvinced. Recently I came across this academic study showing why prompt injections can't be stopped: https://role-confusion.github.io/ - the point being that the attempt to make LLMs more than just predictive text, by separating prompt input into different category roles, are themselves so very weakly implemented they can be easily circumvented. Guardrails don't, in short. Obviously I'm missing something? But what?

Replies (2)

  • @muddle@infosec.exchange 2026-07-28 00:50

    @conniptions@mastodon.social @lzg@mastodon.social "Guardrails don't, in short." I think I've used those exact words myself at some point. The very idea of "guardrails" is basically a lie. What it means, in practice, is "massaging the data" (in a few different senses of that expression).

    Open ##4151138

  • @lzg@mastodon.social 2026-07-28 01:24

    @conniptions@mastodon.social the predictive text thing is fundamentally true but not enough. they are definitively not "intelligences", and I think that's where the hype is most annoying. but they are capable of containing an enormous amount of accumulated context of human culture (at least of humanity not excluded from the training data) and of reflecting it back to us. isn't that something?

    Open ##4332646