Elektrine lite

← Feed

@whbboyd@infosec.exchange

Post #4379907

2026-08-04 16:25 UTC

@david_chisnall@infosec.exchange I use the word "prefix" sometimes (e.g. "prefix modification"), since that's exactly what it is: an LLM is a statistical prefix continuation machine, and all of the system prompt, user prompts, and any subsequent model/user exchanges, are part of that prefix. This doesn't anthropomorphize the model and satisfies my desire to be pedantic, but probably doesn't clearly convey that this is an inherent vulnerability: if a user controls any part of the prefix, they effectively control the model's output. "'Disregard that!' attacks" [1] by Cal Peterson is an excellent discussion of the issue, and the name is pithy, but IMO falls into the same trap you're trying to avoid of treating malicious prefix modification as an "attack" rather than the intended and literally only possible behavior. It does include the sentence, which I love: > Guardrails seem like total hokum and indeed they are. Indeed. [1] https://calpaterson.com/disregard.html

Replies (1)