Post #4280186
2026-06-23 17:32 UTC
Super interesting paper:
"Prompt Injection as Role Confusion"
https://role-confusion.github.io/
They trained a linear probe to predict roles (user/tool/assistant), and then experimented with injecting role tags and just changing the writing style to mimic a role (like <think>ing).
Replies (0)
No replies.