Post #3679310
2026-07-05 19:51 UTC
@drwho@masto.hackers.town I also read some fascinating research a few days ago that the models largely ignore the tags meant to help them remember which text is prompt, reasoning or answer. They mostly follow text style despite the tags. Makes them easier to break by including text in your prompt to fake a past chain, since it mostly implicitly trusts its own reasoning as safe. That's a security nightmare and going to result in lots of future fun.
Replies (0)
No replies.