Post #2662911
2026-04-03 19:01 UTC
@markgritter@mathstodon.xyz @nicole@tietz.social anthropic has a paper where they basically discover their LLM has a "persona space". it has clusters for things like "text from a helpful expert", and by messing with weights you can steer responses toward or away from the "helpful expert" direction.
of course this is just *sounding like* a helpful expert. there's another dimension for "text resembles memorized training data" (which is still not the same as truth, of course. RAG gets a little closer, but still misses)
Replies (0)
No replies.