Elektrine lite

← Feed

@gray17@mastodon.social

Post #2662911

2026-04-03 19:01 UTC

@markgritter@mathstodon.xyz @nicole@tietz.social anthropic has a paper where they basically discover their LLM has a "persona space". it has clusters for things like "text from a helpful expert", and by messing with weights you can steer responses toward or away from the "helpful expert" direction. of course this is just *sounding like* a helpful expert. there's another dimension for "text resembles memorized training data" (which is still not the same as truth, of course. RAG gets a little closer, but still misses)

Replies (0)

No replies.