Post #1769216
2026-04-20 21:28 UTC
interesting paper on how training models on (partial) output from other models transfers unexpected (and difficult to track) additional representational content: this can amplify unwanted signals/behaviours even through training data that seem wholly unrelated to that preference/behaviour
this occurs due to the interaction between gradient descent as a training procedure, shared initialisation, and the fact that internal representations involve super-position #LLM
https://www.nature.com/articles/s41586-026-10319-8.pdf
Replies (0)
No replies.