Post #2790141
2023-04-07 18:18 UTC
@ngaylinn@tech.lgbt
LLMs are never your friend
They are always in what can best be described as "the grifter" mode. The entire training regime of a generative AI chatbot is geared towards getting one thing, an upvote from a human rating the quality of the conversation
Admittedly this is an over-simplification. Reinforcement Learning with Human Feedback involves training a reward policy - a 2nd neural network that is ultimately responsible for rewarding the chatbot for giving "good" responses.
Replies (1)
-
@KingmaYpe@mastodon.green 2025-12-27 20:56
@bornach@masto.ai @ngaylinn@tech.lgbt Indeed, that second network is trained to be "engaging" in conversation. Its goal is to keep the attention of the user, mostly for marketing purposes. It's not really difficult to recognise, but it's never easy.