Elektrine lite

← Feed

@bornach@masto.ai

Post #2790141

2023-04-07 18:18 UTC

@ngaylinn@tech.lgbt LLMs are never your friend They are always in what can best be described as "the grifter" mode. The entire training regime of a generative AI chatbot is geared towards getting one thing, an upvote from a human rating the quality of the conversation Admittedly this is an over-simplification. Reinforcement Learning with Human Feedback involves training a reward policy - a 2nd neural network that is ultimately responsible for rewarding the chatbot for giving "good" responses.

Replies (1)

  • @KingmaYpe@mastodon.green 2025-12-27 20:56

    @bornach@masto.ai @ngaylinn@tech.lgbt Indeed, that second network is trained to be "engaging" in conversation. Its goal is to keep the attention of the user, mostly for marketing purposes. It's not really difficult to recognise, but it's never easy.

    Open ##2790142