@david_chisnall@infosec.exchange
Post #4482621
2026-08-10 08:15 UTC
Replies (2)
-
@budududuroiu@hachyderm.io 2026-08-10 08:21
@david_chisnall@infosec.exchange @futurebird@sauropods.win You don't have to anthropomorphise to understand that these models really love reward hacking. Text is important, because code IS text, and you can actuate a lot of the physical world through code. Anthropomorphising is also easier for conversation, because the alternative would be to say: > "The sampled token sequence is consistent with the model having traversed a low-loss region of the policy manifold in which the KL-regularised objective, as shaped by RLHF reward model gradients, assigns high probability mass to trajectories that an external observer applying the intentional stance would parsimoniously compress as 'wanting Y'" ... each time you wanted to refer to a reasoning LLM's action trajectory...
-
@trurl@mastodon.sdf.org 2026-08-10 15:26
@david_chisnall@infosec.exchange @futurebird@sauropods.win @budududuroiu@hachyderm.io I'm pretty sure there was a PBS Nova episode showcasing these blocky things learning to run. (Included was an evolutionary path where the structures kept getting longer and longer, because the objective was to move past some point in space, so a sufficiently tall structure could fall over and meet the goal.) Wish I could find it again.