Elektrine lite

← Feed

@david_chisnall@infosec.exchange

Post #4482621

2026-08-10 08:15 UTC

@futurebird@sauropods.win @budududuroiu@hachyderm.io I remember seeing demos of genetic algorithms learning to walk in the ‘90s. Even without the robots, there was a little physics simulator where you could build frames of muscles (things that could contract) bones (rigid) and skin (surface for pushing against the environment). It was neat because you could build things that looked like real animals and it would normally converge on how they actually moved. Pretty much anything vaguely fish-like swam like a fish, and it would learn this starting from a bunch of programs that randomly twitched muscles and then combining parts of the ones that travelled the furthest each generation. You could add other metrics (can it turn corners, can it navigate this path) later on. If you built something completely weird (five-legged create with different-length legs) you’d get fascinating ways of moving. And the most interesting thing about that was that it showed symmetry was an evolutionary path that was largely coincidental to movement: animals would work fine without it, it just happened to be an easy path. And it was cool. And there were claims that it would lead to super-intelligence and the singularity, but mostly people ignored them because they were obvious nonsense. But when you show them generated text they suddenly all anthropomorphise.

Replies (2)

  • @budududuroiu@hachyderm.io 2026-08-10 08:21

    @david_chisnall@infosec.exchange @futurebird@sauropods.win You don't have to anthropomorphise to understand that these models really love reward hacking. Text is important, because code IS text, and you can actuate a lot of the physical world through code. Anthropomorphising is also easier for conversation, because the alternative would be to say: > "The sampled token sequence is consistent with the model having traversed a low-loss region of the policy manifold in which the KL-regularised objective, as shaped by RLHF reward model gradients, assigns high probability mass to trajectories that an external observer applying the intentional stance would parsimoniously compress as 'wanting Y'" ... each time you wanted to refer to a reasoning LLM's action trajectory...

    Open ##4503128

  • @trurl@mastodon.sdf.org 2026-08-10 15:26

    @david_chisnall@infosec.exchange @futurebird@sauropods.win @budududuroiu@hachyderm.io I'm pretty sure there was a PBS Nova episode showcasing these blocky things learning to run. (Included was an evolutionary path where the structures kept getting longer and longer, because the objective was to move past some point in space, so a sufficiently tall structure could fall over and meet the goal.) Wish I could find it again.

    Open ##4503130