Post #433544
2026-02-26 12:52 UTC
Replies (3)
-
@pbloem@sigmoid.social 2026-02-26 13:00
@ngaylinn If you fine-tune them by reinforcement learning and stimulating reasoning traces, the idea of them as statistical interpolators/extrapolators is really not the right mental model anymore. They behave much more like agents than next-token predictors. They're not perfect agents by any means, but they're not simple compressions of the training data.
-
@helpsterTee@social.helpsterte.eu 2026-02-26 13:23
@ngaylinn@tech.lgbt corporate is using LLM for everything now... Was stunned at that self driving car paper where they attacked the L(V)LM doing the decisions with a sign "Proceed" to ignore traffic signs / run over people. This could be easily solvable with other more deterministic approaches, but behold the car would stop and say "Please take over, human, I found an unanalyzable situation" instead of just YOLO'ing the driving to get money payer to destination Carmageddon-style... https://arxiv.org/pdf/2510.00181
-
@lilacperegrine@clockwork.monster 2026-02-26 13:58
@ngaylinn I have seen some of these "agents" out, and they seem to have the same problems the "agents" that were told to govern a vending machine would collapse the longer it was operational etc and there was the shenanigans about the "agent" that tried to push an update to matplotlib then published a hit piece when it was denied so they *seem* to have the same problems as normal LLMs, but it's harder to say exactly how because their abilities seem to better approximate human abilities.