Post #3942303
2026-06-21 19:38 UTC
An LLM is basically lossy compression for the google cache.
"Training" is the data compression. "Tokens" are decompressing a zoomed in fractal slice of morphed together training data, with lots of brownian motion jiggling to make it work and a ton of autocorrect applied to the jpeg artifacts.
"Agents" feed the output back into the input a bunch of times (adding zillions of prompts to the decompressor's starting dictionary) to take lots of slices until the metrics add up to something sellable.
Replies (1)
-
@landley@mstdn.jp 2026-06-21 19:44
This is the data version of growing snowflakes in a lab. No two alike, results unreproducible. https://en.wikipedia.org/wiki/Simulated_annealing "Training" is how you put new data into a model. ChatGPT-3 (circa 2020) didn't know about Barbenheimer, let alone the impact of the Hormuz quagmire on gas prices. You can never STOP training, the same way news sites can never stop putting out new stories, so AI business models separating that out are lying. This is why https://meta.wikimedia.org/wiki/Wikipedia_on_CD/DVD never really took off.