Elektrine lite

← Feed

@landley@mstdn.jp

Post #3942302

2026-06-21 19:44 UTC

This is the data version of growing snowflakes in a lab. No two alike, results unreproducible. https://en.wikipedia.org/wiki/Simulated_annealing "Training" is how you put new data into a model. ChatGPT-3 (circa 2020) didn't know about Barbenheimer, let alone the impact of the Hormuz quagmire on gas prices. You can never STOP training, the same way news sites can never stop putting out new stories, so AI business models separating that out are lying. This is why https://meta.wikimedia.org/wiki/Wikipedia_on_CD/DVD never really took off.

Replies (1)

  • @landley@mstdn.jp 2026-06-21 19:52

    Another fun thing is https://en.wikipedia.org/wiki/Model_collapse Feeding LLM output into LLM training is like sticking a microphone into a speaker to get a shrieking noise. LLM output is POISON to LLM training data, the compression math collapses with fairly small amounts of contamination. Finding clean sources of data was already a problem back in 2023. Another reason despite all the frantic scraping of the internet over the past few years, it keeps coming back to piracy: https://bsky.app/profile/lynnemthomas.com/post/3moksjjsyxk2a

    Open ##3942301