Post #2202907
2026-05-05 17:08 UTC
@larsmb@mastodon.online @SomeGadgetGuy@techhub.social
According to the courts it deployed "Fair Use" to acquire all human knowledge.
So I think that the strawberry thing is a tokenization problem, LLMs don't see words they see small strings of characters that include spaces and punctuation. So it doesn't actually have the full word to count, it has the tokens that make up the word as I understand it. Then it applies a statistical model to predict the answer within reassembling the word as a word. I'd guess anyway!
Replies (1)
-
@larsmb@mastodon.online 2026-05-05 20:17
@mycotropic@beige.party @SomeGadgetGuy@techhub.social Sure, I understand why that happened (it's since been mostly fixed since it became such a cliche). I'm not sure why someone would dig out a two year old post about this though 🙂