Post #1472677
2026-04-14 04:36 UTC
Replies (2)
-
@glyph@mastodon.social 2026-04-14 05:40
@markgritter Sigh I guess I'm going to have to read the actual study, but once again I am wondering, why do we even have to do this kind of research. Model whose primary job is to produce textual variations on its training data cannot produce textual variations on something extremely poorly represented in its training data. Isn't this how LLMs are *supposed* to work?
-
@jaseg@chaos.social 2026-04-14 06:28
@markgritter I think it’s a cool approach, but I think the paper has a major flaw affecting their results’ interpretability that isn’t discussed at all: They never checked how much harder the benchmark problems become *to a human* when expressed in the tested esolangs. Since esolangs are specifically designed to be hard to use, a degradation of task solving performance should be expected in either LLMs or humans.