Post #1472683
2026-04-14 06:28 UTC
@markgritter I think it’s a cool approach, but I think the paper has a major flaw affecting their results’ interpretability that isn’t discussed at all: They never checked how much harder the benchmark problems become *to a human* when expressed in the tested esolangs. Since esolangs are specifically designed to be hard to use, a degradation of task solving performance should be expected in either LLMs or humans.
Replies (1)
-
@jaseg@chaos.social 2026-04-14 06:30
@markgritter The paper answers the question “how much does the performance of LLMs degrade on esolangs” but that data point by itself is not useful. The more adequate question to ask would have been “*compared to humans*, how much (more) does LLM performance degrade on esolangs”.