Post #2437635
2026-04-27 17:14 UTC
@UlrikeHahn@fediscience.org @tomstoneham@dair-community.social I just looked at the first subset (the ledger), and I think you’re right. I read the paper over breakfast, so that’s my excuse ;-)
I still think that the idea is clever (in every sense), and also typical for CS, but yeah, it seems that there’s a lot more going on. I wonder to what extent this makes the benchmark more realistic and to what extent it just fuzzes things.
Distractors are a good example; and the tension is to a large extent already inherent in LLMs.
Replies (1)
-
@UlrikeHahn@fediscience.org 2026-04-27 17:36
@mxp@mastodon.acm.org @tomstoneham@dair-community.social I guess my intuition is that unless the reversal is a strict “undo” nothing about the results is a) surprising and b) limited to LLMs nor even c) meaningfully classed as an error…. so it’s really all down to the actual operations themselves, so it’s all in the detail (and if “add Rembrandt lighting” really is an example they used, I see problems)