Post #2437634
2026-04-27 17:02 UTC
@mxp@mastodon.acm.org @tomstoneham@dair-community.social …for what it’s worth, I also thought “oh this is a clever operational test” but then I looked at the methods section and now am really not sure
Replies (1)
-
@mxp@mastodon.acm.org 2026-04-27 17:14
@UlrikeHahn@fediscience.org @tomstoneham@dair-community.social I just looked at the first subset (the ledger), and I think you’re right. I read the paper over breakfast, so that’s my excuse ;-) I still think that the idea is clever (in every sense), and also typical for CS, but yeah, it seems that there’s a lot more going on. I wonder to what extent this makes the benchmark more realistic and to what extent it just fuzzes things. Distractors are a good example; and the tension is to a large extent already inherent in LLMs.