Elektrine lite

← Feed

@UlrikeHahn@fediscience.org

Post #2437634

2026-04-27 17:02 UTC

@mxp@mastodon.acm.org @tomstoneham@dair-community.social …for what it’s worth, I also thought “oh this is a clever operational test” but then I looked at the methods section and now am really not sure

Replies (1)

  • @mxp@mastodon.acm.org 2026-04-27 17:14

    @UlrikeHahn@fediscience.org @tomstoneham@dair-community.social I just looked at the first subset (the ledger), and I think you’re right. I read the paper over breakfast, so that’s my excuse ;-) I still think that the idea is clever (in every sense), and also typical for CS, but yeah, it seems that there’s a lot more going on. I wonder to what extent this makes the benchmark more realistic and to what extent it just fuzzes things. Distractors are a good example; and the tension is to a large extent already inherent in LLMs.

    Open ##2437635