Elektrine lite

← Feed

@mxp@mastodon.acm.org

Post #1722157

2026-04-27 07:15 UTC

“In this work, we conduct a large-scale simulation of how users might delegate work to LLMs across 52 professional domains. We find that current LLMs are unreliable delegates: even frontier models corrupt an average of 25% of document content over long workflows, with sparse but severe errors that silently compound over time.” Good to see the issue addressed explicitly, even though the results aren’t surprising—why would anyone expect LLMs to be reliable!? https://arxiv.org/abs/2604.15597

Replies (2)

  • @mxp@mastodon.acm.org 2026-04-27 07:16

    Rhetorical question, obviously.

    Open ##2437621

  • @UlrikeHahn@fediscience.org 2026-04-27 10:58

    @mxp@mastodon.acm.org @tomstoneham@dair-community.social interesting paper….am a bit perplexed by the reversability constraint as implemented: instructions such as “add Rembrandt style lighting” and “remove Rembrandt style lighting” seem too vague to me to guarantee uniqueness?

    Open ##2437622