Elektrine lite

← Feed

@BenjaminHan@sigmoid.social

Post #2839067

2026-05-24 00:06 UTC

Given a problem queue and a token budget, can an LLM plan which to attempt, in what order, and how much to spend on each — before any execution feedback? TRIAGE tests 20 frontier and open-source LLMs. Most plan worse than random. Reasoning-trained modes systematically lose to standard ones. Even when shown its own per-problem budget, the best complier respects it on 37% of attempts. https://benjaminhan.net/posts/20260523-triage-metacognitive-control/?utm_source=mastodon&utm_medium=social #Paper #AI #LLMs #Metacognition #Evaluation #AgenticSystems

Replies (0)

No replies.