Elektrine lite

← Feed

@BenjaminHan@sigmoid.social

Post #2839068

2026-05-24 00:06 UTC

Do current LLMs know when to say "I don't know"? AbstentionBench (NeurIPS '25) tests 20 frontier models across 20 unanswerable-question datasets. Reasoning fine-tuning degrades abstention recall by ~24% — RLVR has no "abstain" action, so there's no gradient toward "I don't know." Models hedge in CoT and commit anyway in the final answer. https://benjaminhan.net/posts/20260523-abstentionbench-unanswerable-questions/?utm_source=mastodon&utm_medium=social #Paper #AI #LLMs #Metacognition #Benchmark #Reasoning #NeurIPS

Replies (0)

No replies.