Elektrine lite

← Feed

@BenjaminHan@sigmoid.social

Post #2839074

2026-05-23 01:51 UTC

Are some frontier LLMs better than others at knowing when they're wrong? And is some knowledge harder to self-monitor than other knowledge? An atlas of 33 models × 6 MMLU domains: Anthropic clusters at the top with tight ranges, Gemma trails widely. Applied/Professional is reliably the easiest domain across the panel; Formal Reasoning and Natural Science the hardest. Looking at only aggregate scores per model would hide this. https://benjaminhan.net/posts/20260522-metacognition-atlas/?utm_source=mastodon&utm_medium=social #Metacognition #LLMs #Evaluation #AI

Replies (0)

No replies.