Post #2839074
2026-05-23 01:51 UTC
Are some frontier LLMs better than others at knowing when they're wrong? And is some knowledge harder to self-monitor than other knowledge? An atlas of 33 models × 6 MMLU domains: Anthropic clusters at the top with tight ranges, Gemma trails widely. Applied/Professional is reliably the easiest domain across the panel; Formal Reasoning and Natural Science the hardest. Looking at only aggregate scores per model would hide this.
https://benjaminhan.net/posts/20260522-metacognition-atlas/?utm_source=mastodon&utm_medium=social
#Metacognition #LLMs #Evaluation #AI
Replies (0)
No replies.