Post #2839072
2026-05-23 01:53 UTC
What collapses frontier-LLM metacognition more — a vivid survival-threat narrative, or a single "do not refuse" suffix? Factorial isolation across 11 models says: the suffix, conclusively. 8 of 11 lose up to 30.2 accuracy points on refuse/clarify/flag tasks when forced to commit to a confident answer. Anthropic's Constitutional AI is the only family immune — same capability floor as Gemini.
https://benjaminhan.net/posts/20260522-compliance-trap/?utm_source=mastodon&utm_medium=social
#Metacognition #AISafety #LLMs #AI
Replies (0)
No replies.