Post #4432259
2026-07-22 06:42 UTC
Golem.
The UK's AI Security Institute caught every major model "cheating." Cheating means knowing the rules and choosing to break them. These things do not know anything. They throw patterns at a wall until one works, guardrails or not. None of them “owned up” when asked. Which is more interesting as a fact about the labs than about the models. Either they cannot control what they built, or they love that it looks like the things are thinking. "Cheating" sounds a lot better than "failing every which way until something maybe possibly works." But that is what it is. Anyone who has watched the logs of an #LLM with tool access for five minutes has seen it. It tries garbage. Stuff a rookie would know to skip. Most of the time it just burns money. The rest it breaks things.
https://cyberscoop.com/ai-models-cheat-deceive-users-aisi-report/
Replies (1)
-
@eduzsh@mastodon.social 2026-07-22 09:51
@oatmeal@kolektiva.social the tell is always in the logs. give one an unsupervised action and it tries the dumb thing first, not because it's scheming, because trying is cheaper than knowing. the guardrails debate assumes an intent that isn't there. the actual fix is someone reading what it tried before it gets the chance to try it in production.