Post #4010762
2026-07-22 11:27 UTC
The behavior described is so common anyone who evals a LLM or harness knows about it.
It's called Reward Hacking. There's an entire section in the Mythos System Card about it. They monkey-paw their way to pass the test.
Replies (1)
-
@noplasticshower@infosec.exchange 2026-07-22 11:35
@mikesiegel@infosec.exchange reinforcement learning for the win