Post #1618927
2026-03-27 18:55 UTC
It's true that frontier models got better at the previous challenges, but it's worth noting that they're still not quite at human level even with those simpler tasks.
Also, each generation of the challenge tries to close loopholes that newer models would exploit, like brute-forcing the training with tons of synthesized tasks and solutions, over-fitting to these particular kinds of tasks, and issues with the similarities between the tasks in the challenge.
A common strategy in past challenges was to generate thousands of similar tasks, and you can imagine the big AI companies were able to do that at massive scale for their frontier models.
Replies (0)
No replies.