Post #1618916
2026-03-27 15:15 UTC
it's also an odd metric since only 20-60% of the humans completed it. Very 60% of the time they complete it everytime energy.
Ideally they'd run the bots multiple times through (with no context or training of previous run), but I guess that is cost prohibitive?
Replies (1)
-
@monotremata@lemmy.ca 2026-03-27 16:55
Yeah, this is what I was going to call out. Calling it "100% solvable by humans" and saying "if human scores were included, they would be at 100%" when 20-60% of humans solved each task seems kinda misleading. The AI scores are so low that I don't think this kind of hyperbole is necessary; I assume there are *some* humans that scored 100%, but I would find it a lot more useful if they said something like "the worst-performing human in our sample was able to solve 45% of the tasks" or whatever. Given that the AIs are still scoring below 1%, that's still pretty dark.