Post #1594712
2026-04-23 19:45 UTC
on a serious note: designing benchmarks is hard.
the consensus has been that creating verifiable benchmarks is surprisingly difficult and the ones that are difficult (like HLE) only get included in these benchmark images when new higher scores are achieved.
its just soooo nice seeing a 99% score on a tool calling benchmark which literally just tests for if the model can generate proper json
people are trying their best designing benchmarks.
Replies (1)
-
@TotallynotJessica@lemmy.blahaj.zone 2026-04-24 00:34
The best measure for AI is the productivity and accuracy of the work people do with the models. It doesn’t matter if the tech is good at anything if people don’t use it properly. Just like any tool, there are right and wrong ways to use them.