Post #2281796
2026-05-07 11:42 UTC
"We evaluate 9 LMs and find that none fully resolve any task, with the best model passing 95% of tests on only 3% of tasks."
RE: https://social.lansky.name/@hn50/116532766804414096
Replies (0)
No replies.