Post #2700257
2026-05-15 16:37 UTC
"We evaluate 9 LMs and find that none fully resolve any task, with the best model passing 95% of tests on only 3% of tasks. Models favor monolithic, single-file implementations that diverge sharply from human-wrintten code."
https://arxiv.org/html/2605.03546v1
Replies (0)
No replies.