Post #2022547
2026-05-04 07:29 UTC
> The only task that didn’t degrade across most models was Python.
Yeah, after 20 cycles of unsupervised iteration on the task. Gemini 3.1 Pro doing as well as it did under that experiment setup is quite remarkable actually.
The paper does not show what you are arguing.
Replies (0)
No replies.