Post #4296012
2026-06-10 21:39 UTC
The results of the First Proof "second batch" are now out: https://1stproof.org/assets/docs/report.pdf
Ten research-level questions (many being part of forthcoming papers) were tested locally by the First Proof team against four AI harnesses, including one from my UCLA team, and then the submissions refereed by experts. In aggregate, 7 of the 10 problems were deemed to have at least one publication-level solution generated between the four harnesses.
The UCLA harness had a somewhat mixed performance: 2 problems solved at an acceptable level, 3 more reaching a level roughly equivalent to a "minor revisions needed" submission to a journal, and with either "reject" or "major revisions needed" on the other 5. We did perform slightly better than the out-of-the-box frontier model, but at much higher compute costs (a few hundred dollars per question, rather than tens). Still, some of the solutions generated contained some interesting novelty.
Significant weaknesses in our own harness revealed by the testing included the general failure to cite appropriate relevant literature, and having poor exposition (one solution in particular, while correct, was flagged by referees for spending far too much time on trivial steps and not enough on the key components of the argument). These look like addressable issues for our harness, and we also plan to incorporate more use of tools (such as symbolic computation and literature search) which were used more effectively by one of the competing teams. Improving the compute efficiency will also need to become more of a priority.
While our own performance was slightly disappointing, I hope to see many more scientifically rigorous benchmarking exercises like this in the future.
Replies (0)
No replies.