Elektrine lite

← Feed

@0xabad1dea@infosec.exchange

Post #4208531

2026-07-29 14:46 UTC

I'm going to mute this thread because I spent an unhealthy amount of time attempting to resolve a claim that one of my statements was wrong, with the conclusion "I really don't think I'm wrong, but can't conclusively prove it with a smoking gun quote". This is a summary of the dispute: 1) my claim was that the OP of the LLM-generated buggy proof did not know it was buggy when they posted it, and was misled by the LLM but was acting in good faith. 2) Someone counterclaims that the OP of the buggy proof knew perfectly well that it was buggy when they posted it (because they are an expert on theorem provers in general), yet chose not to disclose this up-front and let everyone else figure it out. I think 2) sounds like a rather dickish thing to do, but also, going over the github issues, community threads etc, everything reads to me as if the OP sincerely did not realize it was buggy when they posted it, but is gladly cooperating with figuring out and fixing all the bugs uncovered so this won't happen again. HOWEVER, if you have proof that OP knew it was buggy when they posted it and chose not to say anything up front, feel free to link it and others can check the replies. I don't think it materially changes the point that an LLM can come up with solutions that really seem like they check out but are relying on devastating bugs in other software that you won't spot.

Replies (2)

  • @0xabad1dea@infosec.exchange many many years ago, i filed quite a few issues against coqchk, Rocq's proof checker. it had a lot of unsafe code. (I wanted to put theorems on the Ethereum blockchain so that they self-reward the people who find them.) anyway, although my issues were eventually resolved and coqchk is better as a result (I filed some of the fix PRs!), i was eventually pointed to a previous experiment, which basically concluded with humans proving False quite a few times

    Open ##4208534

  • @0xabad1dea@infosec.exchange 2026-08-02 10:39

    Apologies if you see this post twice in your timeline, however, since "followers only" metadata got added to the reply chain and I can't remove it, my reply has DRM on it that could prevent people from seeing the followup, so here is my reply to the person who originally brought this up: ----- I apologize that it took me a few days to reply to you: I had muted the thread because it got way too much Engagement and it wasn't good for my mennal helths, and I came back when I felt better to see if there was ever a demonstrable verdict on the disagreement about this detail. I agree that the post you linked is pretty convincing that the OP who posted the buggy proof was already aware of the bug and chose not to say anything, as a "funny" way to disclose a bug. I think that was, in fact, dickish behavior: if you know it's a bug, then it should be clearly labeled as such even if it's formulated in a humorous way. I do not think it materially alters the point of my post at all, which is that LLMs can trick you by exploiting bugs to falsify results. All it changes is that OP, specifically, was already aware that the LLM was trying to trick them, and chose to pass it on unaltered and let other people puzzle it out instead of filing a clear bug report, with the nature of the prank being clarified later. But I will boost this and edit the original post to add this clarifying detail.

    Open ##4333743