Elektrine lite

← Feed

@nolan@toot.cafe

2026-04-08 02:34 UTC

I have an incredibly stupid solution for reviewing vibe-coded PRs: - Run Cursor Bugbot - Run Claude - Run Codex - Ask each of them to find functional bugs - Ask another Claude to review the 3, check the code, and rule out hallucinations Sometimes a true bug will be found by all 3, sometimes only by 1 (often by Codex for some reason). It finds incredibly subtle bugs, and avoids the agent fixing a bug only to introduce a new one. It sometimes takes a few rounds, but it's surprisingly effective.

Replies (2)

  • @nolan@toot.cafe 2026-04-08 02:36

    This is essentially inspired by this article, which was a lot more scientific than my methods: https://milvus.io/blog/ai-code-review-gets-better-when-models-debate-claude-vs-gemini-vs-codex-vs-qwen-vs-minimax.md The weirdest part is that originally I tried to write a bash script to do the above, and it was too flaky and error-prone. Then I just wrote a Claude skill, so effectively the script is markdown. (The meta-reviewer is the main agent, the Claude reviewer is the sub-agent.) This proved much more resilient.

    Open ##2619465

  • @njoseph@social.masto.host 2026-04-08 03:25

    @nolan@toot.cafe Have you considered rejecting vibe-coded PRs? In this approach, the maintainer spends more than the original vibe-coder. You can save yourself some tokens/money/time by just rejecting the PR.

    Open ##2619469