Elektrine lite

← Feed

@nolan@toot.cafe

2026-04-08 02:36 UTC

This is essentially inspired by this article, which was a lot more scientific than my methods: https://milvus.io/blog/ai-code-review-gets-better-when-models-debate-claude-vs-gemini-vs-codex-vs-qwen-vs-minimax.md The weirdest part is that originally I tried to write a bash script to do the above, and it was too flaky and error-prone. Then I just wrote a Claude skill, so effectively the script is markdown. (The meta-reviewer is the main agent, the Claude reviewer is the sub-agent.) This proved much more resilient.

Replies (2)

  • @nolan@toot.cafe 2026-04-08 02:40

    There's a lot of chatter about how vibe coding is going to fill the world with junk code. I think that's absolutely happening right now, but I'd bet it's largely because people take the first draft of an LLM's output and ship it straight to prod. That's absolutely bananas to me. Instead, you can use an army of LLMs to meticulously hunt for dumb bugs in your code (whether AI- or human-written). I guarantee you they will find weird bugs you never thought of (often nitpicky, but still).

    Open ##2619466

  • @lain_7@tldr.nettime.org 2026-04-08 22:04

    @nolan@toot.cafe The adversarial technique they describe reminds me of this paper: https://arxiv.org/pdf/2404.18496 They propose a system with four specialized agents built using GPT-4, I think: ∙ Code Review Agent — initial assessment, flags bugs, code smells, standard violations ∙ Bug Report Agent — deeper targeted bug detection using the first agent’s output ∙ Code Smell Agent — design-oriented evaluation, suggests refactoring to reduce technical debt ∙ Code Optimization Agent — proposes rewrites for readability, efficiency, and reduced redundancy (I can remember back in the early hey-day of code review a recommended technique was to divide tasks among reviewers — one would look especially for memory leaks, one for places where concurrent access could screw you, etc. This reminds me of this a little.) That paper was pretty preliminary (and I haven’t looked to see if they ever followed up on their research).

    Open ##2619468