2026-04-29 17:40 UTC
Replies (2)
-
@danluu@mastodon.social 2026-04-29 17:45
@michael For this thing, someone tried auditing with an LLM about a month ago and found it was producing a ton of false positives. It's possible my triage process (nothing fancy, just getting the LLM to produce a realistic repro) on the top of the audit is the trick and the input could be either fuzzing or an audit, but my raw false positive rate before the triaging is lower than what this other person experienced, which maybe (?) helps the overall false positive rate.
-
@agwa@follow.agwa.name 2026-04-29 18:44
@michael Curious what model/harness you're using? I tried to use Claude Code and Opus 4.6 with the prompt in https://sockpuppet.org/blog/2026/03/30/vulnerability-research-is-cooked/ to audit some software that I nervously rely on, and it kept telling me I was violating their acceptable use policy.