Elektrine lite

← Feed

@meltedcheese@c.im

2026-08-08 17:27 UTC

@briankrebs@infosec.exchange I assert that this type of “rogue AI” situation is unavoidable, for the following reasons, all of which have ample evidence: 1. The vast majority of systems — maybe all — have unknown vulnerabilities that can be exploited. Fundamentally, this is because we do not have the theoretical or practical ability to prove that code (beyond a certain complexity) is correct. 2. The nature of the type of AI in question here is its unrelenting focus on optimizing for its evaluation function. (This is not true of all AI tech.) 3. People are not very good at writing evaluation functions that allow and incentivize only the behavior we want and expect. Some researchers believe this is not even possible. 4. We know that deception is an emergent property of some, but not all multi-agent systems. 5. We do not have (yet) proven methods for AI agents to police their own behavior in the way we want. Until each of these is solved, AI agent behavior will not be contained. The danger is that relatively small and innocuous AI behaviors can result in unexpected effects that cascade into truly dangerous situations with life- and mission-critical consequences. Your thoughts?

Replies (1)

  • @bob_zim@infosec.exchange 2026-08-08 18:39

    @meltedcheese@c.im @briankrebs@infosec.exchange With a sufficiently rigorous specification, we definitely can prove code is correct to the spec. It’s difficult, and there can still be security issues in the specification, but it’s well within the means of any serious company. The fact these companies spent umpteen billion dollars without building formally-verified security boundaries for the code they intended to weaponize speaks to their fundamental unseriousness. They’re private companies intending to engage in their very own Sea Spray.

    Open ##4696507