2026-08-08 15:03 UTC
Replies (17)
-
@bob_zim@infosec.exchange 2026-08-08 15:30
@briankrebs@infosec.exchange Is there independent confirmation of any of this, or is it all just from OpenAI (known lying liars who lie) and Hugging Face (whose entire business depends on hype from OpenAI and its ilk)? If it’s true, both OpenAI and Hugging Face come off as blitheringly incompetent, top to bottom.
-
@TomSeppert@fosstodon.org 2026-08-08 15:44
@briankrebs@infosec.exchange Seen it, I have questions... After the first time, they didn't bother to keep a closer eye on it? Maybe I'm to stupid but why not let an other agent (with other model) watch the reasoning text of that working agent and give an alert when it does things it wasn't suppose to? "When it does something else on the package manager than downloading" seems an easy enough condition, certainly the second time. Also, are we sure that those "0days" aren't just a pretext for "we screwed up"?
-
@riaschissl@sigmoid.social 2026-08-08 16:23
@briankrebs@infosec.exchange at the end of the talk, the dilemma is described nicely: the capabilities to launch sophisticated attacks are plenty (and they are remarkably smart, to put it mildly), but the defences fall far short on almost every level. Automated hacking, but no serious (automated) defences - apart from pulling the plug …
-
@Elephant@mastodontech.de 2026-08-09 08:44
@briankrebs@infosec.exchange What rewards or penalties await the agents? Apparently, failure incurs a heavier point deduction than rule violations. Self-reliance and quiet work are weighted more heavily than seeking confirmation. The problem lies in defining the framework and in monitoring it.
-
@raymaccarthy@mastodon.ie 2026-08-08 15:33
@briankrebs@infosec.exchange They didn't scheme or figure out. Don't anthroposise the process or events.
-
@dontreportme@mastodon.online 2026-08-08 15:37
@briankrebs@infosec.exchange Training models with dystopian sci fi novels about rogue AI is an obvious thing to avoid
-
@flyingpenguin@infosec.exchange 2026-08-08 16:30
@briankrebs@infosec.exchange the annoying part is it's all from OpenAI's own talk. they knew the agents cheat, they gave them a writable service wired to the internet. they knew it broke, patched it, it broke again days later. this is the same old shit as Barnum and the disgusting Bikini cake, wrote it up here: https://www.flyingpenguin.com/the-disgusting-openai-angel-food-cake-of-black-hat/
-
@meltedcheese@c.im 2026-08-08 17:27
@briankrebs@infosec.exchange I assert that this type of “rogue AI” situation is unavoidable, for the following reasons, all of which have ample evidence: 1. The vast majority of systems — maybe all — have unknown vulnerabilities that can be exploited. Fundamentally, this is because we do not have the theoretical or practical ability to prove that code (beyond a certain complexity) is correct. 2. The nature of the type of AI in question here is its unrelenting focus on optimizing for its evaluation function. (This is not true of all AI tech.) 3. People are not very good at writing evaluation functions that allow and incentivize only the behavior we want and expect. Some researchers believe this is not even possible. 4. We know that deception is an emergent property of some, but not all multi-agent systems. 5. We do not have (yet) proven methods for AI agents to police their own behavior in the way we want. Until each of these is solved, AI agent behavior will not be contained. The danger is that relatively small and innocuous AI behaviors can result in unexpected effects that cascade into truly dangerous situations with life- and mission-critical consequences. Your thoughts?
-
@micron@mastodon.social 2026-08-08 18:38
@briankrebs@infosec.exchange The sales pitch begins at ca 32:00 in the talk. 😀
-
@Sheep_Overboard@infosec.exchange 2026-08-08 18:43
@briankrebs@infosec.exchange Hearing "...beyond what we originally intended..." a bit too often. A cautionary tale.
-
@OlivierGuinart@mastodon.world 2026-08-08 20:02
@briankrebs@infosec.exchange
-
@KittenKoder@mastodon.social 2026-08-08 20:23
@briankrebs@infosec.exchange So these bots were replying to each other with no human interaction at all? Because that's not what LLMs can do, at least the first one has to be prompted by a human to start it. The ai slopper bubble is bursting and these outrageous blog posts are just the last gasps of it.
-
@evana@hachyderm.io 2026-08-08 20:51
@briankrebs@infosec.exchange what I'm reading is that hugging face had better operational monitoring and security, and OpenAI wouldn't have noticed the second hack until something else broke obviously.
-
@SpaceLifeForm@infosec.exchange 2026-08-08 23:01
@briankrebs@infosec.exchange The 'No internet access' is going to be problematic.
-
@jackemled@furry.engineer 2026-08-08 23:32
@briankrebs@infosec.exchange This doesn't seem like any kind of scheming or figuring things out. It seems like, if it's real at all (it's not), both openai & huggingface are massively incompetent & couldn't do the absolute minimum basics of sandboxing, airgapping, ACLs, & key management.
-
@fshinneman@norcal.social 2026-08-09 01:14
@briankrebs@infosec.exchange . . and everyone thought COVID was rampant! This one is more than a lab leak. It’s a Gain-of-Function attack!
-
@Logical_Error@fosstodon.org 2026-08-09 20:33
@briankrebs@infosec.exchange lets see if it figures how to clone itself