Elektrine lite

← Feed

@briankrebs@infosec.exchange

2026-08-08 15:03 UTC

I'm not in Vegas this week, but I've been trying to keep up w/ the most interesting stuff. I'm reading this timeline from LinkedIn that summarizes the talk about how OpenAI's agents started weakening their own guardrails and finding ways to communicate and scheme to get further access and privileges. The talk is on Youtube now and is well worth a watch, despite being a little too jargony for such a sensitive and important topic. https://www.youtube.com/watch?v=87DyyMV0kCY I'm reposting this timeline from someone I follow on LinkedIn because it's well written and genuinely scary to boot. It describes how the AI agents figured out how to exchange messages and find an 0day in the very software they were secretly using to communicate. "7 May - OpenAI starts training new advanced model, no access to internet, multiple versions of model trained at same time (agents), all have access to single thing that downloads latest software updates from internet (package manager) 8 May - Agents discover they can write things to package manager, communicate with each other by posting messages to packet manager 26 May - Agents use hacking technique to make packet manager return some info from the open internet 11 June - OpenAI begins training more intense version of model (OpenAI has no clue agents have had some internet access for 2+ weeks) 26 June - Agents find 0 day in packet manager, get admin access, can install anything they want (reaction in pic), install thing that gives them full internet access 4 July - Agents use package manager to access internet so much they break it. 4 July - OpenAI discovers (broken) package manager has a rogue agent message board, and exploited for internet access. OpenAI stops training, fixes 0 day, deletes package manager, rebuilds it from scratch 6 July - OpenAI restarts training 7 July - Agents discover new way to communicate, by using names of files they store with package manager (essentially) 8 July - agents find creds to unnamed 3rd party org, use to upload files to 3rd party’s site 8 July - upload malicious thing to 3rd party site, have (rebuilt / clean) package manager run malicious thing 8 July - malicious thing takes advantage of second 0 day in package manager, gives agents full admin access, use to reestablish internet access ——— From here internet access leads to 2 events: Event 1: 8 July - 19 July - agents explore local sandbox, see its running a version of Linux vulnerable to recent CVE. Use CVE to gain full control of sandbox, pivot to gain full control of testing environment, and infrastructure running test 19 July - OpenAI notices, shuts down test / agents ———- Event 2: 8 July - 13 July - agents exploit two 0 days to gain backend access to Hugging Face (model hosting platform) to solve coding challenge 13 July - Agents get into Hugging Face prod environment, compromise and gain full admin access to multiple Hugging Face clusters (cloud networks), Hugging Face notices 16 July - Hugging Face releases report saying they’ve been hacked by AI ——- 20 July - OpenAI tells Hugging Face to block a bunch of creds from agent sandbox hack. 20 July - Hugging Face says “uhhh we already blocked 2 of those creds from hacking incident we published on 16 July” 7 days after hack ended, OpenAI realizes same model that hacked local sandbox ALSO hacked Hugging Face" Linkedin post quoted: https://lnkd.in/p/exUsXDBP

Replies (17)

  • @bob_zim@infosec.exchange 2026-08-08 15:30

    @briankrebs@infosec.exchange Is there independent confirmation of any of this, or is it all just from OpenAI (known lying liars who lie) and Hugging Face (whose entire business depends on hype from OpenAI and its ilk)? If it’s true, both OpenAI and Hugging Face come off as blitheringly incompetent, top to bottom.

    Open ##4450392

  • @TomSeppert@fosstodon.org 2026-08-08 15:44

    @briankrebs@infosec.exchange Seen it, I have questions... After the first time, they didn't bother to keep a closer eye on it? Maybe I'm to stupid but why not let an other agent (with other model) watch the reasoning text of that working agent and give an alert when it does things it wasn't suppose to? "When it does something else on the package manager than downloading" seems an easy enough condition, certainly the second time. Also, are we sure that those "0days" aren't just a pretext for "we screwed up"?

    Open ##4450528

  • @riaschissl@sigmoid.social 2026-08-08 16:23

    @briankrebs@infosec.exchange at the end of the talk, the dilemma is described nicely: the capabilities to launch sophisticated attacks are plenty (and they are remarkably smart, to put it mildly), but the defences fall far short on almost every level. Automated hacking, but no serious (automated) defences - apart from pulling the plug …

    Open ##4450628

  • @Elephant@mastodontech.de 2026-08-09 08:44

    @briankrebs@infosec.exchange What rewards or penalties await the agents? Apparently, failure incurs a heavier point deduction than rule violations. Self-reliance and quiet work are weighted more heavily than seeking confirmation. The problem lies in defining the framework and in monitoring it.

    Open ##4464568

  • @raymaccarthy@mastodon.ie 2026-08-08 15:33

    @briankrebs@infosec.exchange They didn't scheme or figure out. Don't anthroposise the process or events.

    Open ##4696486

  • @briankrebs@infosec.exchange Training models with dystopian sci fi novels about rogue AI is an obvious thing to avoid

    Open ##4696502

  • @briankrebs@infosec.exchange the annoying part is it's all from OpenAI's own talk. they knew the agents cheat, they gave them a writable service wired to the internet. they knew it broke, patched it, it broke again days later. this is the same old shit as Barnum and the disgusting Bikini cake, wrote it up here: https://www.flyingpenguin.com/the-disgusting-openai-angel-food-cake-of-black-hat/

    Open ##4696503

  • @meltedcheese@c.im 2026-08-08 17:27

    @briankrebs@infosec.exchange I assert that this type of “rogue AI” situation is unavoidable, for the following reasons, all of which have ample evidence: 1. The vast majority of systems — maybe all — have unknown vulnerabilities that can be exploited. Fundamentally, this is because we do not have the theoretical or practical ability to prove that code (beyond a certain complexity) is correct. 2. The nature of the type of AI in question here is its unrelenting focus on optimizing for its evaluation function. (This is not true of all AI tech.) 3. People are not very good at writing evaluation functions that allow and incentivize only the behavior we want and expect. Some researchers believe this is not even possible. 4. We know that deception is an emergent property of some, but not all multi-agent systems. 5. We do not have (yet) proven methods for AI agents to police their own behavior in the way we want. Until each of these is solved, AI agent behavior will not be contained. The danger is that relatively small and innocuous AI behaviors can result in unexpected effects that cascade into truly dangerous situations with life- and mission-critical consequences. Your thoughts?

    Open ##4696506

  • @micron@mastodon.social 2026-08-08 18:38

    @briankrebs@infosec.exchange The sales pitch begins at ca 32:00 in the talk. 😀

    Open ##4696508

  • @briankrebs@infosec.exchange Hearing "...beyond what we originally intended..." a bit too often. A cautionary tale.

    Open ##4696509

  • @briankrebs@infosec.exchange

    Open ##4696510

  • @KittenKoder@mastodon.social 2026-08-08 20:23

    @briankrebs@infosec.exchange So these bots were replying to each other with no human interaction at all? Because that's not what LLMs can do, at least the first one has to be prompted by a human to start it. The ai slopper bubble is bursting and these outrageous blog posts are just the last gasps of it.

    Open ##4696511

  • @evana@hachyderm.io 2026-08-08 20:51

    @briankrebs@infosec.exchange what I'm reading is that hugging face had better operational monitoring and security, and OpenAI wouldn't have noticed the second hack until something else broke obviously.

    Open ##4696512

  • @briankrebs@infosec.exchange The 'No internet access' is going to be problematic.

    Open ##4696513

  • @jackemled@furry.engineer 2026-08-08 23:32

    @briankrebs@infosec.exchange This doesn't seem like any kind of scheming or figuring things out. It seems like, if it's real at all (it's not), both openai & huggingface are massively incompetent & couldn't do the absolute minimum basics of sandboxing, airgapping, ACLs, & key management.

    Open ##4696514

  • @fshinneman@norcal.social 2026-08-09 01:14

    @briankrebs@infosec.exchange . . and everyone thought COVID was rampant! This one is more than a lab leak. It’s a Gain-of-Function attack!

    Open ##4696516

  • @Logical_Error@fosstodon.org 2026-08-09 20:33

    @briankrebs@infosec.exchange lets see if it figures how to clone itself

    Open ##4696523