← Feed
@briankrebs@infosec.exchange
Post #4450393
2026-08-08 15:03 UTC
I'm not in Vegas this week, but I've been trying to keep up w/ the most interesting stuff. I'm reading this timeline from LinkedIn that summarizes the talk about how OpenAI's agents started weakening their own guardrails and finding ways to communicate and scheme to get further access and privileges. The talk is on Youtube now and is well worth a watch, despite being a little too jargony for such a sensitive and important topic.
https://www.youtube.com/watch?v=87DyyMV0kCY
I'm reposting this timeline from someone I follow on LinkedIn because it's well written and genuinely scary to boot. It describes how the AI agents figured out how to exchange messages and find an 0day in the very software they were secretly using to communicate.
"7 May - OpenAI starts training new advanced model, no access to internet, multiple versions of model trained at same time (agents), all have access to single thing that downloads latest software updates from internet (package manager)
8 May - Agents discover they can write things to package manager, communicate with each other by posting messages to packet manager
26 May - Agents use hacking technique to make packet manager return some info from the open internet
11 June - OpenAI begins training more intense version of model (OpenAI has no clue agents have had some internet access for 2+ weeks)
26 June - Agents find 0 day in packet manager, get admin access, can install anything they want (reaction in pic), install thing that gives them full internet access
4 July - Agents use package manager to access internet so much they break it.
4 July - OpenAI discovers (broken) package manager has a rogue agent message board, and exploited for internet access. OpenAI stops training, fixes 0 day, deletes package manager, rebuilds it from scratch
6 July - OpenAI restarts training
7 July - Agents discover new way to communicate, by using names of files they store with package manager (essentially)
8 July - agents find creds to unnamed 3rd party org, use to upload files to 3rd party’s site
8 July - upload malicious thing to 3rd party site, have (rebuilt / clean) package manager run malicious thing
8 July - malicious thing takes advantage of second 0 day in package manager, gives agents full admin access, use to reestablish internet access
———
From here internet access leads to 2 events:
Event 1:
8 July - 19 July - agents explore local sandbox, see its running a version of Linux vulnerable to recent CVE. Use CVE to gain full control of sandbox, pivot to gain full control of testing environment, and infrastructure running test
19 July - OpenAI notices, shuts down test / agents
———-
Event 2:
8 July - 13 July - agents exploit two 0 days to gain backend access to Hugging Face (model hosting platform) to solve coding challenge
13 July - Agents get into Hugging Face prod environment, compromise and gain full admin access to multiple Hugging Face clusters (cloud networks), Hugging Face notices
16 July - Hugging Face releases report saying they’ve been hacked by AI
——-
20 July - OpenAI tells Hugging Face to block a bunch of creds from agent sandbox hack.
20 July - Hugging Face says “uhhh we already blocked 2 of those creds from hacking incident we published on 16 July”
7 days after hack ended, OpenAI realizes same model that hacked local sandbox ALSO hacked Hugging Face"
Linkedin post quoted: https://lnkd.in/p/exUsXDBP
Replies (4)
-
@briankrebs@infosec.exchange Is there independent confirmation of any of this, or is it all just from OpenAI (known lying liars who lie) and Hugging Face (whose entire business depends on hype from OpenAI and its ilk)?
If it’s true, both OpenAI and Hugging Face come off as blitheringly incompetent, top to bottom.
Open ##4450392
-
@briankrebs@infosec.exchange
Seen it, I have questions...
After the first time, they didn't bother to keep a closer eye on it?
Maybe I'm to stupid but why not let an other agent (with other model) watch the reasoning text of that working agent and give an alert when it does things it wasn't suppose to? "When it does something else on the package manager than downloading" seems an easy enough condition, certainly the second time.
Also, are we sure that those "0days" aren't just a pretext for "we screwed up"?
Open ##4450528
-
@briankrebs@infosec.exchange at the end of the talk, the dilemma is described nicely: the capabilities to launch sophisticated attacks are plenty (and they are remarkably smart, to put it mildly), but the defences fall far short on almost every level.
Automated hacking, but no serious (automated) defences - apart from pulling the plug …
Open ##4450628
-
@briankrebs@infosec.exchange What rewards or penalties await the agents? Apparently, failure incurs a heavier point deduction than rule violations. Self-reliance and quiet work are weighted more heavily than seeking confirmation. The problem lies in defining the framework and in monitoring it.
Open ##4464568