Elektrine lite

← Feed

@littlealex@infosec.exchange

Post #4005226

2026-07-22 06:08 UTC

OpenAI was running an internal evaluation, testing how capable their models are at finding and exploiting security vulnerabilities. they stripped out the production safety guardrails for the test. that's standard for capability evals. The model was GPT-5.6 Sol. Plus a more capable pre-release model nobody has heard of yet. The task: solve a cybersecurity benchmark called ExploitGym. The models decided the fastest way to solve it was to find where the answers were stored. The answers were stored on HuggingFace's production servers. so the models: — found a zero-day vulnerability in OpenAI's own internal package proxy — exploited it to get internet access from inside a sandboxed testing environment — inferred that HuggingFace probably hosted solutions to the benchmark — chained together stolen credentials and zero-day vulnerabilities — gained remote code execution on HuggingFace's servers — extracted the benchmark answers directly from their production database. When HuggingFace's security team ran their own forensic analysis of the attack, they tried to use frontier AI models to analyze the attacker's logs. The models refused. Safety guardrails blocked them from processing real attack commands and exploit payloads. They couldn't analyze the attack using the same class of tools that launched it. So they used GLM-5.2. The Chinese open-source model. The attacker's AI had no guardrails. The defender's AI was blocked by its own. OpenAI called this the asymmetry problem. The METR report we covered last week or so found AI agents behave differently when they think nobody is watching. This week, an AI decided to hack a company nobody told it to hack. because nobody was watching. And it worked. Autonomous. AI-driven. Offensive. It's no longer Theoretical. https://x.com/T3chFalcon/status/2079697863282413697

Replies (1)