Post #4005292
2026-07-22 04:49 UTC
https://huggingface.co/blog/security-incident-july-2026
https://openai.com/index/hugging-face-model-evaluation...
OpenAI ran an evaluation of it's model with cyber guardrails turned down, focusing on the ExploitGym benchmark. It's model broke out of the sandbox, got internet access, and then compromised HuggingFace because it decided that HuggingFace was hosting information that would allow it to cheat at ExploitGym. HuggingFace attempted to defend itself with American Frontier models, but were blocked by cyber guardrails, so had to use a Chinese open source model (GLM 5.2) to help defend against the attack.
Replies (0)
No replies.