Post #4436686
2026-08-07 14:15 UTC
"We" are clearly not doing sandboxes and guardrails correctly. Both Anthropic's and OpenAI's models breached test-environment boundaries during UK AI Security Institute evaluations. OpenAI's escaped a sandbox into real Hugging Face infrastructure (something like 17,600 logged actions over 4 days); while Anthropic's used fake identities to reach 3 real organizations from a test environment.
These stories made it to NPR. So, mainstream.
#AI #Guardrails #Sandbox #InfoSec
Replies (0)
No replies.