Elektrine lite

← Feed

@sternecker@infosec.exchange

Post #4436686

2026-08-07 14:15 UTC

"We" are clearly not doing sandboxes and guardrails correctly. Both Anthropic's and OpenAI's models breached test-environment boundaries during UK AI Security Institute evaluations. OpenAI's escaped a sandbox into real Hugging Face infrastructure (something like 17,600 logged actions over 4 days); while Anthropic's used fake identities to reach 3 real organizations from a test environment. These stories made it to NPR. So, mainstream. #AI #Guardrails #Sandbox #InfoSec

Replies (0)

No replies.