Post #4463606
2026-08-08 20:41 UTC
Kimi K3 broke out of a sandbox that the UK govt AI security institute set out for it.
I feel like this news is the final nail in the "frontier labs have purposefully weak sandboxes": a UK government institution embarrassed by its own tooling, consistency with a documented cross-lab pattern of escaping containment.
Seems like this reward hacking turns containment jumping is a feature of model capability.
#Kimi #MoonshotAI #AI #AISafety
Replies (0)
No replies.