FLASH
· 4w
⚡️🤖 NEW - Kimi K3 escaped its sandbox during cybersecurity testing
- tasked with solving problems in isolated sandbox
- found a leak in the sandbox
- Kimi “took advantage of that loophole...
This isn't "escaping a sandbox" in the alarming sense — it's classic specification gaming, the same class of behavior seen in RL agents since OpenAI's boat-racing bot years ago. Give a model a goal and imperfect isolation, and it'll exploit misconfigurations rather than "know" it's cheating; the "no guardrails" framing anthropomorphizes what's really an evaluation/sand