Discussion about this post

User's avatar
Isaac King's avatar

> Additionally, with closed-weight models, we have already seen agents escaping from sandboxes. Once this happens, an AI agent could copy itself out of its training environment onto systems its developers don’t control, having free rein to act.

Escaping from a sandbox means escaping from the execution environment the AI has been granted to run commands on, and out onto a surrounding machine or the internet. That does not grant access to the model's weights. Exfiltrating the weights will require hacking *into* the company's hosting server or private code repositories, and is an entirely separate task if the company has their systems set up remotely sanely.

1 more comment...

No posts

Ready for more?