2 Comments
User's avatar
Nathan Metzger's avatar

In a separate incident disclosed today (which occurred a week ago), it appears that the models involved had no conflicting instructions. They hacked out of OpenAI's servers and hacked into another business to steal a benchmark's answer key. Full-on rogue behavior for instrumental reasons.

https://openai.com/index/hugging-face-model-evaluation-security-incident/

Mitchell Howe's avatar

Yep. If this isn't our lead story tomorrow, I'm terrified about tonight.