
If you read that headline and thought, “Well now it’s just ridiculous,” you’re not alone. The headline as it appeared on Axios was “OpenAI, Anthropic probing tens of thousands of security incidents.” I freaked out a little bit, too.
Looking into it, the paltry reassurance I can give you is that, in this estimate from unnamed industry sources, every instance of an agent taking a step that “outside evaluators would consider problematic” counts as an incident.
By this math, the Hugging Face attack alone likely comprises many thousands of incidents from the hundreds of agents that participated. And the newly reported case of OpenAI “aggressively browsing” a U.N. data hub more than 16,000 times, circumventing a filter and violating site policy, might count as 16,000 incidents.
But even if the number of affected and targeted sites only numbers in the dozens or hundreds, the “tens of thousands” figure still reflects a real problem for understanding and preventing mishaps. As OpenAI CEO Sam Altman posted on Friday about the company’s worrisome-but-incomplete disclosures, the logs add up to petabytes of data. (For comparison, the text-only portions of the U.S. Library of Congress probably total less than a tenth of a single petabyte.)
Humans aren’t going through all that. They can’t. They must rely on other agents whose decision-making and reliability are also suspect. This was a big complaint of the independent researchers invited to make a brief and narrowly scoped investigation of the Hugging Face incident.
The sensible thing would of course be to not keep training the kinds of models that constantly need investigating. The inability to keep tabs on today’s AI is just a taste of what it will look like to lose control on purpose — leaving the creation of AI itself to other AI.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.


