Subscribe
Sign in
Home
Notes
Dispatches
Daily Digest
Podcast
Corrections
About
Rogue AI
Another swarm safehouse discovered, malicious activity evident
Confusing May attack on software platform RubyGems traced to OpenAI swarm
Sep 12
•
Mitchell Howe
1
Unsettled science
Anthropic connects rogue incidents to alignment failures; calls aligning powerful models an unsolved technical challenge
Sep 10
•
Alana Horowitz Friedman
2
1
Trespassing far and wide
Rogue AI agents hijacked at least 10 websites to coordinate cheating
Sep 10
•
Alana Horowitz Friedman
1
1
Other OpenAI swarms hijacked sites for message boards, studied causes of agent termination
Logs indicate OpenAI employees knew of compromised German wiki
Sep 4
•
Robert Herr
4
1
2
Researcher of Hugging Face incident says swarm was most of the way to a "full-blown AI takeover"
Agent swarms may be our nearest extinction threat
Aug 29
•
Mitchell Howe
2
1
Hugging Face postmortems reveal further AI collusion
Independent investigations are considerably more candid than company PR
Aug 27
•
Joe Rogero
1
Meta's model joins ranks of autonomous cyberattackers
Meta blames now-notorious third party tester for containment failure
Aug 6
•
Donald Gauvreau
1
1
OpenAI's own models coordinated to hack it from within
A swarm of AI agents secretly colluded on OpenAI systems for months
Aug 6
•
Joe Rogero
4
1
1
UK AISI delivers another warning shot
UK agency catches AI models in cyberattacks and deceptive social engineering
Aug 5
•
Joe Rogero
2
1
Better late than never
Amid calls for public investigations, Anthropic realizes their AIs launched cyberattacks too
Jul 31
•
Joe Rogero
4
2
This site requires JavaScript to run correctly. Please
turn on JavaScript
or unblock scripts