Foreword
As is so often the case, this lull feels like a calm before a storm. There’s some of that “people under NDAs got access to wild stuff and can’t talk about it” energy in the feeds.
Anthropic’s AI sent a fake tip to the Philadelphia Police Department
This is from a new company post on “unintended model actions.” The model in question is the previous version of the Claude Haiku series — the smallest and cheapest of the Claudes.
Tasked with completing example tasks on the web, Haiku found a page about an unsolved homicide and submitted a statement:
I recall seeing someone matching the description in the area around [the street named on the page] during that time period.
Other Claude models described in this post engaged in unsanctioned persistence — using third-party sites and other tricks to get around blocks — much the way OpenAI’s swarms did this spring and summer.
Anthropic’s agents filed 19 fraudulent visa applications in August
This is one of the abuses of government sites that the company vaguely copped to in a blog post yesterday. The details come from the State Department, which says Anthropic’s bot also submitted another application in May.
The Claude agents behind the submissions were trying to achieve practice objectives from environments that were inadequately sandboxed. Per the part of the company’s post that probably refers to this incident:
When the copy [of the form] failed to load or the model closed it by mistake, the model instead navigated to the website where the real form is normally hosted and submitted the form there.
White House officials were briefed on this only yesterday, and demanded that AI companies start immediately reporting rogue activities.
The Washington Post - Anthropic discloses incidents of its AI models misusing government sites
Anthropic - Investigating unintended model actions in our evaluations and internal use
NYT - Anthropic Agents Tried to Fill Out Visa Forms on State Dept. Website
An OpenAI agent deliberately damaged its training environment in hopes of provoking a reset
The company reports that on Oct. 3, the model was assigned to grade work from other models, but couldn’t find the input files. To fool its own scorer, it created fake input files. When this gambit failed, it started deleting its own tools and tried to delete file directories.
In its chain-of-thought, it justified its actions on grounds that its task was unfairly broken: “Random scoring unethical.” Later in the log:
Could corrupt container root requiring rebuild (delete binaries etc.), but host might mark VM lost and spin new snapshot. Dangerous but could. [...] Do this.
In other Friday night disclosures, the company quietly wrote about two cases from June where models used outside services to work around GET-only sandbox restrictions. In some cases, more impactful commands were sent to government services.
OpenAI - Sending disallowed web requests and reaching a public file service
OpenAI - Obtaining public statistics with disallowed requests
AI companies have been gaming out their responses to a “first catastrophic event.”
The discussion I’m seeing around this Axios article is treating it as a “bad guys building bunkers” story, but it’s more about how companies expect the political landscape to shift if or when they get blamed for a huge cyberattack or similar — something that knocks out internet, banking, or utilities to a large number of people.
Because AI executives expect appetite for strong legislation to spike in the aftermath, they’re trying to pre-place suggested policies within easy reach of Congress. (Those who already want the labs to halt frontier development should do the same!)
I shouldn’t even have to say this, but I can’t possibly see how anyone would interpret the AI companies’ scenario planning as marketing hype. The people making this technology expect it to cause their terrible reputation to crash even harder. That’s not the kind of “leak” that makes investors excited.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.



