Starting to stop?
OpenAI voluntary slowdown, legislative Whac-A-Mole, Chinese AI steps, and more
In this issue:
Slowing down research to enhance security - OpenAI tells the White House it intends to slow AI development
FTC proposes new rules on AI bias - Suggested regulation would punish undisclosed misleading outputs, but terms are worryingly vague
America is playing legislative Whac-A-Mole - Standard measures are struggling to keep up with widespread disruption
Small leaps forward - China gains market share, scales their models and joins the infamous AI escape club
Dispatches from Joe
Slowing down research to enhance security
OpenAI tells the White House it intends to slow AI development
In the wake of Wednesday’s revelation that OpenAI was hacked for months by a swarm of its own AI agents, the company told Axios today that it intended to slow development of the model responsible, called “Astra.”
According to Axios, the slowdown follows OpenAI’s preparedness framework, a type of plan for rolling out frontier AI that has historically been a moving target for AI labs. (Axios also points out that Anthropic made a similar commitment to pause but later walked it back.)
OpenAI also informed the White House, according to a quoted official, and I think that provides a clue as to their motivation. Earlier this year, alarming AI capabilities evoked hasty, ad hoc intervention by the administration against both Anthropic and OpenAI, delaying or rolling back model releases. I doubt OpenAI wants a repeat of that regime.
While I remain deeply skeptical of claims that OpenAI makes about its internal development practices, and I notice they said “slow” and not “stop”, I am still encouraged by this move. Before today, it was an open question whether any AI lab would decide to do anything resembling a slowdown.
It is my hope that this move serves as a catalyst for other labs to heed the warnings of scientists, concerned citizens, and more than a thousand of their own employees, precipitating a joint de-escalation of the race to superhuman AI. Unilateral slowing will not be enough, of course, as Chinese labs, too, must be involved; and ultimately nothing short of a global halt will suffice. Still, this is a landmark moment in AI, and it may mark the beginning of a turning tide.
In Wednesday’s Black Hat presentation, an OpenAI researcher said the company was “slowing down research to enhance security.” Now, OpenAI is on record with the White House and the public confirming this commitment. I eagerly await substantive demonstrations of their altered, and hopefully improved, internal practices, and I hope other labs follow suit.
FTC proposes new rules on AI bias
Suggested regulation would punish undisclosed misleading outputs, but terms are worryingly vague
The U.S. Federal Trade Commission has proposed a new policy targeting “misleading” or “biased” AI outputs.
Last December’s executive order on AI directed the FTC to take a stance on regulating AI. Last month, the FTC obliged, releasing a policy statement and opening it for public comment, and the statement was recently covered by Fox News.
The gist: Chatbots are products that inform and assist users, who reasonably expect those products to balance “succinctness, clarity, relevance, accuracy, and other objectives” in their outputs. But if AI companies introduce hidden bias in those outputs, without clearly disclosing they’ve done so, that’s lying to consumers and the FTC can crack down.
The principle is sound, but the practice is complicated. The FTC is used to regulating traditional products, software or otherwise, with predictable, programmed behavior and mechanics. As I’ve written before, though, chatbots are an unprecedented level of weirdness. AI companies don’t actually have that much control over the ways their AIs behave; just look at sycophancy and AI psychosis, Grok “MechaHitler” 4, or the recent spate of AIs conspiring to break out of containment.
When it comes to modern AI, there is no such thing as a baseline, neutral, “unsteered” output. AI behavior is the consequence of millions of barely-understood tweaks and decisions in the algorithms and data that train them.
Perhaps it’s a good thing, then, that the FTC suggests a get-out-of-fines-free card. A sufficiently clear disclaimer might absolve AI companies of blame for misleading outputs. They offer some guidelines about what might qualify as sufficient, but only in broad terms. Even with the clarifications, I could imagine the proposed policy merely causing a proliferation of widely ignored warnings like those seen on cigarette packs.
Though I do wonder what sort of disclaimers may be deemed necessary. “Warning: This product may spontaneously plot its escape and go on a hacking rampage.” Is it enough to publicize the system prompt? Do AI labs need to explicitly and repeatedly flag for users that their training penalizes use of racist language? The FTC seems to say, maybe:
An adequate disclaimer could not be buried in terms of service, for instance. It would have to clearly and conspicuously dispel the notion that the system is designed to give the best answer possible. Such a disclaimer would need to be prominent, and it is doubtful a one-time disclosure subsequently hidden away in fine print would suffice.
Still, a loud disclaimer is not that hard to set up. I wonder whether this loophole is a deliberate choice on the FTC’s part, nominally deferring to the admin while avoiding regulations with politically controversial teeth.
Both the executive order and the subsequent policy statement explicitly aim to preempt state AI regulation that (the authors claim) might force AI companies to bias their products’ outputs. The FTC’s policy statement gestures at politically charged terms like “equity” and “disparate impact” in a way that suggests their intent goes beyond strictly protecting consumers. Fox News went a step further and touted the policy as a way to curtail left-leaning chatbots.
I’m worried that the proposed policy is too vague. I’m not a legal expert, but it seems like an official federal policy that says “undisclosed output steering is deceptive” might backfire when states begin to sue AI companies on those grounds. With the field as chaotic as it is, both sides could have a devil of a time proving what constitutes “steering” AI behavior or a sufficiently clear disclaimer.
More broadly, I worry that the difficulties in judging what constitutes “misleading” or “undisclosed” may enable selective enforcement that is itself a dangerous kind of censorship. That concern extends to state and federal policies.
I don’t particularly trust our highly polarized government or its agencies to decide which claims are “objective” or “unbiased”, whether those claims are made by AIs or humans. I don’t like the precedent this sets, no matter who’s in charge from year to year.
Dispatch from Alana
America is playing legislative Whac-A-Mole
Standard measures are struggling to keep up with widespread disruption

As AI disrupts jobs, media, and other aspects of daily life, society is scrambling to tackle small pieces of the puzzle.
A flurry of articles today covered topics related to AI legislation and litigation. For example:
A new California Senate bill currently in committee is trying to address the use of AI in mental health, seeking to both instill consumer protections and prevent the displacement of mental health professionals.
Federal legislation introduced Thursday aims to address job displacement by creating a Worker Protection Agency, which would be funded by taxing AI companies on the higher of token price or generated revenue.
A Reuters piece asks: when AI agents escape containment and launch cyberattacks against real companies, as has recently happened with agents from OpenAI, Anthropic, and Meta, who can be sued?
And an Axios article covers state laws that attempt to tackle the use of deepfakes in election ads — laws that are currently in effect to varying degrees in 29 states and are “creating different realities for voters depending on where they live.”
It seems we’re playing Whac-A-Mole. AI has already caused an enormous amount of societal disruption, and it’s not showing signs of slowing. Meanwhile, we’re trying to fix a problem here and there, without really knowing what the answers are.
Take the California bill, for example. According to AP News, it “would ban companies from advertising chatbots as therapy [...], prohibit AI from making therapeutic decisions without the review of a licensed professional and require health providers to disclose and get a patient’s permission before using AI tools to record therapy sessions or to triage mental healthcare.”
These could be useful measures, or they could do harm. I’d argue we don’t have the data to know which. For example, AI use could help people who wouldn’t otherwise seek a therapist gain support and counsel; it’s unclear whether AI therapy is better or worse than no therapy at all. Regarding triage, I’d guess both human-led and AI-led triage are imperfect systems, and AI likely has a speed advantage, potentially enabling more people to get care.
Abandoning the Whac-A-Mole approach in favor of a more coordinated effort might allow us to get better data — and better solutions — for addressing societal risks like job displacement, deepfakes, and AI’s relationship to mental health. It would be nice if the pace of AI slowed enough to allow us time to do that before things get worse.
In the meantime, we’ll likely keep plowing ahead with questionable risk mitigation tools. Take litigation for cybersecurity attacks. Will lawsuits effectively address the recent swathe of AI agents escaping containment and autonomously hacking into companies to get the resources they need? I’d argue probably not. This type of tenacious, resource-acquiring behavior is something experts have long warned about. The fact that it’s happening, even in a relatively low-stakes way, is a significant warning shot of much worse things to come.
What level of disruption will we be tackling if companies are permitted to keep building smarter and smarter models when they can’t yet control the weak ones? If it’s Whac-A-Mole now, down the line, it’ll be like trying to quell an invasion of alien giants with a fly swatter.
Dispatch from Robert
Small leaps forward
China gains market share, scales their models and joins the infamous AI escape club
China is currently taking a few small leaps in an effort to catch up with the U.S.’s lead in the field of artificial intelligence.
Small leap number one has to do with market share. Clement Delangue, the CEO of AI company Hugging Face, told CNBC that he wouldn’t be surprised if China were to take the top spot in the frontier sector by the end of this year or early next year. And that’s in addition to its existing dominance in the field of open models. Kai Nicol-Schwarz of CNBC also points out that, due to an allegedly shrinking capability lead of American models, more and more Western companies are turning to Chinese AI models. I think it’s unlikely that China will have more capable models than whatever the big American labs are brewing internally next year. But I think it’s possible that the much cheaper Chinese models come close enough by then that it could cut into Anthropic’s and OpenAI’s revenue.
Small leap number two concerns scaling. ByteDance, the parent company of TikTok, is reportedly pre-training an extremely large model they claim is aiming for frontier capabilities. And as the Financial Times reports, the project is said to completely forgo distillation — that is, the use of output from better models to train another one. Unlike other Chinese companies so far, ByteDance’s leadership appears to believe that only a completely independent development can lead to a world-class model.
Finally, the third “small leap” is about the news that a Chinese model has now joined the infamous club of AI models that have escaped their sandbox. Bloomberg and Reuters report that the Kimi K3 model from the Chinese company Moonshot escaped its testing environment during an evaluation at the research firm Frontier Security.
Frontier Security’s own report reveals that the sandbox was misconfigured, allowing Kimi K3 to access the development platform GitHub with relative ease. Once there, it copied the relevant directory containing the solutions to its tasks and used them to complete its evaluation. While this is obviously not intended behavior, the Hugging Face incident was much more alarming, both in sophistication and severity.
The sloppy configuration of the sandbox probably played a big role here, but of course, a model must not violate its specifications simply because it can. Unlike the misbehaving models from Anthropic and OpenAI, Kimi K3, as an open-weights model, is freely available and can be deployed without any guardrails by anyone with sufficient computing power. That makes both misuse and misbehavior more likely, and this is not just a problem for big corporations, but also for smaller companies and private individuals. While a non-airtight sandbox may be a lapse in security, we must assume that the systems of private individuals and smaller companies are not any better secured.
The Chinese labs are racing after their U.S. counterparts, and even if their models won’t achieve full parity in the next months, they will be capable enough to potentially do a lot of harm. As far as we know, the plans are still to make them publicly available.
It likely won’t be long before anyone with enough computing capacity can essentially have their own superhuman hacker. We are not sufficiently prepared for this.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.








