In this issue:
Citizen suits under attack by Musk, Trump administration - If nobody is willing or able to enforce a law, does it really exist?
Are AI microdramas the fast track to frontier robotics? - Op-ed argues that China’s dominance in video models gives it a strategic advantage
Dispatches from Mitch
Citizen suits under attack by Musk, Trump administration
If nobody is willing or able to enforce a law, does it really exist?

When Americans go looking for the reasons why it has become so much harder to build housing and infrastructure in the U.S., the finger is usually pointed at regulations, especially environmental regulations. But such regulations are typically more of a scab that has formed over the actual obstacle: the ability of citizens to sue. Any aggrieved party can potentially halt a project if they can point to a law that is being violated.
Elaborate regulations around permitting spring up as safe corridors against such litigation. A permitting process is an agreed-upon way to show agencies and the courts that a project is in compliance with the law. This is a main reason industries often ask to be regulated.
If the government wants to make life easier for builders, then, relaxing permitting requirements can actually be counterproductive. Permitting is slow and expensive, but without it, builders are more exposed to litigation, and face greater uncertainty about whether a project will go through at all. Investors tend to hate uncertainty more than they hate a lengthy permitting process.
So if you are the White House, and you want to make it easier for data centers to be built and operated in ramshackle fashion — e.g., powered by noisy and polluting mobile gas turbines in populated areas, as Elon Musk’s AI company is doing in Tennessee — you have to do more than instruct agencies to relax permitting requirements. You have to limit the ability of citizens to sue.
You could do this by working with Congress to repeal laws plaintiffs can point to as being broken (like the Clean Air Act, in this case), or to pass laws providing legal immunity for certain activities. But this is politically complicated, and may not be possible if Congress doesn’t want to go along.
So you might try something bolder, and challenge the very idea of citizen suits.
As reported in the Associated Press, Elon Musk’s xAI (now part of SpaceX), with the administration’s support, is arguing that Congress never had the constitutional authority to allow anyone to enforce federal laws except the President.
This is seen as a legal long shot, made remotely possible only by the current composition of the Supreme Court and pressure from the White House. If successful, this would shake up the nation’s legal landscape far beyond the data center buildout. Citizen suits have successfully punished oil and gas companies that have contaminated water supplies and caused other harms, and presumably deterred greater corporate misbehavior.
In the meantime, the gas turbines at SpaceX’s Memphis data center remain in operation, because the Trump administration has been telling the courts that the facility supports the Department of War and is thus vital to national security. The administration could presumably try to make this claim about any U.S. data center, but at some point the courts may scoff.
Observers note how backwards it is for the executive branch, charged with enforcing laws, to instead be shielding companies from them. Citizen suits have been serving as a check against this pattern. Should they become impossible, many laws might as well not exist. Those tracking the many ways the AI race could potentially harm America should be following this story.
Are AI microdramas the fast track to frontier robotics?
Op-ed argues that China’s dominance in video models gives it a strategic advantage

Fine, I’ll bite.
Bloomberg columnist Catherine Thorbecke’s latest piece argues that China’s dominance in AI video generation is positioning it for dominance in robotics. I’m skeptical, but with just enough uncertainty to not dismiss the argument as a creative way to hype the U.S.-China race narrative.
It’s at least an interesting argument. It goes: 1) China has nine of the ten highest-rated text-to-video models, and is likely to continue leading in this area. 2) Video generation models are precursors to so-called “world models.” And 3) world models are prerequisites for useful general-purpose robots.
Let’s break it down.
First: Yes, China does indeed dominate in video models. Competition for video models within China is largely driven by the appetite for microdramas. If you’re not familiar with microdramas, perhaps you’ve heard of Fruit Love Island, an offbeat example of the genre. A typical microdrama is a series of dozens of video episodes that might each be only a minute or two long, optimized for viewing on phones and sharing across social media. Microdramas aren’t always AI generated, but the AI variety is cheaper to produce. Chinese customers have been willing to spend serious money for microdramas, and Chinese video models have gotten very good.
Next, the “world model” part. Last month, we described a world model as an AI that is “explicitly trained to predict features of an environment rather than the next fragment of text in a passage.” Thorbecke’s argument assumes that video models are much closer to true world models than language models like ChatGPT or Claude, putting China ahead of the U.S. in this area.
This claim is more suspect.
That’s partly because ChatGPT and Claude aren’t pure language models. While we don’t know exactly what they are trained on, we know they are multi-modal, training on images as well as text, and video is just a series of consecutive images.
But also, the reason Thorbecke thinks robots need world models doesn’t actually give an obvious advantage to video models. She gives the following example:
A robot cannot fold laundry or stock a supermarket merely by recognizing what objects are; it must anticipate what will happen when [it] moves or interacts with them.
I agree such understanding is necessary. But the general reasoning skills language models acquire by learning to predict text and solve a wide variety of problems imply a much better model of the physical world than Thorbecke may realize.
And while a video model might train to predict the next frame of a video where a robot hand is putting a box of cereal on a shelf, a language model learns to understand what a supermarket is, what’s in the box, and why someone might buy it. If you want a robot that is more like a flexible human employee, this broader understanding may matter at least as much as hand-eye coordination. Both types of model are going to need robot-specific training, and it’s not clear to me which would take to it quicker or have the higher ceiling.
It’s outside the scope of Thorbecke’s argument, but I think we should also back up to ask why American companies are no longer leaders in text-to-video. I don’t think it’s because Americans aren’t quite as into microdramas, though this may be a factor. I think it’s because the AI companies have decided to focus on models that can solve the problems of AI researchers, in hopes of speeding the development of still-more-powerful models. If you are racing to superintelligence, video is a distraction, and an expensive one at that: Video is computationally very expensive to generate compared to text, and there are only so many chips to go around.
OpenAI, remember, discontinued its Sora family of video generators in March amid reports of compute shortages and loss of market share to Anthropic, which never dabbled in video at all. Sora and Sora 2 were both briefly considered frontier video models at the time of release.
Superintelligence, by definition, would have no trouble piloting robots in economically useful ways. So even if Thorbecke is right that video models are the fast track to robotics, American CEOs are betting their companies — and our lives — that their language models can make an end-run to robotics and everything else without the AIs ending up in control. It’s the stuff of drama, but it is not micro.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.



The robotics argument can be made stronger in this way: Yes, a superintelligence would by definition be able to predict the physical world and its reaction to particular actions. But this is not how we humans fold laundry. Doing this with our explicit reasoning skills is very cumbersome and simply too slow for many tasks. Instead, we use something similar to a smaller model that is faster and much less accessible to explicit reasoning. So even if a superintelligence would be able to steer robots with its explicit reasoning skills, that would probably be inefficient and slow. It might, of course, be able to develop such a smaller model quite fast, and I guess this is what the US frontier companies are betting on.
In addition, getting to superintelligence and using it will require much compute, which in turn might become available faster if efficiently operated robots are available...