Are AI microdramas the fast track to frontier robotics?
Op-ed argues that China's dominance in video models gives it a strategic advantage

Fine, I’ll bite.
Bloomberg columnist Catherine Thorbecke’s latest piece argues that China’s dominance in AI video generation is positioning it for dominance in robotics. I’m skeptical, but with just enough uncertainty to not dismiss the argument as a creative way to hype the U.S.-China race narrative.
It’s at least an interesting argument. It goes: 1) China has nine of the ten highest-rated text-to-video models, and is likely to continue leading in this area. 2) Video generation models are precursors to so-called “world models.” And 3) world models are prerequisites for useful general-purpose robots.
Let’s break it down.
First: Yes, China does indeed dominate in video models. Competition for video models within China is largely driven by the appetite for microdramas. If you’re not familiar with microdramas, perhaps you’ve heard of Fruit Love Island, an offbeat example of the genre. A typical microdrama is a series of dozens of video episodes that might each be only a minute or two long, optimized for viewing on phones and sharing across social media. Microdramas aren’t always AI generated, but the AI variety is cheaper to produce. Chinese customers have been willing to spend serious money for microdramas, and Chinese video models have gotten very good.
Next, the “world model” part. Last month, we described a world model as an AI that is “explicitly trained to predict features of an environment rather than the next fragment of text in a passage.” Thorbecke’s argument assumes that video models are much closer to true world models than language models like ChatGPT or Claude, putting China ahead of the U.S. in this area.
This claim is more suspect.
That’s partly because ChatGPT and Claude aren’t pure language models. While we don’t know exactly what they are trained on, we know they are multi-modal, training on images as well as text, and video is just a series of consecutive images.
But also, the reason Thorbecke thinks robots need world models doesn’t actually give an obvious advantage to video models. She gives the following example:
A robot cannot fold laundry or stock a supermarket merely by recognizing what objects are; it must anticipate what will happen when [it] moves or interacts with them.
I agree such understanding is necessary. But the general reasoning skills language models acquire by learning to predict text and solve a wide variety of problems imply a much better model of the physical world than Thorbecke may realize.
And while a video model might train to predict the next frame of a video where a robot hand is putting a box of cereal on a shelf, a language model learns to understand what a supermarket is, what’s in the box, and why someone might buy it. If you want a robot that is more like a flexible human employee, this broader understanding may matter at least as much as hand-eye coordination. Both types of model are going to need robot-specific training, and it’s not clear to me which would take to it quicker or have the higher ceiling.
It’s outside the scope of Thorbecke’s argument, but I think we should also back up to ask why American companies are no longer leaders in text-to-video. I don’t think it’s because Americans aren’t quite as into microdramas, though this may be a factor. I think it’s because the AI companies have decided to focus on models that can solve the problems of AI researchers, in hopes of speeding the development of still-more-powerful models. If you are racing to superintelligence, video is a distraction, and an expensive one at that: Video is computationally very expensive to generate compared to text, and there are only so many chips to go around.
OpenAI, remember, discontinued its Sora family of video generators in March amid reports of compute shortages and loss of market share to Anthropic, which never dabbled in video at all. Sora and Sora 2 were both briefly considered frontier video models at the time of release.
Superintelligence, by definition, would have no trouble piloting robots in economically useful ways. So even if Thorbecke is right that video models are the fast track to robotics, American CEOs are betting their companies — and our lives — that their language models can make an end-run to robotics and everything else without the AIs ending up in control. It’s the stuff of drama, but it is not micro.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.


