In this issue:
Good at noticing / Bad at stopping - New scorecard finds no frontier lab has a working mechanism to halt a rogue AI model
What automation threats teach us about AI capabilities - Mathematicians, doctors, and many others confront the jagged frontier of AI
Despite PR efforts, Americans still oppose datacenters, and AI - Community pledges and funding don’t seem to be swaying public opinion by much
Dispatch from Donald
Good at noticing / Bad at stopping
New scorecard finds no frontier lab has a working mechanism to halt a rogue AI model
New AI safety watchdog Guidelight has published its first assessment of whether the frontier labs can control the AI systems that they are building. Spoiler: Guidelight is not impressed.
Some backstory: Guidelight, which launched this past May, is run by two people who used to work on safety at OpenAI: Page Hedley, of OpenAI’s policy and ethics, and Steven Adler, a safety researcher at OpenAI until 2024. Prior to this assessment, Guidelight published three standards laying out what the frontier labs ought to do, in order to have something on which to measure their performance. (The assessment discussed today judges just one of those standards.)
Guidelight evaluated five companies: Anthropic, Google, Meta, OpenAI, and xAI. They judged the companies on various criteria — which were individually scored from 0 to 5 — and assigned an overall letter grade based on their average score.
Nobody received a higher score than 3 for any criterion; for their letter grade, Anthropic and OpenAI were tied at C+. (This is a strange letter grade for them to get, by the way. Guidelight says that their letter grades map to the U.S. GPA scale, but GPAs run to 4.0. Anthropic and OpenAI averaged scores of 2.5 out of 5.0, though, and on the scale Guidelight links, a score of 50% would warrant a letter grade of F.)
The companies were graded on six practices: (1) Logging, or whether the company tracks what their AI systems are doing; (2) Monitor Efficacy, or the actual effectiveness of that tracking at catching misaligned behavior; (3) Gated Actions, or whether the monitor must judge and approve certain kinds of risky actions before they take effect; (4) Circuit Breaking, or whether the company’s AI systems can be automatically halted — without human intervention — in response to misbehavior; (5) Third-party Review, or whether an outside party has been allowed to examine the company’s control regime and judge its adequacy; and (6) Containment Plan, or whether there’s a plan to handle an AI model that may go rogue, by restricting it and/or shutting it down.
Only publicly available information was used, but I won’t hold that against the assessment. The labs don’t get credit for any plan that amounts to, “trust me, bro.” These aren’t nuclear-response plans; if a frontier lab has a plan it isn’t sharing, then it probably isn’t confident in that plan.

(The scoreboards on the website are interactive and give additional details.)
Guidelight says that the frontier labs are strongest in detection and third-party assessment. In other words, the labs may not be able to do much to keep us safe, but at least they’ll (probably) be able to let us know when the horse has left the barn. Here, xAI scores outright zeros: there is no evidence that it logs internal use of AI, no evidence that it measures whether its monitoring systems work, and no stated plan to address either of these things. The company has also not taken part in any third-party assessments, so far as Guidelight was aware.
In every lab that has a “circuit breaker” to block unsafe actions by AI systems (and neither Google nor Meta seem to have one), there are at least three shortcomings: (1) the circuit breaker is not applied universally throughout the company; (2) the circuit breaker does not review chain-of-thought reasoning; and (3) the circuit breaker applies to individual agents or even sub-agents, depending on the trigger, which means that, if a sub-agent is halted by a circuit breaker, the AI agent that spawned that sub-agent can keep making new sub-agents and trying again without the circuit breaker ever tracing the unsafe behavior to its root.
Anthropic received a 0 in the Containment Plan criterion because — I am horrified to read — “We have found no evidence that Anthropic has a containment response plan, or has any plans to adopt one.” (Postponing and limiting Mythos and Fable don’t improve Anthropic’s score because Mythos and Fable weren’t trying to break out.) The highest score (3) was achieved by OpenAI. This is perversely due in part to the Hugging Face attack, in response to which OpenAI stated it has temporarily paused the training of new models and restricted internal deployment.
That’s not a containment plan. But it’s the only sane response to recognizing that you don’t have a sufficient containment plan. If you can’t contain the AI models that you’ve already got, then stop building models that are even more powerful.
Dispatches from Joe
What automation threats teach us about AI capabilities
Mathematicians, doctors, and many others confront the jagged frontier of AI
A decade ago, hardly anyone dreamed that AIs would be writing poetry while struggling with basic arithmetic. The advent of large language models flipped that script for a time; in 2020, GPT-3 was doing just that. But the math is back, with a vengeance, and it brought friends. Automation now threatens a wide range of roles, and the sheer diversity of that threat has a great deal to teach us about the capabilities of artificial general intelligence.

Today, mathematicians are seriously talking about losing their entire field to AI. In the wake of several groundbreaking mathematical milestones reached by AI, the Washington Post covers a day-long summit about the future of math. In one talk, University of Toronto professor Daniel Litt argued there’s a chance “human expertise in mathematics is totally lost.” MIT researcher Drew Sutherland added that it might not stop there:
These capable intelligences are going to be applied elsewhere and to other fields. And maybe we’re just the canaries in the coal mine that it’s coming for the mathematicians first.
I suspect Sutherland has come to understand an important fact about AI capabilities: AI is getting better at everything, but not at the same rate. This is a further illustration of what’s been called “jagged intelligence.”

One example of this jagged intelligence at work: Anthropic recently announced progress made by its AI, Claude, in life sciences. Specifically, Claude designed hundreds of new “minibinders”, small proteins that attach to specific target proteins, like connectors for life’s building blocks, and it matched fully equipped human experts in analyzing the content of a chemical sample. The protein design looks like the bigger result, with Claude having found roughly half as many minibinders in this effort as humans have to date. But automating a days-long chemical analysis in under half an hour is nothing to sneeze at, either.
As developers train larger models on new datasets, AI capabilities expand in jagged spurts. Some, like Anthropic’s chemistry progress, make a certain amount of sense when you consider the training data (though the exact rate of progress is still unpredictable). Other capabilities, like social engineering or inducing psychosis in the vulnerable, come as a nasty surprise.
A decade ago, I might have naively expected that AI progress would advance along a predictable track, starting with logic and “pure math” and passing up through physics, chemistry, biology, and eventually psychology, medicine, law, and other messy fields. After all, it’s what sci-fi has primed us to expect of machines. But as often happens, reality has proved weirder than fiction.

In any case, Sutherland is right. AI is already encroaching on fields besides mathematics — like, say, medicine. As reported by Axios, an article in the Journal of the American Medical Association (JAMA) suggests that AIs may soon exceed doctors in many aspects of medicine.
Generative AI “rivals or outperforms” doctors at five cognitive medical tasks, the piece argues: gathering patient information, making diagnoses, selecting tests to establish a diagnosis, prescribing appropriate treatments and managing chronic diseases.
A day later, the AMA and the Digital Medicine Society published a framework covering roles they consider “fundamental to the profession.” I notice a concerning brevity and vagueness in their list of indispensably human responsibilities, and I wonder if next year’s framework will be shorter and vaguer still.
This week alone, we’ve seen several more fields visibly threatened by AI. Today, the Guardian described AI’s growing use in human resources and hiring. Even physical tasks are not exempt, as Reuters and others discuss the latest logistics, manufacturing, and service demos from Chinese robot makers: moving and sorting boxes, packaging phones, and handling household chores.
Meanwhile, Bloomberg catalogues some of the options that are being proposed for after AI replaces all the jobs, like AI dividends or basic income for all.
My own views here are somewhat conflicted. High-quality work in any field is usually a good thing, whether that work is done by machines or humans. But not all AI outputs meet that bar, and a proliferation of slop serves no one. And for various reasons, the end state of the trajectory we’re on does not look very human at all.
Increasingly, though, people seem to be realizing that they need not stand helplessly by while their careers succumb to automation.
A mathematician friend of mine, Xiaoyu He, has taken the alarm in his own field as an opportunity to warn his peers about the threat of extinction AI poses, urging them to turn their attention to understanding and steering the minds of AIs. I commend his initiative and drive.
For my part, I normally recommend that those interested in the future of AI set their sights instead on policy or technical governance, or simply follow Xiaoyu’s example and share their own concerns with their peers.
Despite PR efforts, Americans still oppose datacenters, and AI
Community pledges and funding don’t seem to be swaying public opinion by much
It’s looking like opposition to datacenters will be a major fixture of upcoming U.S. elections.
It’s already a central issue in several gubernatorial races, Axios reports, citing Pennsylvania and Wisconsin as examples. Pennsylvania’s Democratic governor Josh Shapiro and his Republican opponent Stacy Garrity are both taking stances that would limit datacenter development in the state; yesterday Shapiro signed an executive order imposing strict requirements on datacenter construction, and Garrity has been campaigning on a full-blown halt. In Wisconsin, the candidates are each accusing the other of being overly friendly towards datacenter projects.
This view is bolstered by a frankly odd memo to AI companies from the National Republican Senatorial Committee, which represents GOP candidates in campaigns. Also covered by Axios, the letter warns AI companies that datacenters are featuring heavily in a Democratic Senate campaign in Ohio, and “If voters’ perceptions of data centers are not fixed quickly, the campaign against them will expand far beyond Ohio.”
I can’t tell whether this is supposed to be a friendly warning or a threat by the Republican party to oppose datacenters if they lose Ohio, but either way, I don’t think AI companies need the encouragement to try painting datacenters in a positive light. The Wall Street Journal describes an ongoing publicity campaign promoting them; one Georgia effort by OpenAI featured an open house meeting, $150 million in community pledges, and a taco bar.
Datacenter builders like Meta, Microsoft, and Amazon have been making public commitments around electricity costs, water use, and jobs. Meta has also pledged a $1 billion fund for sweetening the deal in affected communities. I’m not getting the sense the communities are buying it, though.
I honestly feel some sympathy for these companies, as they experience a frustration many Americans share with the sea of misunderstandings, environmental reviews, and red tape that often plagues large-scale construction projects. Assuming the builders make and keep commitments to reimburse anyone affected, and offset costs like electric bills for the local community, I don’t think datacenters are innately bad. (Though it’s telling that they seemingly weren’t making such commitments until after the massive backlash, and that many builders are trying to keep construction a secret from communities with nondisclosure agreements.)
At the same time, I think AI companies’ reckless endangerment of our species is more than deserving of ire, and I have to respect the sea of grassroots opposition on display. And it’s not just limited to datacenters: The Washington Post reports on a Pew Research poll finding that when it comes to AI, “more concerned than excited” young Americans now outnumber the inverse five to one, and nearly three-quarters think AI will destroy more jobs than it creates.
I don’t think there’s a community pledge on Earth that is large enough to offset the direction they’re taking humanity’s future.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.




