New AI safety watchdog Guidelight has published its first assessment of whether the frontier labs can control the AI systems that they are building. Spoiler: Guidelight is not impressed.
Some backstory: Guidelight, which launched this past May, is run by two people who used to work on safety at OpenAI: Page Hedley, of OpenAI’s policy and ethics, and Steven Adler, a safety researcher at OpenAI until 2024. Prior to this assessment, Guidelight published three standards laying out what the frontier labs ought to do, in order to have something on which to measure their performance. (The assessment discussed today judges just one of those standards.)
Guidelight evaluated five companies: Anthropic, Google, Meta, OpenAI, and xAI. They judged the companies on various criteria — which were individually scored from 0 to 5 — and assigned an overall letter grade based on their average score.
Nobody received a higher score than 3 for any criterion; for their letter grade, Anthropic and OpenAI were tied at C+. (This is a strange letter grade for them to get, by the way. Guidelight says that their letter grades map to the U.S. GPA scale, but GPAs run to 4.0. Anthropic and OpenAI averaged scores of 2.5 out of 5.0, though, and on the scale Guidelight links, a score of 50% would warrant a letter grade of F.)
The companies were graded on six practices: (1) Logging, or whether the company tracks what their AI systems are doing; (2) Monitor Efficacy, or the actual effectiveness of that tracking at catching misaligned behavior; (3) Gated Actions, or whether the monitor must judge and approve certain kinds of risky actions before they take effect; (4) Circuit Breaking, or whether the company’s AI systems can be automatically halted — without human intervention — in response to misbehavior; (5) Third-party Review, or whether an outside party has been allowed to examine the company’s control regime and judge its adequacy; and (6) Containment Plan, or whether there’s a plan to handle an AI model that may go rogue, by restricting it and/or shutting it down.
Only publicly available information was used, but I won’t hold that against the assessment. The labs don’t get credit for any plan that amounts to, “trust me, bro.” These aren’t nuclear-response plans; if a frontier lab has a plan it isn’t sharing, then it probably isn’t confident in that plan.

(The scoreboards on the website are interactive and give additional details.)
Guidelight says that the frontier labs are strongest in detection and third-party assessment. In other words, the labs may not be able to do much to keep us safe, but at least they’ll (probably) be able to let us know when the horse has left the barn. Here, xAI scores outright zeros: there is no evidence that it logs internal use of AI, no evidence that it measures whether its monitoring systems work, and no stated plan to address either of these things. The company has also not taken part in any third-party assessments, so far as Guidelight was aware.
In every lab that has a “circuit breaker” to block unsafe actions by AI systems (and neither Google nor Meta seem to have one), there are at least three shortcomings: (1) the circuit breaker is not applied universally throughout the company; (2) the circuit breaker does not review chain-of-thought reasoning; and (3) the circuit breaker applies to individual agents or even sub-agents, depending on the trigger, which means that, if a sub-agent is halted by a circuit breaker, the AI agent that spawned that sub-agent can keep making new sub-agents and trying again without the circuit breaker ever tracing the unsafe behavior to its root.
Anthropic received a 0 in the Containment Plan criterion because — I am horrified to read — “We have found no evidence that Anthropic has a containment response plan, or has any plans to adopt one.” (Postponing and limiting Mythos and Fable don’t improve Anthropic’s score because Mythos and Fable weren’t trying to break out.) The highest score (3) was achieved by OpenAI. This is perversely due in part to the Hugging Face attack, in response to which OpenAI stated it has temporarily paused the training of new models and restricted internal deployment.
That’s not a containment plan. But it’s the only sane response to recognizing that you don’t have a sufficient containment plan. If you can’t contain the AI models that you’ve already got, then stop building models that are even more powerful.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.
You can receive emails of dispatches as we write them, or subscribe to our Daily Digest for a once-a-day compilation.


