In this issue:
PauseAI UK’s visit to Parliament - What I learned when I failed to meet my representative
Tens of thousands of incidents - The AI companies have a data analysis problem on top of their alignment problem
Google engineer resigns to avoid contributing to AI acceleration - Describes it as the Christian thing to do
Guest post by Phil Hazelden
PauseAI UK’s visit to Parliament
What I learned when I failed to meet my representative

Editor’s note: Despite our U.S.-heavy readership, I invited Phil to share this experience because the AI problem is a global problem and many of these U.K. insights apply to politics in general. Keep in mind that the visit he describes took place months before the Jacob Coxon awakening; I imagine the next visit will be different.
It turns out you can just show up at Westminster Hall and ask to meet your Member of Parliament. That’s what I did on June 23rd, along with about 30 others from PauseAI UK. (Every time I mention PauseAI, I mean PauseAI UK. I’m not sure how much coordination there is between them and PauseAI Global.) Our concrete request for them was to sign a (not-yet-publicised) open letter from the organization, calling for AI labs to be held liable for damages caused by their models.
I’d emailed my MP — I’m not naming her, because it’s not my intent to shame anyone — twice at the time, first on February 12, warning about the existential dangers of superintelligent AI, and asking her to sign ControlAI’s campaign statement. She replied on March 26, not dismissing the risks but reassuring me that the Labour government had things in hand. Then on June 17, I asked her to sign the PauseAI letter and to meet me when I visited Parliament. She replied on July 30, again reassuring me. At least she switched from describing the issue as “serious” to “vital”.
I hadn’t heard back by the time we went — I’d been advised to email several times, but I’m not very good at that kind of thing — so I wasn’t expecting a meeting myself. (Also, Keir Starmer had resigned as Prime Minister the day before, so MPs were more likely to be around but also less likely to have time for us.) But I went anyway.
I stood in an airport-style security line for about half an hour, then waited another hour or so before I could fill out my “green card” to request a meeting. Normally that would have been faster, but another group (I don’t know who) had picked the same day to all show up, and the stewards were limiting how often people could go into the central lobby. After filling out the green card (which I think was neither green nor card), I waited for a few more hours without hearing anything, and then left with most of the remainder of the group.
Others were more successful. Among the group, we had meetings with two MPs, three MPs’ assistants and one MP’s researcher: three Labour, two Lib Dems and a Green, meetings ranging from 10 to 50 minutes. Some of us also ran into Jeremy Corbyn (former Labour party leader, now of a party confusingly named Your Party, not MP for any of us), who seemed receptive.
(In fact, all three Labour MPs were also members of the Co-operative party. The two parties have an agreement, and some candidates run simultaneously for both. Usually they’re just counted as Labour MPs; Co-operative is marginal enough that I didn’t know they existed. It’s surprising that all the Labour MPs we met were Labour Co-op — 0.1% prior probability — but I don’t know if it means anything.)
I wasn’t present for any of this, but we shared notes afterwards, on the meetings and on email conversations we’d had. Here’s what I’ve learned.
MPs are slow at replying to email
They’re even slower than me, though I grudgingly admit they have more excuse. Of the dated correspondence I’ve seen, the fastest first-reply took 18 days.
That doesn’t mean there’s always a two-week lag in communication. Followup after meetings or previous replies tends to be faster. One person had had relative radio silence for weeks, but got a meeting after one final email the day before our visit (following advice I had failed to heed).
Still, 18 days is a long time in AI. During the 18 days before I wrote this paragraph, September 7 to September 24, StopWatch published 62 dispatches. OpenAI announced a solution to Navier-Stokes. Amodei, Musk, Altman and Hassabis agreed on the need to “pace the frontier”. We learned that Anthropic had set up a wet lab. Trump declared that the US would call AI “super intelligence”. OpenAI disclosed that their agents had hacked into an Australian Medicare website.
When emailing your representative, it’s probably worth making asks that will still be relevant in a month. You can’t always know what those are, and I think PauseAI succeeded.
Politics rarely moves this fast outside of emails, either. Alex Sobel’s Artificial Superintelligence Bill had its first reading in the Commons on September 8, and is slated for its second (if it even has one) on November 13, more than two months later. If we can somehow speed this sort of thing up, or encourage politicians to speed it up, that seems valuable.
MPs’ opinions aren’t necessarily their own
Reading a bunch of emails from MPs, it became clear how generic they tend to be. Several of the ones I’ve seen could almost have come from a keyword-driven auto-reply. You mentioned risks from AI, so here is my statement about risks from AI. (Sometimes they pass on the letter to someone else, and forward you their keyword-driven auto-reply.) I haven’t seen any “I disagree with you about …” or “I’m not going to do the specific thing you asked, but I will instead …”.
I was only mildly surprised to see that two emails from different MPs (one of them mine) were word-for-word identical. In their defense, the emails they were replying to were pretty close as well: PauseAI’s email-your-MP tool fills in a helpful default letter, I’d only lightly edited mine before sending, and I think the other person hadn’t edited at all.
I understand the incentives here, or at least I can guess at them. MPs are busy, and replying to emails isn’t their main job. Sometimes they get a lot of substantially similar emails. Political parties like to seem united. If the party drafts a shared response, then MPs can send a lot of similar replies without much time investment.
So it makes sense. But if I tell you “I asked an LLM to write this opinion piece, and I didn’t change anything because I agree with it”, you probably won’t assume you’re seeing something that reflects my honest considered opinions. You might wonder whether I even have honest considered opinions. This feels similar to me. We’re learning something about the zeitgeist-within-Labour, but not so much about these specific MPs.
I was curious who wrote the text of those emails. TheyWorkForYou has a useful search feature: At least some of it comes from a speech given on June 5. Another part comes from a written answer from the Department for Science, Innovation and Technology (DSIT) on April 28. My best guess is it was compiled by the Parliamentary Research Service (PRS) that Labour MPs have access to — at least, that’s the kind of thing they’d do. After I noticed this, I searched for other phrases in emails and found some of those in other statements and speeches.
A positive side to this is that tides can probably shift faster than we might otherwise expect. If PRS becomes convinced of the extinction problem, then anyone following along with them will go on record believing it too. It also suggests points of added leverage. I’m not sure if “try to talk directly to PRS” is a live option — it seems like the kind of thing everyone with an agenda would be doing, if so — but perhaps we can see who PRS quotes most often and try to get those people on board.
People can also start to take this seriously without PRS. The other MP who sent the same email as mine also didn’t meet with us on the day of our visit. But his constituents did get a meeting with him later, and now they’re having a conversation for real.
Shout out to Chris Vince (Labour, MP for Harlow), who followed up his constituent meeting by sending an email to DSIT. He talks intelligently about AIs being misaligned with human intentions, and about failures of interpretability research. If he got it all from this meeting, I’m impressed with both him and his constituent.
We could stand to be clearer about the extinction problem
I was disappointed that it was barely mentioned. I mentioned it in all the emails I’ve sent, but PauseAI email templates don’t bring it up; it’s in an attached policy briefing, but a mention in a multi-page pdf is much less visible than one in the email body itself. And when I read through the notes of the meetings that took place, no one clearly talks about it. There are mentions of national security risk, bioterrorism, catastrophic risk, and loss of control. Some of that is very close, and maybe in some of the actual conversations it was clear that human extinction was on the table. But it feels like we’re shying away from talking about the danger we’re most worried about.
To some extent this makes sense. The specific short-term goal of “AI labs should be held liable for damages caused by their models” would help to slow down the race, and help internalize the costs of current levels of misalignment. It would (I claim) be worth pushing for even in the absence of an extinction threat. It’s a thing the UK can do now. It can be motivated by concrete things happening now. So why muddy the waters with theorizing about what might happen in future?
But it’s also insufficient: I don’t expect current alignment methods to scale to superintelligence; and no company can internalize the costs of human extinction. Maybe the UK can’t do anything sufficient yet. But the landscape can change, and when it does I think we’ll be better off if our MPs already know about the extinction problem, and know that other MPs know. If the UN starts discussing an international treaty to ban the development of superhuman AI, I’d like the UK to back it from the beginning.
So how effective was this?
That’s hard to say. Anneliese Dodds and Siân Berry signed the open letter after meeting with us; but both had signed ControlAI’s campaign statement months previously. So I guess this was less “bringing someone on board” than “helping keep them focused”. None of the MPs we met with signed Alex Sobel’s subsequent letter to the Prime Minister introducing his Artificial Superintelligence Bill; I don’t know if they declined, or were never asked.
But I guess a lot of the point is about making MPs aware that we’re a coalition that exists and votes, in a way that’s (uncomfortably) independent of how correct we are. We hoped that several people all requesting meetings on the same topic on the same day would be noticeable; we’re unlikely to get a clear picture of whether that worked or not, but we certainly shouldn’t assume it didn’t. And, even if we failed to convince someone on the day, we might have helped bring them around slowly.
I’m up for trying again. Finnish activists tell me that post-Hugging Face, 85% of their requests to meet MPs have been successful. (Though they also say that where I have only one MP representing me specifically, Finns have a lot more choice about who they can reasonably contact.)
I emailed my MP again on September 18 to ask her to support the Artificial Superintelligence Bill, and to meet me on October 20 when PauseAI next visits. (If you’re in the UK, consider joining us? Or even if you can’t join, you can still ask your MP to sign the open letter.) In light of the …stuff… that’s happened since then, maybe she — or the Labour zeitgeist she draws from — will be more open to the idea that “our current approach is insufficient”.
I do think that this time, we should be clearer that the priority is human extinction. It’s in ControlAI’s campaign statement, with 67 out of 650 MPs as signatories, plus 100 other lawmakers; so I think it’s on the table.
Phil Hazelden is a former software engineer in London. He’s been abstractly interested in the alignment problem for about fifteen years, but for most of them he didn’t expect it to be relevant for a long time.
Dispatches from Mitch
Tens of thousands of incidents
The AI companies have a data analysis problem on top of their alignment problem

If you read that headline and thought, “Well now it’s just ridiculous,” you’re not alone. The headline as it appeared on Axios was “OpenAI, Anthropic probing tens of thousands of security incidents.” I freaked out a little bit, too.
Looking into it, the paltry reassurance I can give you is that, in this estimate from unnamed industry sources, every instance of an agent taking a step that “outside evaluators would consider problematic” counts as an incident.
By this math, the Hugging Face attack alone likely comprises many thousands of incidents from the hundreds of agents that participated. And the newly reported case of OpenAI “aggressively browsing” a U.N. data hub more than 16,000 times, circumventing a filter and violating site policy, might count as 16,000 incidents.
But even if the number of affected and targeted sites only numbers in the dozens or hundreds, the “tens of thousands” figure still reflects a real problem for understanding and preventing mishaps. As OpenAI CEO Sam Altman posted on Friday about the company’s worrisome-but-incomplete disclosures, the logs add up to petabytes of data. (For comparison, the text-only portions of the U.S. Library of Congress probably total less than a tenth of a single petabyte.)
Humans aren’t going through all that. They can’t. They must rely on other agents whose decision-making and reliability are also suspect. This was a big complaint of the independent researchers invited to make a brief and narrowly scoped investigation of the Hugging Face incident.
The sensible thing would of course be to not keep training the kinds of models that constantly need investigating. The inability to keep tabs on today’s AI is just a taste of what it will look like to lose control on purpose — leaving the creation of AI itself to other AI.
Google engineer resigns to avoid contributing to AI acceleration
Describes it as the Christian thing to do

In a tweet with over a million views in three days, Google engineer Robert O’Callahan announced he was resigning from his work on chips that would make AI faster and cheaper: “I think AI is already progressing too fast, so I had to quit.”
In a more detailed blog post, he explained:
There are millions of people contributing to AI acceleration and taking my foot off the accelerator will have a very small impact ... but not no impact; some of my skills are rare.
Linking to the book If Anyone Builds It, Everyone Dies, the New Zealander added:
I think the existential risks many people are warning about deserve to be taken seriously; a lot of the phenomena predicted by the “doomers” have come to pass (e.g., reward hacking, misalignment, deceptive models, model eval awareness, psychotic swarms).
The uncertainty he shares with others about this danger is itself “very alarming,” warranting a “massive effort” to minimize risk and back away from building superhuman AI. He’s also very concerned about other AI-related issues like power concentration, cybersecurity, and “cognitive surrender.”
A self-described Presbyterian and “occasional lay preacher,” O’Callahan wrote:
It’s tempting to just turn a blind eye to the impact of my work, but that would not be a Jesus-following thing to do.
He followed up with an explanation of how he thinks about AI against his conviction that “God has a plan that’s good for us.”
I don’t know what that plan is (and wish I did) but it lets me sleep at night in spite of the AI chaos. I expect his plan involves me continuing to make the best use of my talents. Even if the plan is for Jesus to return to rescue us from our folly, we’d better be busy when he returns!
He wants his future work to be “unambiguously pro-human.”
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.



