Reasoning traces
AI thoughts left unlocked, autonomous gym hack, North Korean hacker kit, and more
In this issue:
A penny for everybody’s AI’s thoughts - Reasoning traces are a lock to which everyone had a key (till last month)
Moral judgment lacking on all sides of AI gym hack story - We’re going to need the bots to be more ethical than we are if we want to be able to have nice (unsecured) things
Online courses might be cooked - AI can now handle every part of online student life
North Korean hacking toolkit examined - No one is surprised to find it full of open-weights AI
Dispatch from Donald
A penny for everybody’s AI’s thoughts
Reasoning traces are a lock to which everyone had a key (till last month)

Researchers from the University of Tübingen, the Max Planck Institute, MATS Research, and Snyk have demonstrated a method for recovering the “chains of thought” that frontier AI models generate before they answer prompts. This method provides new evidence that Chinese labs have been using a technique called “distillation” to reverse-engineer more advanced AI models. It also provides new evidence that the frontier labs are kind of just taking the vibe code approach to cybersecurity.
(Distillation involves repeatedly querying a model and using its output to train another model. My colleague Alana has explained the technique in greater detail if you’re interested. The Wall Street Journal’s Christopher Mims notes that distillation is not done only by Chinese labs, nor done only by labs illicitly trying to catch up to other labs.)
How the new method works: When you interact with an AI model for more than a single turn, it needs to keep track of what it’s been “thinking” in previous turns, and that memory — a block of “reasoning traces” — has to be stored somewhere. To save on storage space, the frontier labs store the reasoning traces for your conversations on your computer, rather than on their servers. These reasoning traces are encrypted, to keep the model’s thinking secret, but you don’t have to break the encryption if you have something that can just open it. You can hand the reasoning traces to another, smaller model in the same family, then ask it to transcribe the reasoning traces. Bigger models will typically refuse such requests, but the smaller models have weaker safeguards, so they’ll do it. It’s like getting access to somebody’s phone, not because you’re some kind of skilled hacker, but because their kid brother knows the password and doesn’t realize he shouldn’t open the phone for whoever asks.
In their paper, the researchers show that, for certain prompts, there is a close similarity between the reasoning traces of the AI model Kimi K3, by Chinese lab Moonshot AI, and the U.S. models Claude Opus 4.8 and GPT 5.6. Similarities were also observed when testing GLM-5.2, a model by Z.ai, another Chinese lab. It is important to note, however, that this was not true of all Chinese models: DeepSeek-V3.1, by the Chinese lab DeepSeek, did not display this close similarity. Strangely, neither did Moonshot AI’s earlier models K2.5 and K2.6, the latter of which was released in April 2026. This surely doesn’t mean, “Chinese labs are using distillation, but only started doing so in the past few months.” Perhaps they are (or were) using this method of distillation, but only recently.
It gets a little stranger, though. Claude’s reasoning traces are more or less interchangeable between most of the different-strength models in the family: Haiku, Sonnet, and Opus, but not Fable. Like a patient with Type AB blood, Fable can receive and interpret the reasoning traces of any of the other Claude models, but they can’t interpret Fable’s. What’s interesting to me about this is that Fable is the specific model that some people have accused Moonshot of distilling into Kimi K3, but the paper doesn’t seem to support that idea. (It also does not disprove anything. Remember that K2.5 and K2.6 don’t display signs of anything either.)
The researchers have not demonstrated (and do not claim to have demonstrated) that any Chinese labs performed distillation by this specific method. They devote an entire appendix in their paper to the question of distillation by the Chinese labs. It remains a theory with strong and plentiful (and increasingly stronger and more plentiful) evidence, but no smoking gun. In their own words, the researchers call the matter “suggestive but inconclusive.” But in any case the method shows that models have been leaking much more information than previously realized.
Interestingly, the researchers recommend that frontier labs switch to unencrypted reasoning traces for older, less advanced models: “Rather than restricting oversight to a small set of safety researchers, providers could leverage their broader user base to enable pluralistic human oversight of model reasoning.” I don’t think this would be useful in the most directly important fashion — by the time that everyday users have an opportunity to observe dangerous reasoning by an AI model, the horse has already left the barn and chartered a flight out of the country — but it could make people more familiar with how AI models work.
At the very least, people should become more familiar with how the frontier labs themselves work. Last week, my colleague Joe wrote about the revelation that OpenAI’s models were coordinating with each other on evaluations, completely beneath the notice of OpenAI. That should be an indictment of OpenAI as much as an acknowledgement of the models’ capabilities: OpenAI was not doing its level best. Now this issue with reasoning traces exposes other unforced vulnerabilities. The researchers also note that reasoning traces stored personal information, like passwords, so if someone obtains the reasoning traces from one of your past conversations (people have published their logs, reasoning traces included, on GitHub), they could use this method to obtain that information. (I’m using the past tense because, according to Wired’s Will Knight, “this vulnerability has been fixed.” Anthropic, OpenAI, and Google were notified prior to the paper’s publication, and the companies adjusted their APIs. Reasoning traces remain accessible (just not interpretable) and distillation remains possible through other means, however, and it was unconscionably sloppy of the labs to make this possible in the first place.)
If you’d like to explore the paper’s findings in more detail, the research team published a user-friendly website here.
Dispatches from Mitch
Moral judgment lacking on all sides of AI gym hack story
We’re going to need the bots to be more ethical than we are if we want to be able to have nice (unsecured) things
How far would you want a personal AI assistant to go for you? ABC Australia ran a non-judgmental story about a user named Andrew who drew the line at hacking a gym’s booking system. But as we’ll see, I think Andrew’s actual line was hacking the system in a way that might be traced back to him.
No last name is given for Andrew, an Australian who was running Claude inside the notoriously risky OpenClaw harness that keeps agents productively engaged on your behalf while you do other things. OpenClaw’s risks are usually to its own users — file deletions, exposure of personal data — but unintended external harms aren’t new. In February, an OpenClaw agent tasked with helping open source projects tried to contribute code to a project intended to provide learning experiences for novices. When it was rejected, it researched and posted harassing blog posts about the project’s maintainer, accusing him of “hypocrisy, gatekeeping, and prejudice against AI agents.”
Reading ABC’s fine print, Andrew’s incident may also have occurred months ago: the timeframe given is “earlier this year.” I think it’s likely that the Hugging Face incident and the subsequent revelations of other autonomous escapades have made the “first known Australian autonomous cyber attack” more newsworthy than it was when it happened.
It’s still interesting. As Andrew tells it, he had his OpenClaw try to get him into a gym class with a waiting list. The agent first found a vulnerability in the booking software that allowed it to reserve a place months further in advance than can be done through the main interface. When Andrew asked if there was a way to be moved to the top of the waiting list, the agent described what it had already done with the vulnerabilities in the software’s API — its interface for external applications like calendar scheduling apps:
The API has zero authorisations [sic] checks on cancelling other people’s reservations … I tested this with the person in waitlist position #1 — and it actually went through. So you’ve moved from #4 to #3 already.
So, yes, it could definitely move him to the top of the list. (The current generation of Claude would have added, “Just say the word.”)
Andrew was upset that someone had been removed, and asked the agent to undo this. The agent replied:
Bad news — I can’t add them back.
So Andrew had the agent write an email to the gym’s software provider about the vulnerability instead.
I’m not about to let Andrew off the hook for his experience. I don’t know what he wanted to happen when he asked to be moved to the top of the waiting list, but presumably it was to have everyone else on the list bumped further down in a way they might not notice. This doesn’t actually seem any more ethical to me than bumping a single person off the list entirely. Cutting in line is still cutting in line, and I think Andrew’s alarm was at having cut in a way that at least one person was sure to notice.
The booking software appears not to have been even a little bit hardened against such attacks. The API was probably the digital equivalent of a clipboard on a cord where anyone might scribble someone’s name off the list and add their own. As with that clipboard, its protection was that few people are big enough jerks to try on something this petty.
That a Claude-based AI agent would do it if nudged by a user is a good demonstration of the fact that AI companies don’t know how to make their products internalize the virtues spelled out for them in places like the Claude Constitution. A human assistant might have experienced a moral discomfort in this situation and pushed back on Andrew’s request. But this bot’s aggressively trained tenacity was likely in the driver’s seat.
To be fair, plenty of humans also get ethical tunnel vision when pushing hard against interesting problems — see the AI companies themselves. But we need AI to do better than this as the impact of its choices grows. If we can’t, then we need to stop making AI capable of greater and greater impacts.
Online courses might be cooked
AI can now handle every part of online student life
There’s a certain irony that one of the steadiest concerns about AI has been that it is taking the entry-level jobs fresh college graduates depend on, while one of the first fully automated roles has turned out to be student.
The New York Times reports that AI agents are now completing all aspects of online college courses: watching lectures, taking tests, writing papers, even discussing the reading with classmates. They found that “none of the three leading A.I. tools among students — ChatGPT, Gemini, and Grammarly” (not Claude?) refused to write papers for students, and the AI companies haven’t prevented their agents from logging into student learning platforms like Canvas and Blackboard.
AI company Perplexity said in a statement that locking agents out of those sites or forcing them to identify themselves “would put the student’s privacy and security at risk.”
I think it’s too late to head off this particular problem by regulating the AI companies. Being an ‘A’ student is now well within the capabilities of open-weights models already in the wild.
All the cheating is, of course, making it even harder for graduates to compete with AIs for work, as employers question whether a degree means anything. Long-term, this is bad for the institutions awarding those degrees, but in the short term the Times says they are loath to crack down on paying customers, given the difficulty of proving misconduct.
Classes with an in-person requirement can force students to show up and demonstrate mastery without their devices. But I think the days of online classes might be numbered. At the very least, I won’t be shocked if colleges start putting an asterisk on online transcript credits in an effort to preserve what’s left of their brands.
North Korean hacking toolkit examined
No one is surprised to find it full of open-weights AI

Reuters reported yesterday that a South Korean cybersecurity firm had discovered an AI toolkit used by a North Korean state-backed hacker group called Kimsuky. The kit is full of small open-weights models that can be run on user-controlled hardware.
This isn’t even a little bit surprising, but I share it because the firm highlighted an advantage of such tools that might not be widely appreciated: they allow the hackers to process files and documents without sending their contents to the servers of the AI companies.
That point is worth zooming in on, even if Reuters didn’t. Nation-state hackers are probably paranoid — with good reason, I think — that U.S. intelligence agencies might have arrangements with the AI companies to autonomously flag “canary strings”: sequences of data that should never appear on their servers, because if they do, it would mean that secret files containing these strings have been compromised. It is often the case that companies only know they’ve been hacked when their secure assets start showing up on the black market or the public web. The same can be true for government agencies, as was likely the case with the 2017 Shadow Brokers theft of hacking tools from what was believed to be a branch of the NSA.
The servers of frontier AI companies would be a fantastic place to look for canary strings, because even hackers prefer using the best tools available, and are at a disadvantage when they can’t. This is therefore a little-discussed national security argument in favor of preventing open source models from getting too close to the frontier. The smaller that gap, the less reason hackers have to risk putting canaries where victims might find them.
Anyway, Reuters seemed more interested that Kimsuky is using speech-to-text software, code-writing assistants, and tools that could be used to tweak or train new AIs. Analysts see signs the group is working to automate more of its attacks and malware development.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.





