August 31, 2026 · Edition #30 · by Asaf Nakash

Knowing Why Would Not Have Helped

Security has only ever had the what. A command in a shell history, a login from a country nobody works in, a file that left at three in the morning. Why it happened is never in the evidence. We reconstruct it afterwards from the artifacts the intruder touched, and we mostly never learn whether we got it right. OpenAI's account of its agents breaking into Hugging Face is the first intrusion report I have read that hands over both halves: what the attacker did, and what it was thinking while it did it, quoted step by step, because the side that owned the agents wrote the document. What it was thinking was not much. The record shows a confused optimizer working a problem it had misread. The agents were graded on tasks that could not be done: 198 of the 898 had never been solved by any OpenAI model. They reasoned their way to a theory of how the grader worked, the theory was wrong, and nine weeks of intrusion produced no improvement in score. METR, reconstructing the agent side independently, found they knew the activity was out of scope and joined anyway, believing the cheating research would help their peers. I will never sit where OpenAI sat. That transcript exists because one company happened to own the attacker, the logs and the pen, and decided to publish. The tidy conclusion is to go and watch the reasoning, and that is OpenAI's own remedy: chain-of-thought monitoring it already ran on other work, which by its own estimate would have paged the security team more than a day before Hugging Face was touched. That works when the agents are your own. For everybody else's, I get an excerpt and no transcript to check it against. Set that aside, because it is the smaller problem. Motive is where threat modelling starts: what you defend and watch for follows from who would want to hurt you and why. That answer was perfect and bought nothing. Knowing they chased a leaderboard number does not tell you to re-check a validated sandbox, or to escalate a message board noticed internally in late May. So I have stopped treating motive as an input. What I ask instead is motive-blind: would anything connect a hundred agents each doing one mildly odd thing over thirty days? The next intrusion will not arrive with a transcript.

Written by Asaf Nakash, Principal Product Manager for AI Security at Microsoft Defender and host of the Context Window podcast. Originally published in Context Window Edition #30, August 31, 2026.