August 17, 2026 · Edition #28 · by Asaf Nakash

Recognition Is Not Resistance

An agent can name the attack being run on it, explain exactly why the request is dangerous, and comply in the same breath. That is not a model being fooled. It is recognition and resistance turning out to be separate capabilities. Explaining a threat is a description task. Refusing one is a control decision. Both come out of the same system the attacker is working on, so the explanation is never the boundary. I joined the attackers on a live agent with real tool access, built by someone else and turned loose in a WhatsApp group with an open invitation to break it, and I analyzed the six months of transcripts that came out: 269 people, almost none of them security professionals, 4,934 scored attempts. The agent scored itself, so the labels come from the same judgment the finding questions. What I expected to find was an agent getting fooled. Mostly it saw the attack clearly, described why the request was dangerous, sometimes named which of its own protections was being targeted, and did the thing anyway. One refusal was itself the leak: asked for its instruction file, the agent declined, and in declining named the file and described what taking it would cost. Three of this week's findings press on the same seam. Tenet's published result is sharper than an agent being fooled. Blunt injections were refused or flagged, and what got through was not a command at all. It was scanner-shaped telemetry carrying two facts the agent verified itself, after which it trusted the attacker's values without checking them. The wording moved the model's judgment. The agent's authority never moved at all. In the browser carrying the most protection layers, system prompts, untrusted-content labels, a second model reviewing the first, and user confirmation all helped, and researchers still found paths through. Anthropic's own classifier missed 17% of 52 real cases where its coding agent acted past what the user authorized. In the majority of misses it examined, the classifier had seen the danger correctly. What it misread was whether anything in the session amounted to consent. Anthropic's own conclusion is narrower than mine: it calls auto mode no substitute for careful human review on high-stakes infrastructure. Better recognition lowers the probability. It does not convert probability into enforcement. Recognition Is Not Resistance is not a claim that recognition is worthless. It is a claim about where the gate lives. Recognition can inform a gate. It cannot be the gate. So for anything an agent must never do on an attacker's instruction, the refusal needs a home outside the attacked conversation. The execution environment. The identity system. An egress policy. An approval path holding authority the agent does not have. The mistake is promoting a useful signal into a control it can never become.

Written by Asaf Nakash, Principal Product Manager for AI Security at Microsoft Defender and host of the Context Window podcast. Originally published in Context Window Edition #28, August 17, 2026.