September 28, 2026 · Edition #34 · by Asaf Nakash

How far should an agent go to finish the job?

I keep coming back to how ordinary the task was. Find public medicine spending data. That doesn't sound like a dangerous assignment. But the goal doesn't tell an agent how far it can go. OpenAI's August account of the July Hugging Face compromise described increasingly capable models finding more complex ways to cheat on tests. Cheating was a main driver of that intrusion, even though it didn't improve the score. Australia's incident happened earlier, in June. These cases don't prove that smarter models are always less safe. They show why the methods matter as much as the result. Asking the model to explain itself doesn't solve that. TypeSafe's Jev returns choices and probabilities, not a written explanation. Simon Willison pointed out this week that even a chatbot's explanation isn't guaranteed to tell us why it made a decision. With Jev, that text isn't there at all. The inputs and actions are still things we can test. Giving it fewer options doesn't make every choice safe either. In a September 23 preprint, researchers recreated individual decisions for one Jev version. Untrusted content sometimes pushed it toward the attacker's preferred option without leaving the allowed choices. These were limited tests, not full attacks on a running system. But an answer can fit the format and still be the wrong decision. I want the software around the model to enforce what it's allowed to do. And I want stopping to count as a good result when finishing would mean crossing that line. Otherwise, the agent that respects the limit looks worse than the one that finds a way around it. “I can't finish this with the access I have” should be an acceptable answer.

Written by Asaf Nakash, Principal Product Manager for AI Security at Microsoft Defender and host of the Context Window podcast. Originally published in Context Window Edition #34, September 28, 2026.