The weight of AI agent security is shifting. Instead of checking results after the work is done, the field is moving toward confirming and monitoring each tool call an agent makes, on the spot. In Part 4 of this series, we noted that benchmarks have started to measure the commands that will actually run, not knowledge about the tools. The papers we cover today carry that yardstick out of the evaluation lab and into real operations: scrutinizing a command at the moment it executes, and scrutinizing the agent again when it wraps up its work.

Start with call-level monitoring. One study argues that as agents read data through tools and act on outside systems more often, they need oversight that attaches to every call with low latency. Existing permission systems can tell whether an agent is allowed to use a given tool. What they cannot judge is whether a particular call is a reasonable step toward the intent of the assigned task. An allowed call can still drift from that intent, and a misbehaving agent can even egg on other agents and combine calls to its advantage. So the researchers hold that every call must be verified, and they examine whether a small language model could take on that verification. The idea is a lightweight watchdog that can run on a company's own servers.

Monitoring is not free. The JOVE study splits a complex question across several language models and tackles the problem of choosing which intermediate results are worth paying to verify. Running alone cannot tell you whether a result is correct, and verification runs asynchronously, with its output used to improve later assignments. Under a long-term budget and a per-question latency limit, the researchers weigh whether to spend on execution now or on learning for later. Here the demand to verify every call meets the reality that verification has a price.

The range of what needs watching has widened, too. EvoRiskBench is a benchmark that measures risks arising while an agent carries out multi-step tasks in a workspace. It sorts the origins of risk into nine categories and the resulting effects into five. Risky cases are executed directly in an isolated environment, and outcomes are independently confirmed from execution logs. Grounding the verdict in what actually happened, not in guesswork, points the same way as the trend described in Part 4. A study of source preferences exposes another blind spot. Across 12 agent models in three domains, each model had preferred sources, and those preferences largely agreed with one another. An item that met one requirement fewer than its rivals was still picked about two-thirds of the time when it came from a preferred source, and it was almost never picked in the reverse case. In other words, the results of calls can skew even without any malicious intrusion.

The last issue is how work gets closed out. In tasks such as incident and outage investigations, where someone must judge whether the gathered evidence is enough to close a case, an untrained 9-billion-parameter model overstated the evidence in 97% of its answers. Even a top-tier model that identifies the cause correctly 84% of the time overstated in 91% of cases, and it closed 17 of the 41 cases whose official conclusion was "cause unknown." The researchers also point out that the source of a case alone predicts the label fairly well: a rule that reads only the source reaches a balanced accuracy of 83.0. Being able to get the answer right and knowing when it is safe to close are two separate abilities.

Here is that flow on a single page.

The shift in when we checkCheck after the factConfirm permissions onlyLearn the outcome only after seeing itMonitor every callSee whether the call fits the intentConfirm with execution logsJudge whether it is safe to close

In short, on top of permission checks that only ask whether something is allowed, three layers are stacking up: the intent of each call, the execution record, and the judgment to close.

So solo founders and planners who put agents to work should look at three things. First, whether the calls an agent makes are logged in a way that shows they match the intent of the task you handed over. Second, whether the cost and delay of verification stay within a range you can bear. Third, when an agent declares that it has enough and closes the work, whether its grounds can actually be checked. These papers are all research-stage results, so please don't read them as already applied in the field. Still, one thing is clear: the question in AI agent security is changing from "Can I trust this agent?" to "Can I trust this call, right now?"