The center of gravity in AI agent benchmarks is shifting. The old question was whether an AI knows about a tool. The new one is whether it can produce commands that actually run and carry a job through to the end. In Part 3 of this series, we noted that verification of AI output is moving from a single score to claim-by-claim checks against evidence. Today the same shift shows up in how tool-use ability is evaluated. A high score is no longer enough; people are asking what that score was actually looking at.
Measuring Execution, Not Knowledge
The KaliBench researchers argue that existing evaluations are either knowledge quizzes or tests that hand over an entire task, so neither directly measures the ability to build a command correctly. Security work runs on command-line tools, and one small slip, such as an option value bound to the wrong flag or arguments in the wrong order, can make a command useless. So the team collected 8,504 pairs that translate a natural-language request into a command line, spanning 1,642 tools, 23 capability dimensions, and 5 security stages. They also normalize options that mean the same thing and use alias-aware scoring. The design aims to check whether a command will run, not whether a sentence looks plausible.
Argo-Bench points the same way. Its researchers say existing text-to-SQL evaluations look only at query generation, and that their answer keys are often wrong. So they built a warehouse for an imagined New York delivery platform: 81 million orders from 2024, 235 tables, and 7.5 billion rows. They then posed 210 data science and analytics tasks. These are not single-table problems. They test whether a model can analyze across dozens of tables and act on the results. HumanoidToolBench extends the idea to robots. It evaluated seven policies in simulation and three on real robots across 18 tasks, and found a large gap between choosing the right tool and completing the task. Picking and finishing, in other words, are different abilities.
How Not to Be Fooled by Surface Scores
The clearest case for measuring execution is "Keyword Harnesses Fail Open." An evaluation that only checks whether keywords appear can award points for tool use that never happened. The researchers compared two Spanish-language security models with the same architecture. A 660-million-parameter model and a 1.1-billion-parameter model had nearly identical loose tool-use scores, 0.660 and 0.650. But when the models were asked to reproduce their training examples verbatim, the results split. The smaller model produced a valid tool call on all 6 examples. The larger one produced none of the 6. Looking next at first-token probabilities, the researchers found that the larger model gave the special token that opens a tool call a probability of roughly one in ten thousand to one in a hundred thousand. They interpret this as the web-focused training stage having erased that token's probability.
The diagnosis narrows down in this order.
The lesson: identical scores can hide opposite real-world calling ability, and a few cheap checks are enough to expose the difference.
So What Should You Look At?
When you read an AI agent benchmark score, look at the grading method before the number. The same score means different things depending on whether it checks that keywords match or that the command actually runs, and whether it stops at choosing a tool or follows through to finishing the job. If you are bringing an AI tool into your own work, you are safer running a handful of execution tasks similar to your real work and seeing whether the tool finishes them, rather than trusting the average score in a vendor's brochure. Going forward, watch whether evaluations built on execution-based grading become more common, and whether high-scoring models hold up under checks like these.



