The standard for checking AI results is moving away from a single score and toward confirming each claim against its evidence and source. In Episode 2, we noted that an agent can be fooled by its own score. Today's papers answer the follow-up question: if not the score, then what should we look at? The answer is claims that come with evidence, data that keeps its provenance, and evaluations that can be told apart from a pull toward form over content.
The first problem is how to judge a change after you fix an agent. One study takes issue with the common practice of swapping out a single component, such as a controller or a verifier, and then judging the change by the overall task score. A single score cannot tell you whether an improvement was possible in the first place, which component lost value, or what the agent's own self-checks actually guarantee. The researchers' answer, «Verify Claims, Not Scores», scores the evidence rather than the agent.
The verdict is one of four: supported, unsupported, unresolved, or not evaluated. The key point is that all three of these items are recorded together for every conclusion: the evidence, the verdict, and the boundary where it holds. To get there, the researchers use tools such as a reference policy that measures how much improvement is achievable within a stated range of actions. They also swap in a perfect version of one component at a time. That lets them trace whether a low number comes from the environment or from the way the evaluation was designed.
The same direction shows up in law. ARCCS, a regulatory compliance checking system, breaks regulatory documents into requirement units that cannot be divided further, then checks the document under review against each requirement. Each verdict comes with retrieved evidence, a confidence score, and an explanation a person can read. According to the researchers, in an evaluation on GDPR policy documents, LLM-based judges found the system's verdicts and reasoning legally sound. Instead of a one-line pass or fail, the structure asks for evidence on every requirement.
On the provenance side, there is PACE. For an agent that uses tools, the text it generates turns into real side effects, so a poisoned tool description, a retrieved page, a memory, or a reused skill can steer the next call. The researchers argue that screening what comes in is not enough, because a safe variant and a leaking variant can produce the same screening evidence. So PACE acts at the moment just before a tool call executes, checking the path along which influence has flowed against the authority that came from the request. As the memory and skill reuse we covered in Episode 1 grows, this last checkpoint matters more.
The last piece concerns the evaluators. A study of 137,293 votes from Compar:IA, a French-language LLM arena, finds that human preferences can reflect not only what an answer says but also how it looks. The analysis drew on 137,113 matchups involving 116 models. Bold text appeared as an association with 11.0% higher odds of winning, and the researchers named it one of the two associations that changed least across settings. Still, length, bold text, and lists tend to appear together, so it is hard to separate each one's contribution. The accurate reading is as a warning that formatting can seep into rankings.
In short, instead of one score, we now check three places: evidence, sources, and form.
So there are three things for practitioners who use AI tools to look for. First, check what evidence is attached next to a result and whether it states how far the result holds. Second, check whether an agent that draws on outside information or stored memory has a step that verifies right before execution. Third, be suspicious of whether a ranking or a review was swayed by attractive formatting rather than content. These methods are still at the research stage, but as questions to ask when choosing a product, they are useful today.
