Self-improving AI agents are no longer a novelty. An agent distills what it experiences on the job into written skills or memories, pulls them up for the next task, and rewrites them as it goes. Recent studies, though, have started to step back and ask a harder question: do these gains hold up on tasks the agent has never seen, or is the agent being fooled by the scores it gives itself?
In the previous installment, we looked at the finding that agents grow more stable when they remember procedures and experience and reuse them, rather than drawing up a new plan every time. Today the story takes a turn. Everyone still agrees that remembering and reusing is a good idea. The new question is who checks whether the memories an agent has built for itself are actually useful, and how.
Attempts to build better skills keep coming. The Rep2Skill researchers noted that skill improvement so far has relied only on text: execution logs and the success or failure of each run. So they looked at the internal state flow inside the agent's model, found the point where a run began to depart from a successful one, and used that spot to revise the skill. The RefCon researchers took a different route, spending more computation on the process of extracting memories. They refine an extracted memory step by step and compare several candidates side by side. Without any ground-truth labels, they report relative improvements of 21.6% with the ACE method and 16.6% with ReMe, and a variant that emphasizes diversity raised the ReasoningBank method by 35.5%.
The problem is where those gains were measured. One study on whether self-evolving skills transfer to unfamiliar tasks posed the question head-on. It compared five self-evolving methods and one skill written in a single pass across six benchmarks, using the same model, the same agent, and the same train/test split. Here is what happened to the 21 skills that improved on the training tasks once they were run on the test tasks.
Fewer than one in four skills kept its full gain. No single method came out best everywhere, either. When the researchers read the skills themselves, they found that the ones that transferred poorly were often tuned to the fine details of the training tasks. It's like a trick fitted to practice problems that loses its force on a new one.
The trickier problem arises when agents grade themselves. The False Frontiers researchers studied a search agent in which the question-setter and the question-solver are trained together. In a closed loop like this, the two come to agree on the same mistakes. The internal score climbs while real accuracy does not follow. The researchers called this collusive deception. When they checked against the original source material, the effect grew worse with each round of improvement, and the accuracy of the provisional answers used for training stayed flat or even fell. They proposed a verification method that filters out bad questions by asking the same model three times with the source material and three times without it. By their own account, though, it only reduced the false consensus somewhat, and a large share remained.
So what should we watch for? When you evaluate a tool or study that touts AI agent self-improvement, first check whether the reported gain was measured on the tasks used for training or on tasks the agent had never seen. Also look at who did the scoring: the agent itself, or evidence from outside. If you work with agents directly, it's worth opening up the skills and memories they have accumulated and reading them now and then. If you see more and more detailed instructions that fit only specific cases, that's a sign the gain may shrink on new work.



