The question behind AI agent memory management is changing. It used to be how much memory an agent could pile up. Today's papers ask what to keep, when to shrink it, and how an agent improves the more it reuses the same material. The emphasis has moved from storing to choosing.
This picks up where the last installment left off. Part 1 argued that an agent becomes more reliable when it remembers the procedures and past experience of repetitive work and reuses them. Today starts from the next problem: that memory cannot grow without limit. The issue is which stored memories are actually useful, and when to tidy them up.
Start with a case that does not shrink anything. VISTA lets the model see its environment directly and gives it a visual memory that keeps past observations in their original form, with nothing lost. When needed, the model pulls those observations back up itself and rebuilds its input from them. According to the researchers, Claude Opus 5.0's relative human action efficiency score on ARC-AGI-3 rose from 40.68 to 100.00. It finished all 25 public games, using 57.4% fewer actions than a person playing for the first time. The point is that the design worked by keeping the originals, rather than trimming them into summaries, and retrieving them on demand.
On the other side is research that wants the agent to learn when to shrink. AutoCompact targets coding agents that work across an entire code repository over long tasks. As work proceeds, what the agent examined earlier goes stale, so the researchers treat context management as a matter of judgment, not just avoiding overflow. They train three things as part of the agent's policy: when to compress, which task state to preserve, and how to continue after compressing. The training data is built in the following sequence.
In short, the agent does the compression itself, and the method teaches it by having a separate judge check whether that compression was right and correcting it.
That leaves the question of what to use when choosing which memories to keep. Causal Memory Policy points to a gap here: if a memory is never retrieved, there is no way to know whether it helps. Tinkering with the memory store alone produces identical results either way. The researchers intervene directly in the retrieval step. They set aside a portion of the context slots and fill them with memories drawn at known probabilities, which lets them estimate each memory's usefulness. It is a study that confronts head-on the fact that you cannot tell whether an unused memory is safe to discard.
Finally, SourceLearn looks at the situation where the same material is used again and again. Existing approaches treated reusing a source as mere repeated access. The researchers see it as a chance to deepen understanding of that source. They propose building a source model that captures how its knowledge is organized, interpreted, and used, and refining it over time.
The four studies head in different directions. One protects the original, one learns when to compress, one measures the usefulness of the memories worth keeping, and one builds understanding from the same material. What they share is a view of memory as a problem of selection, not of storage capacity.
That leaves three things for practitioners to watch. First, check whether a tool that advertises AI agent memory management makes its criteria visible: what it keeps and what it cuts. Second, check whether work continues after compression, meaning whether the summaries are validated by results. Third, check whether performance really improves when the same material is used repeatedly. One caveat: every case today comes from a research setting. Whether the same effect shows up in your own work is a separate question, so start by finding out how the tools you use handle memory.



