The solution
Long-horizon language-model agents keep accumulating interaction history, which raises computational cost and makes relevant information harder to preserve and reuse. Existing context management focuses on how to compress or retrieve history, largely leaving open whether the model itself already represents the need for those memory operations before they occur.
The authors studied the hidden state immediately before each agent action and found compression and recall needs are already encoded in the model's internal representations. The signals cannot be explained by simple context length or interaction progress and exhibit distinct formation patterns across model depth. They also found most memory-decision information survives in a compact recent context, with selectively restored historical evidence covering the long-range dependencies recent context misses.
Based on these findings, PaMER combines state-guided compression with evidence retrieval, and PaMER+ adds step-level evidence selection. On WorkBuddyBench, across multiple context-management baselines and model backbones, the framework substantially reduced context consumption while maintaining competitive task performance.
Why it worked
Steering memory from internal states acts before the context degrades, not after.
The signals are distinct from context length, so they carry information length-based heuristics miss.
Recent context already holds most memory decisions; retrieval only needs to fill specific gaps.
Step-level selection keeps restored evidence proportional to what the current step requires.
What can be applied
Internal states often anticipate the operations a system will need; predicting the need beats reacting to surface symptoms like context length.
Aftermath
On WorkBuddyBench the framework substantially reduced context consumption while maintaining competitive task performance across multiple context-management baselines and model backbones.
FOLLOW THE EVIDENCE