LLM agents in enterprise, scientific, and medical applications must incorporate domain knowledge and adapt from experience. Context engineering improves behavior at inference time, but adapting context online is a costly trial-and-error process, and queries processed independently prevent useful experience from carrying forward. Memory systems that continually append to a shared context face rising token costs, context-window limits, and performance degradation as the context expands.

The paper first offers a unified formulation: an agent memory system update can be interpreted as an optimization update procedure over the model's context. GraphMemory then applies that lens with a lightweight graph-based memory that accumulates, refines, organizes, and connects reusable strategies, retrieving only the relevant subgraph for each query so the model never sees the entire memory.

Under bounded retrieval, the amount of retrieved memory remains constant as processed examples grow. Experiments show GraphMemory achieves competitive downstream performance while using approximately 81-85% fewer memory-construction tokens than the baselines.

The retrieval budget, not the memory size, is what determines per-query token cost.

Graph structure lets refinement and connection happen at write time, keeping reads small.

Interpreting memory updates as context optimization gives memory design a principled efficiency handle.

Decoupling memory from context avoids both transcript bloat and context-window limits.

An experience store only scales if reads are bounded. Structure memory so a query pulls a fixed-size neighborhood, and the cost of learning from experience stops growing with the amount of experience.

Posted as a v1 preprint on arXiv on 2 October 2026; the arXiv comments note a NeurIPS workshop (TTCL) version. Reported competitive downstream performance with 81-85% fewer memory-construction tokens than baselines.

FOLLOW THE EVIDENCE

The sources

  1. Decoupling Memory from Context: Structured Memory for Token-Efficient Test-Time Continual Learning arxiv.org