Large language models (LLMs) are increasingly deployed in enterprise, scientific, and medical domains, where they must incorporate domain‑specific knowledge and adapt from experience. Context engineering offers a practical alternative to weight updates by supplying instructions, strategies, and evidence at inference time, but online adaptation typically requires costly trial‑and‑error and queries are often processed independently, preventing useful experience from being carried forward.
Memory systems can retain information across interactions, yet conventional approaches continuously append new data to a shared context. This leads to exploding token counts, context‑window limits, and performance degradation as the context grows.
We present a unified formulation of context optimization and show that updating a memory system can be interpreted as an optimization step over the model’s context, providing a principled framework for studying memory design and efficiency.
Building on this perspective, we propose GraphMemory, a lightweight graph‑based memory that accumulates, refines, organizes, and connects reusable strategies. For each query, GraphMemory retrieves only the relevant subgraph, avoiding exposure of the entire memory to the model and dramatically reducing token usage.
Under bounded retrieval, the amount of retrieved memory remains constant as the number of processed examples increases. Experiments demonstrate that GraphMemory achieves competitive downstream performance while using roughly 81%–85% fewer memory‑construction tokens compared to baselines.
Review