When large language‑model agents interact with external environments such as web browsers over long horizons, the accumulated interaction history quickly exceeds the context window. To remain reliable, an agent must preserve factual information, remember execution errors and corrective feedback, and track progress across actions. Existing solutions either rely solely on reasoning and action histories, learn a fixed‑size memory via an overwrite mechanism, or periodically summarize the trajectory. Although overwrite memory can in principle store everything an append‑only memory can, it must carry each fact through every rewrite, which is hard to learn from sparse outcome rewards; in web‑agent experiments the learned overwrite memory tends to delete key facts and environmental feedback.
AMBER (Append‑only Memory Bank for Evidence Retention) introduces a simple, scalable framework: the agent simultaneously reasons, acts, and writes free‑form entries to an append‑only memory, and the append‑only rule guarantees that information is never overwritten. This allows end‑to‑end reinforcement learning using only outcome rewards, without the need for large curated instruction‑fine‑tuning datasets. On WebArena Lite, AMBER improves average success by 4.09 percentage points over overwrite‑based memory, raises the fraction of tasks solved across five repeated runs by 4.8 points, and matches an overwrite baseline that required substantially more expensive supervision. At the same time, AMBER stays within a practical token budget, striking a strong balance between context efficiency, task performance, and reliable long‑horizon execution.
Review