Long‑running LLM agents compress past interactions into persistent memories that are later reused as premises for new tasks. This creates a distinct derivation problem: does a memory truly follow from what the interaction history supports? Relevant evidence may be scattered across earlier turns, while compression can introduce relations or event statuses never established by the history. Consequently, a seemingly valid memory may appear unsupported because its citations omit crucial evidence, and individually supported facts can be combined into a stronger claim that the history never proved. We characterize the issue with three coupled requirements: evidence scope, compositional validity, and admission reliability. The core question is whether the interaction history available at write time supports what enters persistent memory. To address this we introduce DerivAudit, a framework that audits whether a memory is actually backed by the history at the moment of writing. The audit separates three questions: whether supporting evidence lies beyond writer‑provided citations, whether the composed memory introduces unsupported meaning, and how write‑time admission decisions affect later memory use. Audits on two natural memory corpora, using broader pre‑write history, recover support for roughly 60% of memories that appear unsupported from citations alone, while 17%‑21% remain unsupported after expansion. However, broader evidence alone does not make admission reliable: unsupported memories are still frequently admitted across verification models, and evidence expansion even worsens the issue on two backbones.
Review