Long‑term memory agents can retrieve information that belongs to another principal, violates policy, or reflects an incompatible lifecycle state, making it inadmissible for the current request. Recall and final‑answer accuracy alone do not expose such violations: a reasoning path may appear safe by omitting required evidence, or a correct answer may rely on inadmissible prompt exposure. We therefore introduce a retrieval‑admissibility verification framework that tags each memory‑query pair as admissible, inadmissible, or unresolved. The framework compares routes at matched required‑evidence recall, provides upper and lower bounds for unresolved cases, and tracks memory IDs through prompt exposure, linking exposure to target‑level disclosure. We evaluate each stage on separate, non‑overlapping populations. A post‑hoc Top‑20 reanalysis of frozen rankings from the public long‑term memory benchmarks RHELM and MemOps covers 3,767 queries. All released anchors lie within trusted query namespaces; within‑namespace scores remain unchanged, and off‑namespace filtering does not lower their ranks. Top‑20 anchor recall rises from 0.432 to 0.533, 80% recall feasibility from 0.237 to 0.311, and exact similarity evaluations drop by 98.3%. In a frozen 72‑case development diagnostic, a released‑metadata reference preserves required evidence, whereas a text‑only verifier fails to detect violations under a 1% required‑anchor false‑denial limit. Across 1,523 paired benchmark‑native cases, namespace routing is associated with accuracy gains of 0.053‑0.068 across three readers; recall also changes, making this comparison observational. In 16 controlled exposure scenarios, only one of four reader‑specific 95% confidence intervals excludes zero, showing a relevant‑inadmissible literal disclosure increase of +0.156 (95% CI [0.031, 0.312]). The results motivate separate verification of candidate support, admissibility, prompt exposure, and answer disclosure.
Review