The reliability of a persistent agent memory hinges on when an assertion is written to the store. If a weakly supported claim is retained and later reused, it may be mistaken for an established fact. This work asks whether the retention decision should be governed by a confidence bar conditioned on the semantic category of the claim rather than a single global threshold. Well‑evidenced categories such as factual statements can be kept liberally, while categories with unreliable inference—values or beliefs—should be abstained from more aggressively.
We evaluate the idea in a cold‑start memory pipeline across 100 synthetic personas. Among 4,715 candidate assertions, only $77.9\%$ of value and belief claims are backed by their source, compared with $96.2\%$ for all other categories. A global confidence threshold cannot separate the two: it either admits many unsupported value claims or discards many well‑evidenced ones. Conditioning the threshold on category applies a stricter bar to values, reducing unsupported retentions from $6.2\%$ to $4.0\%$ (≈$36\%$ relative reduction) consistently across folds. At comparable retention, coverage improves by about 13 percentage points (95% CI $9.8\sim16.0$) relative to a global threshold.
These results suggest that reliable retention depends on the type of assertion rather than confidence alone. A category‑conditioned threshold acts as a simple, effective form of selective prediction at the write boundary, filtering low‑trust information.
Review