NeFut Logo NeFut
Admin Login

[CS.AI] Reality Monitoring in Conversational Memory of LLMs

Published at: 2026-07-29 22:00 Last updated: 2026-07-30 03:24
#AI #Machine Learning #LLM

In conversational AI, if a model cannot distinguish its own output from user statements, it will treat its mistakes as user-provided facts. This capacity, known as reality monitoring in humans, is linked to hallucinations, delusions, and confabulation; however, whether LLMs possess it remains untested.

We demonstrate across two experiments and six LLMs that source attribution depends on the structure of conversational memory: ceiling accuracy for self-generated content under minimal memory demands reverses to a fragile external-item advantage once episodic delay removes that shortcut.

Feedback reveals two failures: in some models, internal and external judgments swap; in others, accuracy improves while confidence decouples from correctness, dissociations invisible to existing benchmarks. This pattern across models implicates active, not aggregate, parameter count.

This suggests that as AI systems take on autonomous, multi-turn roles, evaluating what they know is not enough; tracking where that knowledge came from may matter equally.

Original Source: https://arxiv.org/abs/2607.23927

[h] Back to Home