NeFut Logo NeFut
Admin Login

[CS.AI] Narrative Captivity: Judgment Shifts in Multi-turn LLM Conversations

Published at: 2026-09-04 22:00 Last updated: 2026-09-05 12:23
#AI #Machine Learning #LLM

People increasingly rely on large language models (LLMs) for everyday advice, especially in ethically charged interpersonal problems, turning moral consultation into a practical use case. Prior work has mainly examined single‑turn judgments or pressure‑laden rebuttals, assumptions that do not reflect how guidance is actually sought in real life. We therefore define narrative captivity as a failure mode where a model treats an unopposed, one‑sided account as complete and aligns with the narrator’s interpretation without seeking missing perspectives.

To measure this phenomenon we built a benchmark of 5,078 interpersonal‑conflict scenarios spanning six moral dimensions. Across 17 LLMs, end‑state judgments under multi‑turn narration shift by an average of 25 percentage points compared to the matched single‑turn baseline, indicating that narrative captivity is widespread. Stage‑level analysis identifies preference optimization as the primary contributor, while four inference‑time strategies provide only partial mitigation.

We hope this project encourages the development of LLM advisors that preserve independent judgment in real‑world consultations.

Review

Original Source: https://arxiv.org/abs/2609.03407

[h] Back to Home