We build a simulation framework grounded in the World Values Survey (WVS) where culturally diverse LLM agents with distinct communication styles engage in longitudinal, value‑laden discussions. The study covers roughly 4,000 conversations, 1,200 personas, 15 topics, and three models: GPT‑4o, Gemini‑2.5‑Flash, and Gemma‑4‑E4B. We measure value faithfulness, value drift, and conversational realism. More than half of the personas fail to express their assigned WVS profiles from the first turn, and 2%–7% drift after repeated interactions. Ablations that remove demographic details improve faithfulness for some models but do not alter the overall pattern: simulated value distributions still deviate systematically from the assigned profiles. Compared with human dialogues, simulated exchanges show a different trade‑off between stylistic consistency and semantic diversity, often yielding content‑rich but stylistically repetitive conversations. These findings suggest that while current LLM agents can produce plausible dialogues, they remain limited proxies for faithfully representing and preserving diverse human value profiles over time.
Review