Faithful user simulation underpins the construction, evaluation and improvement of interactive AI at scale. Plausible single‑turn replies alone do not guarantee that simulated users reproduce the evolution of intent and outcomes seen in real interactions. We introduce TRACER, a multi‑turn user simulator that explicitly models the dynamic change of user intent and learns to align simulated behavior with real interaction trajectories.
TRACER is trained in two stages. First it undergoes supervised fine‑tuning on authentic user dialogues. Then it is refined through multi‑turn reinforcement learning. The RL phase combines hierarchical outcome rewards with trajectory rewards and incorporates deviation‑aware advantage modulation, which alleviates reward sparsity and resolves credit assignment in long conversations.
Evaluated on real customer‑service sessions organized into reference cohorts, TRACER‑7B outperforms the strongest baseline by 11.4 conversion F1, achieves the lowest group‑level conversion‑rate error and semantic trajectory distance, and generalizes well to out‑of‑distribution scenarios. Human Turing tests show identification accuracy near chance, confirming the perceived naturalness of the generated dialogues.
Building on this simulator we propose the Dynamic Marketing Benchmark, which jointly assesses persuasion effectiveness and response quality of LLMs. Results reveal that higher response quality does not necessarily translate into higher conversion rates.
Review: TRACER’s intent modeling and hierarchical reward design deliver a marked improvement in behavioral consistency, offering a more reliable platform for user simulation and shedding light on new directions for LLM research in marketing contexts.