Large language models are turning e‑commerce from static recommenders into interactive shopping assistants, yet real‑world shopping demands session‑level decision support—users gradually reveal and revise constraints, coordinate multiple goals, and expect product‑grounded recommendations throughout a full conversation.
Existing benchmarks mainly focus on outcome or execution metrics, leaving this evolving decision process under‑evaluated. We introduce REALWORLDSHOP, a benchmark built on a catalog of 3.28 M grounded products, structured shopping episodes, a profile‑grounded and action‑controlled user simulator, and role‑play evaluation.
Our analysis shows that current systems produce locally plausible responses but struggle with state tracking, constraint updating, and grounded convergence, especially in ambiguous intent, bundle, and multi‑intent scenarios.
To address these gaps we propose REALSHOP_AGENT, an executable session‑control framework featuring explicit state management, shopping‑flow control, catalog‑grounded retrieval, and runtime guards. Experiments demonstrate that REALSHOP_AGENT consistently outperforms strong baselines on REALWORLDSHOP.
Review