We introduce Fuse, a multi‑agent simulation framework for studying user‑mediated social reasoning. In Fuse, a target agent with a hidden motive interacts with other agents, including a user agent who consults the evaluated LLM assistant to infer the motive. Because the motive is predefined in the simulation, the ground truth is verifiable by construction. A human study with 24 k annotations confirms the fidelity of the simulation. We then apply Fuse to 12 LLMs, systematically isolating key factors and observing that:
- User mediation compounds the inherent difficulty of social reasoning;
- LLMs are systematically sensitive to biased user framing;
- Models often need more detail than humans to arrive at a correct prediction;
- Longer conversations do not always improve performance despite offering clarification opportunities.
Fuse and a dataset of 21 k examples are released as open source.
Review