NeFut Logo NeFut
Admin Login

[CS.AI] Verifiable Social Reasoning with Fuse for LLM Assistants

Published at: 2026-09-17 22:00 Last updated: 2026-09-18 00:46
#Machine Learning #LLM #Artificial Intelligence

We introduce Fuse, a multi‑agent simulation framework for studying user‑mediated social reasoning. In Fuse, a target agent with a hidden motive interacts with other agents, including a user agent who consults the evaluated LLM assistant to infer the motive. Because the motive is predefined in the simulation, the ground truth is verifiable by construction. A human study with 24 k annotations confirms the fidelity of the simulation. We then apply Fuse to 12 LLMs, systematically isolating key factors and observing that:

  1. User mediation compounds the inherent difficulty of social reasoning;
  2. LLMs are systematically sensitive to biased user framing;
  3. Models often need more detail than humans to arrive at a correct prediction;
  4. Longer conversations do not always improve performance despite offering clarification opportunities.

Fuse and a dataset of 21 k examples are released as open source.

Review

Original Source: https://arxiv.org/abs/2609.17496

[h] Back to Home