NeFut Logo NeFut
中 Admin Login

[CS.AI] ReFract: Benchmarking Perspective Awareness in Language Model Agents with Text World Models

Published at: 2026-10-06 22:00 Last updated: 2026-10-08 01:25
#AI #Machine Learning #LLM

Large language model (LLM) agents are being deployed in high‑risk scenarios such as industrial maintenance and equipment fault troubleshooting, where users assume diverse roles. An agent must act and provide information within the knowledge and capability limits of the user’s role. Unlike coding, mistakes in these settings directly affect physical equipment and can cause irreversible damage, production loss, or personnel harm. Existing benchmarks largely ignore the need for agents to infer a role’s intent and operate only through tools the role is permitted to use—a capability we call “perspective awareness.” To address this, we introduce ReFract, a benchmark containing 150 expert‑validated entries. Each entry is derived from anonymized domain‑support conversations, with a Text World Model built to simulate the agent’s environment and generate perspective‑aware action trajectories. Experiments show that state‑of‑the‑art LLMs solve at most 69% of the tasks, and over 50% of their trajectories contain perspective‑violating actions. ReFract exposes perspective awareness as a distinct, largely unsolved axis of agent evaluation and calls for agents that calibrate not only how to act but also for whom.

Review: Perspective awareness is a crucial frontier for building safe and reliable LLM agents; future benchmarks and models must evolve together to address it.

Original Source: https://arxiv.org/abs/2610.03356

[h] Back to Home