NeFut Logo NeFut
中 Admin Login

[CS.AI] Choosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use Agents

Published at: 2026-10-05 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #LLM

Large language models (LLMs) often rely on long‑horizon sequences of tool invocations to solve complex tasks, where each call changes the task state and influences later decisions.

In such long‑horizon tool use, rewards based only on the final outcome provide weak credit assignment across the extended interaction trace. Step‑level rewards can give more targeted feedback, yet obtaining reliable step supervision usually requires human or LLM judgment, or extra rollouts to estimate the downstream effect of an intermediate decision.

We argue that an effective tool‑use agent should estimate the long‑horizon value of a candidate next tool invocation before actually executing it. This objective demands comparative supervision over alternative invocations under the same context, while logged trajectories contain only the invocation that was taken.

To address this, we introduce Comparative Inference for Tool‑use Agents (CITA). CITA trains a Comparative Inference Model (CIM) from paired signals that combine observed tool behavior, scalable supervision from a Bayesian tool‑graph simulator, and semantic judgments derived from LLM‑based comparisons. The resulting CIM learns to predict how likely a prospective next tool call will support final task success given the current context.

Across three tool‑use benchmarks and multiple backbone LLMs, CITA consistently improves Tool F1 and overall task success. Additional analysis shows that CIM provides accurate step‑level value estimates for competing tool choices.

Review

Original Source: https://arxiv.org/abs/2610.02330

[h] Back to Home