Large language model (LLM) agents rely on external harnesses to shuttle information between the model and its environment and to recover from execution errors. Existing evaluations usually report only average task success, which hides an important tension: the same recovery action can rescue a failing trajectory or disrupt one that would otherwise succeed. We formulate recovery as a causal decision problem. Starting from an identical execution state, we compare outcomes with and without recovery, separating rescue from harm, and study how the value of recovery evolves over time. We then introduce the Causal Intervention Router (CIR), a lightweight policy that uses information available before recovery to decide whether intervention is worthwhile. In long‑horizon ALFWorld tasks with Qwen3‑14B, CIR raises success from 70.33% to 73.33%, a gain of 3.00 percentage points, while leaving all trajectories with correct observations untouched. Additional control experiments show that the benefit of recovery cannot be explained solely by the new observation returned by the environment. These results offer a practical way to evaluate recovery and apply it selectively.
Review