Multi‑agent LLM workflows typically consist of planning, execution, verification and summarization to boost task performance. The utility of each stage depends on the state already produced, so running every stage can waste compute or overwrite a correct intermediate answer. We cast skipping a stage as a counterfactual credit‑assignment problem: full‑workflow logs give the reward of the executed trajectory, while controlled skip interventions reveal the effect of omitting a future step. On this basis we introduce Learning What to Skip (LW2S), which learns action‑specific safety models from the interventions and combines held‑out calibration with domain‑native guards to decide skips. If an early skip is rejected, the controller can continue and reconsider later components. Experiments on mathematical reasoning, multiple‑choice QA and code generation, using two families of instruction models, show that LW2S reduces recorded token cost while matching or improving overall accuracy. Scale‑up and second‑topology studies examine component redundancy, and shared‑error cases demonstrate that agreement alone is insufficient for skip selection. The findings link efficient workflow execution to learning the conditional utility of individual components.
Review: LW2S leverages counterfactual interventions to assess component value on the fly, achieving notable compute savings without sacrificing performance, and thus offers a practical route for deploying multi‑agent LLM systems at scale.