Large language models (LLMs) are increasingly building multi‑agent workflows that decompose a complex task and assign specialist agents from a pool. Determining the granularity of decomposition, which agent handles each subtask, and when to introduce a new specialist are critical decisions that must be made beforehand. Whether a subtask succeeds is unknown until the workflow runs, and improving the workflow is costly. Fault localization typically requires a reference answer, a graded outcome, or a trained assessor, and fixing the workflow involves re‑execution, re‑search, or retraining. We introduce InFlowOp, which assigns a label‑free cost to every decision, weighing how well an agent’s competence matches a subtask’s demand against the agent’s runtime. Before execution, InFlowOp bidirectionally determines task granularity and agent assignment based on this cost rather than a fixed template. During execution, the same cost guides the cheapest correction of any fault. To address workflow‑level evaluation, we present Braid, a benchmark whose tasks require coordination beyond a single agent’s capability. Across diverse domains and model backbones, InFlowOp outperforms single‑agent baselines by up to $+11.97\%$, achieving $+9.64\%$ improvement with in‑flow optimization. Project page: https://xhguo7.github.io/InFlowOp/.
Review