EvoSteer introduces an online self‑evolving graph orchestration paradigm. Existing multi‑agent orchestrators typically revise the team only after a trajectory finishes, which leads to coarse credit assignment and uncalibrated skill admission. EvoSteer builds the team on‑the‑fly, using execution features and a learned value estimate to repair plausible but failing steps.
Key technical components are:
- Anchored Trajectory Balance (AnchorTB): a regression‑style flow‑matching loss. It balances sub‑trajectories against a frozen reference trajectory, assigning a coefficient to each orchestration action. Formally, $$L_{\text{AnchorTB}} = \sum_{(s,a)} \left( f_{\theta}(s,a) - \hat{f}_{\text{ref}}(s,a) \right)^2$@@@MATH_BLOCK1@@@f{\theta}@@@MATH_BLOCK2@@@\hat{f}{\text{ref}}$ the reference flow.
- Validated Skill Admission: a candidate skill is trialed before promotion; it is admitted only if paired evidence passes a sequential test within a shared nominal testing budget. This ensures that only effective skills are retained and prevents wasteful skill accumulation.
AnchorTB also incorporates task‑level reference reward statistics together with prefix‑dependent corrections, allowing credit assignment to respect both global reward distribution and fine‑grained local decisions.
Experiments on twelve datasets span question answering, mathematical reasoning, code generation, and interactive decision‑making. EvoSteer consistently outperforms baselines, achieving up to a 15% gain on challenging reasoning tasks.
The code is released at https://github.com/beita6969/evosteer
Review: EvoSteer’s real‑time repair mechanism and rigorous skill validation make multi‑agent collaboration more reliable, offering a promising direction for practical deployments.