In industrial settings, matching a query to an agent often fails when topical relevance is mistaken for executable capability, especially for long‑tail and boundary‑sensitive requests. We formalize the annotation task as capability‑bound process supervision and instantiate it with the Debate-to-Skill framework. The framework leverages reusable decision principles, structured deliberation, verifier‑based verdict extraction, and disagreement‑driven refinement to supervise the capability‑critical decision process directly. Experiments on an industrial Query2Agent benchmark compare Debate-to-Skill with direct‑label supervision, reasoning‑SFT, and several structural ablations. The results indicate that, in grey‑zone cases where semantic relatedness diverges from executable capability, supervising the decision process itself yields a notable accuracy boost.
Review