Agent harnesses dictate how large language models (LLMs) collect context, invoke tools, verify outputs, maintain state, and decide when to stop, which heavily influences overall performance. Different tasks have heterogeneous requirements: a mechanism that helps one task may add overhead or distract context in another, making a single global harness sub‑optimal. We attribute this sub‑optimality to a mismatch caused by fixed mechanism choices and motivate the construction of task‑specific harnesses.
Generating harness code for every task incurs generation and debugging costs, and execution risks grow as more mechanisms are produced. To mitigate these issues, we mine reusable harness primitives from failed task trajectories; each primitive has a clear application scope and composition contract. Building on these primitives, we introduce STITCH, a framework that selects suitable primitives based on task information and compiles them into a task‑specific harness at test time, avoiding on‑the‑fly code generation or repair.
Extensive experiments show that STITCH markedly improves harness adaptability and robustness. As the primitive library expands, task success rates increase by up to 12 points over fixed‑harness baselines and surpass human‑crafted harnesses such as Codex CLI. Moreover, STITCH adds only a 2.7% overhead at test time, making it 638× more efficient than generating task‑specific harnesses from scratch.
These findings demonstrate that building task‑adaptive harnesses benefits diverse task completion, and that developing reusable primitives is a promising route toward that goal.
Review