Automated harness optimization can markedly improve LLM agents by iteratively refining prompts, tool interfaces, and control logic based on execution feedback. Existing approaches mainly focus on how the harness itself is updated, while keeping the training scenarios that generate feedback largely fixed. As the harness evolves, the most informative scenarios also shift, suggesting that the training curriculum should co‑evolve with the harness. We formalize this missing dimension as an automated curriculum learning problem and introduce ActiveSaddler. ActiveSaddler models the evolving curriculum as a non‑stationary multi‑armed bandit whose arms correspond to dynamically instantiated optimization targets. It abstracts recurring failures into reusable failure‑pattern arms, estimates the potential learning gain from further targeting each pattern, and adaptively balances revisiting known weaknesses with exploring unseen scenarios. Optimization outcomes continuously update both the set of discovered failure patterns and their priorities, allowing the curriculum to co‑evolve with the harness. Experiments on GAIA2 and Terminal‑Bench 2.0 show that ActiveSaddler consistently discovers stronger harnesses, improving test Pass@1 by 4.4 and 7.5 percentage points respectively compared to a harness optimizer with a fixed scenario order. Ablation studies confirm that these gains stem from dynamically constructing optimization targets, estimating their evolving utility, and balancing continued optimization with new failure discovery. Together, the results establish automated curriculum learning as a crucial new dimension for harness optimization.
Review