TimeEvo is a failure‑driven self‑evolution framework for time‑series question‑answering agents. Existing pipelines fix a tool library before execution, which creates two problems: human‑agent‑tool misalignment— even a curated set of 21 expert tools can hurt anomaly detection on certain tasks; silent harm— a generic self‑revision round changes 147 answers, breaks 56 of them, yet the overall score moves by less than a point. The root cause is that a tool’s usefulness is only known at runtime, while it is judged beforehand by a single average metric. TimeEvo clusters diagnosed failures into capability gaps, designs a measurement for each gap, synthesizes evidence‑only tools to fill them, and admits candidates through a paired admission gate. Experiments on ten time‑series QA tasks and three backbone models show that starting from an empty library, TimeEvo improves accuracy on every task and backbone, and a library grown on a cheap model still yields gains when transferred to stronger ones. The code is publicly available.
Review: By coupling runtime failure analysis with evidence‑based tool synthesis, the approach overcomes the rigidity of pre‑defined tool sets and offers a scalable, adaptive solution for time‑series analytics.