We introduce AREX-2, an effort to enhance the self‑improving capability of large language model (LLM) agents, defined as the ability to iteratively refine solutions at test time. This capability rests on two complementary mechanisms: reflection, which generates a solution better than the current one, and long‑horizon execution, which keeps the iteration effective over many rounds. We hypothesize that both mechanisms are domain‑agnostic and can therefore be learned in supervision‑friendly settings. Accordingly, we synthesize long‑horizon improvement trajectories from machine learning and algorithmic programming tasks, which provide verifiable feedback and reward sustained iteration. Trained on this data, our agent built on Qwen3.8‑27B achieves 81.8 on MLE‑bench Lite, 70.7 on Frontier‑CS, and transfers to deeper research tasks with scores of 84.0 on BrowseComp, 52.6 on HLE, 92.2 on GAIA, and 93.8 on DeepSearchQA, continuing to improve as the budget of rounds grows. These results demonstrate that long‑horizon reflective data is an effective route toward self‑improving agents.
Review