NeFut Logo NeFut
中 Admin Login

[CS.AI] ACLArena: Agent Continual Learning in Multi-stage Post-training

Published at: 2026-09-24 22:00 Last updated: 2026-09-28 00:49
#AI #Machine Learning #Neural

Building general‑purpose agents for industrial deployment requires stitching together multiple capabilities, each typically acquired at a distinct training stage. Currently there is no well‑established recipe for Agent Continual Learning (ACL), and the trade‑offs among existing integration paradigms remain poorly understood. To bridge this gap we introduce ACLArena, a framework for comprehensive study, analysis, and evaluation of ACL.

We first construct a sequential training pipeline and conduct an in‑depth analysis of forgetting and generalization from two complementary perspectives: the model level (parameter drift) and the token level (semantic representation stability). Guided by these insights we systematically compare three representative approaches—multi‑teacher on‑policy distillation, self‑distilled fine‑tuning, and model merging—assessing their ability to recover previously learned capabilities while preserving newly acquired ones.

Experiments reveal predictable transfer patterns across stages: distillation is robust to forgetting but adapts slowly to new tasks; self‑distillation adapts quickly yet suffers from forgetting; model merging can balance both under certain conditions.

Based on the findings we propose a new ACL recipe: offline replay of high‑quality trajectories combined with a routed network of multiple LoRA experts, each specialized via reinforcement learning for a sub‑task. This design markedly improves the agent’s ability to learn across multiple domains.

Extensive experiments on four reasoning and agentic tasks, evaluated both in‑domain and out‑of‑domain, confirm the value of our analysis and the effectiveness of the proposed approach.

Review: ACLArena offers a rigorous analytical foundation and a practical LoRA‑routing solution that together advance the state of continual learning for multi‑stage agents, demonstrating strong potential for real‑world deployment.

Original Source: https://arxiv.org/abs/2609.23989

[h] Back to Home