NeFut Logo NeFut
中 Admin Login

[CS.AI] Adversarial Closed-Loop Curriculum for Evolving Role-Playing Agents

Published at: 2026-09-25 22:00 Last updated: 2026-09-28 00:49
#AI #Machine Learning #LLM

Role‑playing agents built on large language models are widely used for personalized assistance and social simulation. Existing reinforcement‑learning approaches usually train on a fixed scenario pool collected before learning starts, creating a distributional bottleneck: as the agent improves, the hard scenarios shift while the training distribution stays static. To address this, we introduce AdvRole, an adversarial context‑rewriting framework that turns role‑playing RL into a closed‑loop curriculum. AdvRole alternates between an Actor that learns to role‑play and a Rewriter that edits character profiles and dialogue contexts into actor‑specific hard scenarios. The Rewriter is optimized with a performance‑gap reward, favoring rewrites that lower the current Actor’s score relative to the original scenario. Consequently, the scenario pool evolves together with the Actor, continuously targeting under‑mastered regions of the character‑context space. Experiments on three role‑playing benchmarks covering English and Chinese, plus a newly released multilingual benchmark, demonstrate that AdvRole consistently outperforms baseline methods.

Review

Original Source: https://arxiv.org/abs/2609.28609

[h] Back to Home