NeFut Logo NeFut
中 Admin Login

[CS.AI] EVOL: Simulator-Guided Evolutionary Expert Synthesis for Deployment-Free Learning Path Recommendation

Published at: 2026-10-06 22:00 Last updated: 2026-10-08 01:25
#algorithm #AI #Machine Learning

Reinforcement learning for learning‑path recommendation faces two coupled challenges. First, a policy must commit to a sequence of L concepts without intermediate feedback, causing a super‑exponential search space and delivering reward only at the final step. Second, ideal expert paths could alleviate sparse rewards, yet educational logs capture what learners did, not what they should have done. To overcome both, we borrow the simulator‑based demonstration paradigm from robotics. Using a knowledge‑tracing simulator, we first synthesize per‑learner expert demonstrations via evolutionary search, then train a deployment‑free policy that distills these demonstrations into a feed‑forward learner. The EVOL framework realizes this pipeline with an asymmetric actor‑critic: the actor plans blindly during both training and inference, while the critic exploits privileged simulator states only during training. Across ASSIST15, Junyi, and EdNet (39‑189 concepts) and path lengths 5, 10, and 20, EVOL outperforms eight baselines spanning heuristic, sequential, RL, graph‑enhanced RL, and LLM‑enhanced methods. Comparing three imitation objectives—behavior cloning (BC), advantage‑weighted regression (AWR), and DAPG—shows that final performance is driven by the quality of the evolutionary experts rather than the specific imitation loss.

Review

Original Source: https://arxiv.org/abs/2610.03273

[h] Back to Home