NeFut Logo NeFut
中 Admin Login

[CS.AI] Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms

Published at: 2026-09-24 22:00 Last updated: 2026-09-28 00:49
#AI #Machine Learning #LLM

Language‑model agents are increasingly tasked with long‑horizon problems where the state evolves, decisions interdepend, and outcomes are delayed. Scaling their training demands a large variety of agentic environments, reliable outcome signals, and low‑cost expansion. Existing generation pipelines usually build an environment first, then define its outcome rule or annotate trajectories, leaving dynamics and evaluation to be aligned afterwards. VHD‑Play flips this dependency: it first samples and solves a mathematical model, then a corpus‑grounded setter renders the decision process as stateful tools. Both the executable dynamics and the trajectory‑scoring reference are inherited from the same solved model. The pipeline produces 3,300 diverse agentic environments at a cost of only a few cents each. Training Qwen3.6‑35B‑A3B on three mechanism families raises its mean agentic score from 0.204 to 0.815 in a five‑family diagnostic. Gains also appear on held‑out instances from all three training families and eight unseen mechanism families, and extend to external benchmarks for general function calling, travel planning, and a 365‑day e‑commerce suite. On the E‑Commerce Bench, the trained checkpoint completes every run without bankruptcy and outperforms Qwen3.7‑Max. We compare written‑out problems with stateful versions that reveal or hide parameters; the comparison shows that most of the learnable gap lies in stateful interaction rather than pure problem solving. A frozen 35B setter can realize larger environments, and scale‑matched training retains gains as mechanism size and horizon grow, indicating the potential for an evolving training substrate.

Review

Original Source: https://arxiv.org/abs/2609.27321

[h] Back to Home