NeFut Logo NeFut
Admin Login

[CS.AI] Recursive Synthesis for Long-Horizon Terminal Tasks

Published at: 2026-08-07 22:00 Last updated: 2026-08-08 01:08
#AI #Machine Learning #Recursive Synthesis

Recursive Synthetic Terminal Tasks (RST) is a recursive verified synthesis framework for constructing long-horizon terminal-agent tasks. Starting from verified seed tasks, RST extends the reference solution, realigns the verifier and instruction to the new workflow, validates the result in a fresh sandbox, and reuses accepted tasks as seeds for subsequent rounds. Across fifteen recursive rounds, RST produces 37,484 synthesized terminal-agent tasks at roughly $0.05 per task. Task difficulty increases substantially over rounds: the median reference solution grows from 67 to 374 lines, the median number of executed commands grows from 40 to 244, and DeepSeek-V4-Pro pass@4 drops from 90% at $R_1@@@MATH_BLOCK1@@@R{15}$. To demonstrate training utility, we collect rejection-sampled Qwen3.5 trajectories on the synthesized tasks and use them for supervised fine-tuning. Fine-tuning on these trajectories improves Qwen3.5-27B and Qwen3.5-122B-A10B by up to 10 points on Terminal-Bench~2, Terminal-Bench Hard, and Long-Horizon Terminal Bench, while agentic PPO lifts Qwen3.5-27B to 49.44%, 32.00%, and 22.07% on the three benchmarks, corresponding to relative gains of 20.0%, 41.2%, and 21.9% over the base model. Moreover, after 15 rounds, the recursion shows no ceiling: synthesis yield and validation rates remain stable as difficulty keeps climbing, indicating that the process can continue well beyond the scale reported here. Blogger's Review: Recursive synthesis provides an efficient and scalable solution for constructing long-horizon terminal tasks, with significant potential for improving task difficulty and model performance, making it worthy of further exploration and application.

Original Source: https://arxiv.org/abs/2608.05466

[h] Back to Home