We introduce the Recursive Self‑Rewrite (RSR) framework, which leverages a single base model, Qwen‑3.8‑27B, to discover successful solutions under diverse specialized harnesses and reconstruct them as training trajectories under a general harness. RSR consists of three components: a planner extracts successful solutions into executable runbooks; a critic screens for verifier or solution leakage and guides recursive revisions; an executor runs qualified runbooks in fresh sandboxes. Experiments on roughly 3 K self‑curated terminal tasks show that three harnesses together solve 759 tasks, a 34.3% gain over the strongest individual harness in the pool. RSR expands 2 001 source trajectories into 11 094 rewritten trajectories for supervised fine‑tuning. Fine‑tuning on these trajectories outperforms both the base model and direct trajectory SFT across several benchmarks: pass@3 on Terminal‑Bench 2 rises from 57.0% to 74.2%, on Terminal‑Bench 4 from 1.5% to 9.1%, on our hard self‑curated set from 39.0% to 63.0%, and on the Software set from 3.0% to 6.0%. Process reward on Long‑Horizon Terminal‑Bench increases from 0.21 to 0.29. These results demonstrate that diverse harness‑assisted experiences can be reconstructed into reusable capabilities for a model operating under a general harness.
Review: RSR provides a practical route to boost model generality and performance by unifying trajectory rewriting, offering a promising direction for future self‑supervised task learning.