This paper focuses on the pathways for recursive self‑improvement (RSI) toward general intelligence, contrasting macro‑level language‑model scaling with the interaction‑driven principles of the “Era of Experience”. Regardless of the macro approach, the underlying reinforcement‑learning (RL) update engine must support cumulative adaptation. Recent algorithmic self‑discovery produced Disco103, which outperformed PPO to achieve state‑of‑the‑art benchmark results, yet its internal update machinery remains a black box.
We conduct the first causal mechanistic audit of a self‑discovered RL rule, structuring the analysis around the five pillars of the Era of Experience: extended horizon, grounded reward scales, continuous streams, within‑lifetime change, and exploration depth. The audit manipulates recurrent states by pinning, freezing, or transplanting them while keeping meta‑parameters fixed, thereby testing when learning history serves as an asset or a burden.
The audit yields three key findings:
- Recurrent history actively expands usable reward scales, extending the effective window from three decades under zero‑pinning to six decades.
- Decoupling historical content from its maintenance shows that the penalty of mismatched history stems mainly from perpetual clamping; allowing imported state to evolve naturally attenuates this burden.
- Under environmental change, controlling replay retention reverses the apparent adaptation advantage over DQN, indicating that external data turnover can confound internal plasticity.
These results are validated through capability thresholds and successfully transferred to a second rule (OPEN), grounding macro‑RSI ambitions in micro‑level learning dynamics and establishing an audit standard for next‑generation self‑evolving RL algorithms.
Review