In the in-depth study of delayed generalization, or grokking, we identify an exactly solvable late-time relaxation mechanism for linear models trained with full-batch heavy-ball optimization and weight decay, extending this to nonlinear neural networks.
Our analysis reveals a distinguished population-active component called the grokking subspace. Along this subspace, training predictions remain unchanged, with weight decay as the sole restoring force leading to slow dissipative relaxation governed by exact discrete-time and continuous-time laws.
We show that only this subspace contributes to the slow asymptotic decay of population risk and derive explicit iteration-scale predictions for grokking time, recovering the familiar scaling of $(1-\eta)/(\eta\lambda)$ in the weak-regularization regime.
The theory further predicts distinct effects of optimizer choice, distinguishing coupled $L_2$ regularization from decoupled weight decay, and yields causal predictions for interventions modifying the grokking component.
We verify all theoretical identities without fitted parameters in a synthetic model where every subspace and relaxation rate is computable in closed form. Furthermore, we observe genuine delayed generalization in modular addition, where the measured delay follows the predicted scaling, and the late-time relaxation closely agrees with the theoretical clock.
Blogger's Review: This paper delves into the mechanisms behind the grokking phenomenon, presenting new mathematical models and predictions that hold significant theoretical and practical implications. The analysis of optimizer choice provides fresh insights into algorithm design and points future research in a promising direction. Overall, this study marks a milestone in understanding the generalization capabilities of deep learning models.