Compositional reasoning is essential for solving real-world problems: because training data are inevitably limited, models must generalize by recombining learned skills in novel ways. Post‑training methods such as reinforcement learning have markedly improved the reasoning abilities of language models, yet their impact on compositional reasoning remains unclear. We introduce a dependency‑graph framework that formalizes compositional reasoning and defines three progressively complex levels of compositionality. Empirically, we instantiate the framework with data‑structure tasks, which offer deterministic reward computation and a clear compositional structure. Our findings reveal a consistent decomposed‑to‑composed asymmetry: training on decomposed skills does not reliably transfer to composed tasks, whereas training on composed tasks more readily transfers back to decomposed ones. We provide a theoretical explanation for this asymmetry and further evaluate compositional generalization under length extrapolation, structural distribution shift, and transfer to tasks requiring unseen skills. Finally, a pilot study on real‑world tool‑calling benchmarks shows preliminary evidence that the decomposed‑to‑composed asymmetry extends to practical settings.
Review