When selecting mathematical training data for large language models, a common organizing principle is topic, e.g., providing probability examples for probability targets. An alternative is to group by reasoning approach, i.e., choosing worked solutions that share the target’s method even if the mathematical domain differs. We compare these two matching criteria after fine‑tuning. Two counterbalanced $2\times2$ designs are used: the primary design crosses probability with combinatorics, paired with invariant reasoning and double counting (2,000 problems); the secondary design crosses number theory with geometry, paired with complement and pigeonhole reasoning (800 problems). Each cell is held out in turn as the target: same‑approach (SA) sources keep the target’s method but change the topic, while same‑topic (ST) sources keep the topic but change the method. Every source appears once in each role, so additive source‑quality effects cancel in the equally weighted aggregate contrast. Across five base models and three training seeds per design, SA outperforms ST in all 40 seed‑pooled model‑target comparisons. Model‑level gains range from 8.2 % to 16.2 % in the primary design (mean 10.8 %) and from 12.0 % to 16.0 % in the second design (mean 14.3 %); all ten 95 % confidence intervals exclude zero. Although ST sources are more similar to targets under embedding and lexical metrics, the SA advantage runs opposite to this measured resemblance. These findings identify reasoning approach as a more effective matching criterion than topic for mathematical transfer across the evaluated topic‑approach combinations.
Review: The work offers solid empirical evidence that prioritizing solution strategies over subject domains can substantially improve mathematical fine‑tuning of LLMs, guiding future dataset construction.