Reinforcement learning (RL) has been shown to boost the reasoning performance of large language models (LLMs) on complex math and programming tasks. This improvement, however, comes with systematic length misallocation: models spend excessive reasoning steps on easy questions while terminating early on harder ones, degrading inference efficiency with negligible accuracy gains. Many length‑adaptive methods mitigate this by allocating token budgets according to question difficulty, implicitly assuming that harder questions benefit monotonically from longer reasoning. We find that the impact of reasoning length on accuracy is concentrated on partially solvable questions. Further analysis reveals that explicit length rewards can induce unintended training dynamics. Motivated by these findings, we propose CARE (Contrastive Accuracy Reward Estimation), which compares the beneficial length adjustment per question from online sampled responses and applies adaptive length rewards within Group Relative Policy Optimization, without extra hyper‑parameters or inference cost. Experiments across multiple reasoning benchmarks demonstrate that CARE improves Pass@1 by up to $4\%$ while reducing reasoning length by $37\%$, achieving higher token efficiency. Code will be released upon paper acceptance.
Review: CARE’s contrastive reward estimation provides fine‑grained control over reasoning length, simultaneously enhancing success rates and cutting computational overhead, offering a practical pathway to balance efficiency and effectiveness in real‑world reasoning applications.