Reasoning models often produce excessively long reasoning traces, making inference costly. Existing solutions either add early‑stopping mechanisms at inference time or incorporate length penalties during training (e.g., via reinforcement learning) to encourage shorter reasoning. This work introduces a different supervisory signal—confidence—trained in a self‑supervised manner. Specifically, using only 600 training problems, we fine‑tune models to predict their answer confidence at intermediate points along their own reasoning trajectories. Confidence serves solely as a training target; the loss contains no term for reasoning length, efficiency, or stopping. At inference, the fine‑tuned models follow the standard generation procedure without any confidence elicitation or early‑stopping logic. Experiments on Gemma, Qwen, Nemotron, and GPT‑OSS across mathematical, scientific, and coding reasoning benchmarks show that self‑supervised confidence fine‑tuning reduces generated tokens by up to $25\%$ while preserving accuracy, achieving efficiency gains comparable to methods that explicitly optimize for shorter reasoning. Further analysis reveals that confidence supervision largely preserves the base models' high‑level reasoning composition rather than merely suppressing specific behaviors. These findings suggest that efficient reasoning can emerge as a downstream effect of learning metacognitive signals, without direct optimization.
Review