CounterRoute is an online reinforcement‑learning framework that learns both routing decisions and mode‑conditioned responses directly from a native dual‑mode checkpoint, without any method‑specific SFT warm‑up. The method pairs the current policy with counterfactual rollouts; cross‑mode credit is assigned solely to the routing token, while response tokens are trained within each mode using GRPO. Early training follows a paired‑to‑self‑routed curriculum: forced rollouts from both modes are generated first, then the proportion of self‑routed updates is gradually increased to stabilise exploration and improve autonomous routing. Across nine benchmarks, CounterRoute outperforms heuristic and learned adaptive‑routing baselines in balancing accuracy and efficiency. Compared with always‑thinking checkpoints, it raises macro‑average accuracy while cutting mean generated tokens by 51 % for Qwen3‑8B and 41 % for Qwen3‑14B. On instruction‑following and commonsense tasks where direct answering is strong, think rates drop to as low as 1 % and response quality improves. Although trained only on math and instruction data, its routing behaviour and response quality generalise to held‑out coding, science, knowledge and commonsense benchmarks. Review