Reinforcement learning with verifiable rewards (RLVR) has markedly improved the mathematical reasoning of large language models. Recent work injects search into RLVR rollouts to increase trajectory diversity, yet diversity alone does not guarantee that the search‑induced rollout policy surpasses the current one. To fill this gap, we introduce APIVIS, a training‑time framework that adapts finite‑budget Gumbel search to chunk‑level mathematical reasoning. APIVIS merges direct and searched responses within each rollout group, allowing improvements found by search to generate informative relative rewards. It also applies selective supervision to tokens improved by search, preserving a learning signal when uniform group rewards render GRPO ineffective. We prove that exact value‑guided selection raises the expected verifier reward at each searched state and that this guarantee extends to the full rollout policy; an approximate guarantee holds under bounded value‑estimation error. Experiments on widely used mathematical reasoning benchmarks and across model scales show substantial gains over competitive search‑based methods, confirming APIVIS’s effectiveness.
Review