NeFut Logo NeFut
中 Admin Login

[CS.AI] Budget Boundary Effects in Test-Time Mathematical Reasoning

Published at: 2026-10-01 22:00 Last updated: 2026-10-06 12:11
#algorithm #AI #Machine Learning

A cumulative token cap can intersect a mathematical derivation, forcing a test‑time controller to decide between strict stopping (cut off at the cap) and advisory continuation (let the current attempt finish). We evaluate this boundary choice by replaying 19,200 public traces in pairs, covering 120 AIME, BrUMO and HMMT problems and two archive configurations of a single model. Candidate order and a 16‑attempt cap are fixed, and answer selection is blind to reference answers and correctness labels.

Key findings:

  1. At a 4k token cap, most of the advisory accuracy gain comes from converting abstentions into correct answers; strict stopping loses an unfinished prefix that a completed‑only selector could otherwise exploit.
  2. Comparisons based on realized cost differ from same‑cap comparisons: in low‑budget settings, advisory 4k achieves higher accuracy than strict 8k at comparable mean completion cost, while in high‑budget settings its observed accuracy is 0.42 points below strict 32k despite using only 59% of the mean tokens. These aggregate comparisons do not establish equal‑compute superiority or accuracy equivalence.
  3. Greater candidate coverage does not guarantee higher answer accuracy: a log‑probability selector loses accuracy even as coverage rises, including after a source‑grade consistency repair. Same‑cap majority‑accuracy gaps shrink below 1.3 percentage points at 32k.

Recommendation: Budget curves should jointly report the cap, realized cost, eligible candidates, stopping rule, and selector details.

Review

Original Source: https://arxiv.org/abs/2609.38699

[h] Back to Home