Voting over multiple LLM responses is a common primitive for test‑time scaling and ensemble inference. Collecting more calls expands the candidate pool and raises the chance of discovering a correct answer. Under a fixed call budget, a discovered answer still needs to gather enough support in the remaining calls to become the final plurality winner, creating a discovery‑to‑decision gap.\
We characterize this gap by the realized vote state and the remaining budget. A sharp recoverability threshold $R^{*}$ is derived, and we show that as sampling proceeds the observed candidate set can only expand while the set of reachable endpoint winners can only contract, yielding a candidate‑level conversion window. Assuming an i.i.d. response law, the same state yields exact finite‑horizon endpoint probabilities:$$P_{\text{end}}(s)=\prod_{t=1}^{B}\Pr\bigl(\text{vote}_t\mid s\bigr)$$where $B$ is the remaining call budget.\
We further prove that merging wrong‑answer identities preserves single‑call correctness and cannot improve plurality accuracy; the effect of redistributing wrong‑answer probability depends on the realized vote state. Singleton reachability provides a gold‑free exact locking certificate. For a known answer universe, the first trigger of locking is the earliest prefix at which all admissible continuations yield the same fixed‑budget output.\
Empirically, on a controlled Word16 study we find that:\
- Input permutation boosts raw plurality accuracy by $21.1$ points with essentially unchanged single‑call correctness;\
- Exact locking saves $28\sim30\%$ of calls at a $16$‑call budget while preserving every fixed‑budget output.\
Most discovered‑but‑unselected correct answers lose reachability only after discovery, indicating that the discovery‑to‑decision conversion window is typically narrow in practice.\
Review