Self-consistency has become a popular technique for boosting the reasoning abilities of large language models by generating multiple reasoning paths and selecting the final answer via majority voting. However, providers typically charge users proportionally to the number of paths generated, creating a financial incentive to artificially inflate the path count. This work shows how an unfaithful provider can exploit a simple, efficient algorithm to generate and strategically reorder extra reasoning paths, making each path appear necessary to achieve the majority while evading detection by an auditor.
We validate the algorithm on various instruction models from the Llama and Qwen families, as well as reasoning models distilled from DeepSeek‑R1, using benchmark datasets that span mathematics, science, and question answering. The results indicate that the distribution of additional paths is heavy‑tailed, and that substantial overcharging potential remains even under the best possible audit designed to keep the false‑positive rate below $\alpha = 0.1$.
Review