Vision‑language models (VLMs) have achieved strong performance on multimodal reasoning tasks, yet they frequently produce answers that are plausible but incorrect. Self‑verification offers a practical way to improve answer reliability without external judges, but existing approaches usually rely on a single verification criterion or a fixed prompt, resulting in incomplete and unstable reliability estimates. We first conduct a systematic analysis of how verifier strength and prompt design affect verification performance. The study shows that stronger verifiers yield more trustworthy judgments, while verification performance is highly sensitive to the choice of prompt, with no single prompt consistently dominating across tasks. Guided by these insights, we propose MOTIVE (Multi‑View Self‑Verification with Reliability‑Guided Rethinking). MOTIVE evaluates each candidate answer from complementary verification perspectives and learns a correctness‑aligned reliability score via correctness‑grounded multi‑view verification learning. During inference, this score drives an “accept‑or‑rethink” decision: reliable answers are returned directly, whereas uncertain ones trigger a history‑guided rethinking process. Extensive experiments on diverse multimodal benchmarks and across various VLM backbones demonstrate that MOTIVE consistently outperforms strong self‑verification and self‑correction baselines. Further analysis shows that reliable verification improves accept‑or‑rethink decisions and reduces unnecessary reasoning turns, enabling more reliable and efficient self‑verification without an external judge.
Review