As large language model (LLM) agents take on increasingly complex and long-horizon tasks, verifying their outputs becomes more challenging. This work investigates how to strengthen verification when only a fixed base model is available and no reference answers or grading rubrics can be accessed at test time. Repeated sampling yields multiple rollouts that may contain complementary correct claims, but a reliable verification mechanism is needed to decide which claims to trust. We observe that disagreement among model outputs often reveals correct alternatives, whereas consensus can sometimes hide errors. Motivated by these findings, we introduce VeriHarness, which turns the underlying LLM used by a generator into an agentic verifier by providing a workspace, evidence tools, and reusable verification skills. The system comprises a disagreement resolver that checks competing claims against environmental evidence, and a consensus challenger that probes shared claims and searches for omitted requirements. Their findings guide the selection and revision of the final artifact. Experiments on five long-horizon workspace benchmarks and two frontier models show that VeriHarness achieves the highest selection scores among all baselines. Evidence‑backed revision further improves average performance, adding 6.2 points for Gemini 3.5 Flash and 6.4 points for Claude Opus 4.8. Additional studies demonstrate that verification skills can self‑improve from failure feedback, confirming VeriHarness as a novel and critical approach for scaling long‑horizon agentic verification. We release approximately 26,000 rollouts covering all five benchmarks and both models, at a total cost exceeding $100,000, to support future research on agentic verification.
Review