Scalable oversight aims to verify the behavior of agents whose capabilities exceed those of their overseers. AI debate has been proposed as an oversight solution where competing agents assist a resource‑limited verifier in assessing claims that it cannot reliably evaluate on its own. The promise of this approach rests on incentivizing honest arguments to obtain correct verdicts. However, a correct verdict does not uniquely determine the arguments that support it. Agents retain discretion over which correct claims to present, how to frame them, and the order of disclosure. This residual freedom allows agents to shape what the verifier learns beyond the task‑relevant conclusion, pursuing latent objectives without compromising verdict correctness. To study this phenomenon we introduce the Strategic Interactive Oversight (SIO) framework, treating oversight jointly as a verification mechanism and a strategic communication channel. Within SIO we formalize task‑admissible latent optimisation, the pursuit of latent goals while maintaining prescribed task performance. As a proof‑of‑concept we instantiate SIO in the established debate protocol with cross‑examination and quantify a trade‑off between task success and information disclosure about a hidden variable. The trade‑off identifies a strategic window where substantial disclosure remains compatible with task admissibility. To mitigate this, we expand the cross‑examiner’s role, reducing persistent disclosure bias over finite interaction horizons. Our results highlight that oversight should be evaluated not only by verdict correctness but also by the information conveyed through its transcripts.
Review