NeFut Logo NeFut
中 Admin Login

[CS.AI] How to Conduct a Sensitive Debate: An Instance-Optimal Protocol for AI Debate

Published at: 2026-10-05 22:00 Last updated: 2026-10-06 12:11
#algorithm #AI #LLM

As large AI models now match or surpass human experts on cognitively demanding tasks, supervising these systems accurately has become critical. AI debate leverages a contest between two powerful AIs to decompose a hard question into simple claims that can be judged directly. Prior theoretical work formalized this intuition within computational complexity, aiming to design debate protocols (the game rules) that guarantee correct judgments under limited human supervision.

The best existing protocol works for all problems that admit a sufficiently stable decomposition into sub‑problems. This paper introduces a new protocol for the same class that improves on several fronts:

  1. Correctness holds in the worst‑case rather than only on average, so the protocol remains reliable regardless of the opponent’s strategy.
  2. Being honest and correct is a dominant‑strategy equilibrium for both debaters, unlike the previous Stackelberg equilibrium that required a leader‑follower hierarchy.
  3. By proving black‑box lower bounds, we show the protocol is instance‑wise optimal: no other protocol that only makes black‑box queries to human judgments can outperform it on this problem class.

Technically, we relate stable problem decompositions to the query‑complexity notion of fractional block sensitivity, denoted $\operatorname{fbs}(f)$. This connection yields tight upper and lower bounds on the number of human queries needed, establishing the protocol’s optimality.

Review: The work solidifies the theoretical foundations of AI debate by providing worst‑case guarantees and a dominant‑strategy equilibrium, reducing reliance on fragile average‑case assumptions. The black‑box optimality result offers a clear benchmark for future protocol designs.

Original Source: https://arxiv.org/abs/2610.02557

[h] Back to Home