Adjudicating debates requires tracking how arguments evolve through interaction, yet existing corpora rarely combine fine‑grained transcripts with professional judgments collected under a shared rubric in real competitions. We present a new dataset and benchmark for Chinese competitive debating, covering match, stage, and speaker levels. We organized 182 matches, recruited 120 professional judges, and had each match independently scored by three judges according to a predefined rubric.
After discarding matches with incomplete records, the collection contains 148 matches, 2,698 stages, and 20,542 exchange units, all manually verified for transcription and segmentation. It retains original stage scores, match votes, best‑debater ballots, and adjudication rationales. We define three tasks: predicting winner tendency, predicting stage scores, and predicting the best debater.
Zero‑shot evaluation of several large language models yields a highest winner‑prediction accuracy of 66.2%, a peak Pearson correlation of 0.250 between model stage scores and mean human ratings, and a best‑debater prediction accuracy of 56.8%. The dataset and benchmark provide a testbed for studying large models' understanding of interactive argumentation and their agreement with professional judges. Review