NeFut Logo NeFut
Admin Login

[CS.AI] Beyond Final Decisions: A Process-Centric Benchmark for Transparent AI-Assisted Peer Review

Published at: 2026-09-10 22:00 Last updated: 2026-09-12 06:35
#AI #Machine Learning #Open Source

Peer review is central to scientific quality control. Existing evaluations of AI‑assisted reviewing mainly focus on the overall quality of generated reviews or the accuracy of final decisions, offering little evidence on whether model decisions are backed by sufficient and reliable review evidence. We introduce a process‑centric diagnostic benchmark that uses (x,$z_s$,$z_c$,$z_r$,y) to denote paper content, summary, critique, suggestion, and decision. Heterogeneous review records from PeerRead, NLPeer ARR‑22, and OpenReview‑ICLR are converted into process‑aligned data. The benchmark adopts direct decision prediction from the paper content (Direct) as a baseline, compares the decision value of Gold‑process variables with Predicted‑process variables, and conducts stage‑level evaluation, chain‑consistency evaluation, and interventional sensitivity analysis. Experiments across three datasets and six models show that Gold‑process variables generally yield higher decision value. For the primary analysis model, the Gold‑Predicted gap remains stable across datasets and random seeds, and it is reproduced in most model‑dataset combinations. Although model‑generated intermediate review texts exhibit relatively high local consistency between adjacent stages, the final decisions are not consistently supported by preceding review evidence. Our benchmark targets AI systems designed to assist, not replace, human reviewers, providing a transparent and auditable diagnostic tool for assessing the reliability of their review processes.

Review

Original Source: https://arxiv.org/abs/2609.05947

[h] Back to Home