AI reviewers can generate many specific criticisms, yet a larger number of criticisms does not guarantee a higher‑quality review. A review may miss a consequential flaw or retain an allegation unsupported by evidence. These two error types require opposite corrections, but generation‑oriented systems and aggregate metrics often conflate them.
We therefore recast AI‑assisted review as evidence‑guided refinement of a structured concern set, treating omission and over‑critique as separate risks. Building on this formulation we introduce EquiReview‑R. It resolves existing concerns against localized evidence, searches for missing issues from both independent and review‑conditioned perspectives, and returns one of three actions: stop, continue, or defer.
To expose the failure mode motivating this design we constructed an evidence‑linked trajectory corpus, ReviewTrace. Retrospective analysis shows that nearly all concerns in a high‑recall review lack a definitive evidential disposition, and an earlier refinement step cannot revise them; revision must precede further search.
On a frozen cohort of previously unseen papers, EquiReview‑R meets the pre‑specified non‑inferiority criterion for major omission, reduces major over‑critique from 15.5% to 8.1%, and achieves a one‑sided omission upper bound of 9.9% while stopping on 52.4% of papers. Computation‑matched controls, paired experiments, and ablations demonstrate that the gain stems from the revision process rather than extra inference or shorter output.
We release the ReviewTrace corpus as an evidence‑linked resource for studying review revision, disagreement, and provenance.
Review