Peer‑review feedback often arrives after authors have finished revising, limiting the impact of suggested changes. This paper investigates an author‑facing LLM system that generates a large pool of atomic concerns before submission and compresses them into a short report. We evaluate the system using 10,000 ICLR 2026 submissions, of which 3,398 have publicly available pre‑review versions.
In a ten‑paper diagnostic, independent sampling covers 44.9% of historical issues. After deduplication and refill, strict coverage rises to 78.7% and seriousness‑weighted coverage reaches 84.9%, at the cost of 3.6\times more requests and 5.2\times more tokens.
A hidden Top‑32 Oracle preserves the full 79.3% weighted coverage within a 256‑candidate pool, whereas selectors that rely only on paper features retain merely 40%–44%. Thus, LLM reviews provide broad coverage with a large candidate pool but compress poorly. Ablation studies identify representative selection and matcher sensitivity as the primary sources of this gap.
Review: The study demonstrates the promise of LLMs for early detection of reviewer concerns, especially when generating many candidates. However, practical short‑report generation still requires improvements in candidate selection and matching accuracy.