Systematic literature reviews demand extensive human judgment across thousands of records, yet most existing evaluations of large language models (LLMs) examine only isolated stages of the review process. We introduce SciLitBench, a multi‑stage benchmark that spans title/abstract screening, full‑text screening, and schema‑guided data extraction. The dataset comprises 42,981 retrieved records, 1,012 full texts, and annotations for 888 included papers.
Across 22 open‑weight LLMs from six model families, providing explicit inclusion/exclusion criteria boosts the $F_2$ score of title and abstract screening by 28.8%, while researcher‑authored rationales improve full‑text screening accuracy by 15%. Data extraction reveals a distinct reliability regime: accuracy remains high for publication year (0.97) but drops to a Jaccard overlap of 0.37 for computational approach; even the strongest models recover only 30% of annotated evaluation evidence and 25% of limitations.
SciLitBench identifies a practical boundary between high‑recall screening and evidence‑complete extraction and offers a reproducible resource for evaluating LLM‑assisted evidence synthesis.
Review