NeFut Logo NeFut
Admin Login

[CS.AI] Evaluating Human and LLM Screening Workflows in a Conceptually Complex Scoping Review: Recall–Workload Trade-offs and Run-to-Run Consistency

Published at: 2026-08-30 22:00 Last updated: 2026-09-01 02:31
#AI #Machine Learning #LLM

Large language models (LLMs) are increasingly employed for literature screening in evidence synthesis, yet the risk of discarding relevant studies remains. This preregistered study embedded in a conceptually complex scoping review compares human and LLM title‑and‑abstract screening workflows.

After a conservative title‑only screen, 1,131 records were screened by one review lead, four trained assistants (each handling a non‑overlapping subset), and seven complete LLM runs using different models and processing configurations, including a nominally identical repeat run. We measured retained workload, operational recall against 316 verified eligible records, agreement, run‑to‑run consistency, and procedural burden. Because eligibility was verified only for records advanced to full‑text assessment in the parent review, recall estimates are operational.

No workflow recovered all verified eligible records. Human workflows and two GPT‑5.4 file‑batch runs retained 42.2%‑45.0% of records while achieving 82.3%‑82.9% recall. Gemini 3.1 file batches achieved the highest recall (83.9%) but retained 56.7% of records. All‑at‑once configurations recovered fewer eligible records than their file‑batch counterparts. Two nominally identical GPT‑5.4 file‑batch runs agreed on 91.7% of records but differed on 94 records, including 29 verified eligible records retained by only one run.

The discussion highlights that LLM screening performance depends more on the implemented workflow than on model identity alone. Processing configuration, workload, record‑level variation, and human‑LLM decision integration are substantive properties of deployed systems. For high‑recall tasks, LLMs are better suited to validated, auditable, human‑supervised workflows rather than autonomous exclusion.

Blogger's Review: The study underscores that relying solely on model predictions cannot guarantee completeness in evidence screening. Thoughtful workflow design and human oversight remain essential to balance efficiency with high recall.

Original Source: https://arxiv.org/abs/2608.26885

[h] Back to Home