This study evaluates how the length of context windows affects the quality of literature reviews generated by large language models (LLMs) and examines the role of AI in supporting review writing. We collected source papers from Semantic Scholar and arXiv, prompting an LLM to produce 20 literature reviews under two settings: a short context window (≈2k tokens) and a long context window (≈8k tokens). Two researchers then assessed the reviews across 15 dimensions—accuracy, completeness, originality, citation coverage, among others—using both quantitative scores and qualitative analysis.
Findings indicate that AI‑generated reviews still require human oversight to meet academic publishing standards. While larger windows allow LLMs to incorporate broader information and maintain coherence over longer inputs, they also exacerbate issues such as content repetition, omission of pivotal works, and a tendency toward descriptive rather than synthetic discourse.
We conclude that AI‑generated reviews can serve as useful overviews, but their outputs must be critically evaluated and refined by domain experts. Future research should integrate additional LLMs or fine‑tuned models across various fields and adopt hybrid human‑AI workflows to address the limitations identified here.
Blogger's Review: The paper offers a systematic investigation of context‑window effects, reminding practitioners to apply critical judgment when leveraging LLMs for scholarly writing and to develop collaborative human‑machine review pipelines.