AI‑generated content, often dubbed AI slop, is becoming ubiquitous in academia. Unlike generic text, AI‑written scientific papers usually appear plausible at the surface but collapse in the logical reasoning that links sections, potentially misleading readers. We term this failure scientific slop and define six metrics across Structure, Argument, and Artifacts to quantify it.
Using these metrics we built SciSlopBench, a dataset of 390 AI‑generated papers—mostly computer science but also spanning life, social, and natural sciences—each paired with a human‑written counterpart matched by research problem and contribution type. The metrics identify the AI paper in each pair with 85.9% accuracy, compared to 68.7% for the existing Binoculars detector. Higher slop correlates with lower ICLR scores and distinguishes rejected from accepted papers above chance for every year from 2017 to 2025.
Directly optimizing the metrics does not effectively reduce slop. Therefore we propose SciSlopHarness, a harness‑level framework that guides a fixed LLM to revise only those parts where experimental records support a change. Standard prose revisions leave residual slop, and slop‑aware prompting triggers reward hacking. SciSlopHarness cuts the AI‑human gap by 63% over the strongest revision baseline without requiring human reference targets.
In sum, AI‑generated scientific papers leave detectable traces in their global reasoning, and responsible mitigation demands strict evidentiary grounding rather than mere prose polishing.
Review