NeFut Logo NeFut
Admin Login

[CS.AI] Paper Pilot: A Human-in-the-Loop Expert System for Evidence-Traceable Scientific Manuscript Generation

Published at: 2026-09-02 22:00 Last updated: 2026-09-03 02:56
#AI #Machine Learning #LLM

Large language model (LLM) agents are increasingly embedded in scientific workflows for literature analysis, drafting, and review. Existing systems enable autonomous discovery and manuscript generation but leave a governance gap: ideas, methods, results, and claims can propagate through AI‑assisted pipelines without mandatory human approval or artifact‑level traceability.\ \ This paper introduces Paper Pilot, a human‑in‑the‑loop expert system for evidence‑traceable manuscript generation in applied sciences. It adapts the Collaborative Agent Reasoning Engineering (CARE) methodology to manuscript development, featuring manuscript‑owner approval gates, explicit no‑pass criteria, claim classification, audit logging, advisory LLM review, and evidence‑locked revision control.\ \ The framework defines eight approval gates across the idea‑to‑claim pipeline and distinguishes literature‑grounded from artifact‑grounded claims, requiring reported numbers and interpretations to remain traceable to approved evidence. The system prompt is openly released for deployment in ChatGPT, Gemini, Claude, or institutional LLM environments.\ \ For an initial empirical validation, we evaluate the citation‑grounding layer with a controlled, mechanically scored benchmark using two commercial LLMs and real arXiv papers, without an LLM judge. Under coverage pressure, ungated drafters fabricated up to 25% of citations and never flagged an evidence gap; the same models under Paper Pilot’s evidence‑locked rules produced zero fabricated citations and surfaced planted gaps as explicit placeholders. Preliminary results for result grounding, revision, and adversarial robustness point the same way; full evaluation is left to future work.\ \ Paper Pilot positions LLM‑assisted writing as a controlled human‑AI decision‑support process rather than a fully autonomous authorship pipeline.\ \ Review

Original Source: https://arxiv.org/abs/2608.28596

[h] Back to Home