Large language model (LLM) agents are increasingly embedded in scientific workflows for literature analysis, drafting, and review. Existing systems enable autonomous discovery and manuscript generation but leave a governance gap: ideas, methods, results, and claims can propagate through AI‑assisted pipelines without mandatory human approval or artifact‑level traceability.
This paper introduces Paper Pilot, a human‑in‑the‑loop expert system for evidence‑traceable manuscript generation in applied sciences. The framework adapts the Collaborative Agent Reasoning Engineering (CARE) methodology to manuscript development by adding manuscript‑owner approval gates, explicit no‑pass criteria, claim classification, audit logging, advisory LLM review, and evidence‑locked revision control.
Paper Pilot defines eight approval gates along the idea‑to‑claim pipeline, distinguishes literature‑grounded from artifact‑grounded claims, and requires reported numbers and interpretations to remain traceable to approved evidence. The system prompt is openly released for deployment in ChatGPT, Gemini, Claude, or institutional LLM environments.
For an initial empirical validation, the citation‑grounding layer was evaluated with a controlled benchmark: two commercial LLMs, real arXiv papers, and a mechanically scored assessment without an LLM judge. Under ungated drafting, models fabricated up to 25% of citations and never flagged evidence gaps; under Paper Pilot’s evidence‑locked rules, fabricated citations dropped to zero and all planted gaps appeared as explicit placeholders. Preliminary results for result grounding, revision handling, and adversarial robustness show the same trend, with full evaluation left to future work.
Paper Pilot positions LLM‑assisted writing as a controlled human‑AI decision‑support process rather than a fully autonomous authorship pipeline.
Blogger's Review: Paper Pilot offers a technically sound governance layer through explicit approval gates and audit trails, addressing the credibility crisis of AI‑generated research content. If its scalability and usability are demonstrated across larger real‑world projects, it could become a standard tool in scholarly publishing.