In scientific literature retrieval, overall answer accuracy alone cannot reveal whether a search agent actually encountered a target paper, attempted to inspect it, or returned an accepted answer after inspection. To address this, we introduce decision checkpoints that log observations and tool actions during inference without requiring benchmark labels; later we merge target identities with evaluator labels and assign outcome categories based on the recorded events. Experiments were conducted in a fixed, target‑rich environment on 540 answerable AutoResearchBench deep questions across five conditions. Keyword search achieved 24.6% accuracy, compared with 17.8% for raw search. Under the keyword condition, there were fewer incorrect answers with neither exposure nor inspection, but more errors after the target was exposed yet left uninspected. Compared with keyword search, the read‑first condition recorded 27.4% more evidence‑search calls. Target inspection attempts occurred on 199 questions under read‑first and 191 under keyword search; both conditions reached 24.6% accuracy. The checkpoint protocol makes these question‑level differences explicit, separating target exposure and inspection from aggregate accuracy and total tool usage.
Review: The checkpoint approach offers fine‑grained behavioral tracing, providing a clearer lens for assessing the true capabilities of scientific search agents.