Deep Research agents synthesize retrieved evidence into cited reports, yet even heavily cited reports can reach misleading conclusions. Conventional citation‑correctness checks only verify whether each cited source supports its claim, without assessing whether the adaptive search exposed a representative view of all documents available for evaluation—the candidate pool. Early search results steer later queries, document selections, and stopping decisions, so the set of documents an agent reads becomes a biased sample, a factor rarely accounted for in existing evaluations.
We formalize this as adaptive evidence sampling and introduce Causal Evidence Selection Correction (CESS). CESS predicts the evidence direction (positive or negative) for each candidate document and corrects the candidate‑pool average using the logged probabilities of selecting each document at each search round. The procedure is:
- Compute an evidence score $e_i$ for each document (positive for supporting, negative for opposing).
- Record the selection probability $p_{i,t}$ of document $i$ at round $t$.
- Estimate the pool‑wide evidence direction via a weighted average $$\hat{E}=\frac{\sum_{i,t} p_{i,t}\, e_i}{\sum_{i,t} p_{i,t}}.$$
- Apply shrinkage when the search is short, pulling the estimate toward the overall mean to stabilize it.
- Replace point estimates with confidence intervals for documents that were never sampled, reflecting uncertainty.
We also prove that estimating the average evidence direction of a common pool differs from measuring how a change in search policy alters the evidence read; the latter requires an intervention experiment that forces different retrieval policies on the same pool and compares outcomes.
On the MS2 systematic‑review benchmark, CESS reduces the mean absolute error against the candidate‑pool average by $9.2\%$ and cuts the estimate’s change under opposing document rankings by $39.4\%$ compared to simply averaging evidence scores of read documents. Across trajectories from a public Open Deep Research agent, the reductions reach $60.1\%$ and $87.2\%$, respectively. A further 4,800 paired‑intervention trajectories confirm that correcting a pool estimate and measuring a policy effect are distinct tasks.
Thus, CESS audits whether the evidence direction underlying a report truly reflects the documents available for evaluation, while a separate intervention analysis quantifies the impact of search decisions.
Review