Scientific coding agents generate interdependent code, results, figures, and claims, yet evaluating only the final outputs does not reveal whether the conclusions are scientifically supported. We define evidence‑grounded multimodal scientific analysis, requiring agents to produce executable analyses whose claims are directly backed by results and visualizations from the same run. To this end we introduce the SciRIGOR framework and benchmark, comprising 100 cases drawn from six domains and seventeen sub‑fields. The framework reconstructs typed evidence graphs, separates artifact fidelity from relational validity, and scores complete claim‑support paths while pinpointing the earliest unsupported relation. Source‑grounded alternative paths allow scientifically equivalent analyses and visualizations. We evaluate eleven agent/model configurations. Across full‑benchmark runs, claim agreement with faithful and unfaithful results is nearly identical (91.8% vs 91.0%). No system exceeds a 62.6% soft evidence‑chain score or an 18.0% strict whole‑chain success rate. These findings demonstrate that internal coherence alone does not guarantee scientific correctness; evaluation must verify support along the entire data‑to‑claim chain.
Review