NeFut Logo NeFut
Admin Login

[CS.AI] Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents

Published at: 2026-09-12 22:00 Last updated: 2026-09-15 01:15
#AI #Machine Learning #LLM

Autonomous research agents are expected to retrieve literature, analyse experimental evidence, and generate hypotheses, which demands multi‑step, evidence‑grounded reasoning that progressively acquires, integrates, and verifies evidence before reaching a conclusion. Existing multimodal benchmarks mainly assess final‑answer accuracy and do not verify whether predictions are supported by traceable scientific evidence. To address this, we introduce Sci‑MMR, a benchmark built on structured argument graphs that link scientific claims, citation‑grounded knowledge, visual evidence, and their supporting regions. The dataset comprises 235 multi‑hop reasoning tasks across four scientific disciplines, with an average of nine figure panels per task. Evaluating eight state‑of‑the‑art multimodal models, we find that answer accuracy consistently exceeds the complete‑evidence recovery rate by more than 20%, indicating that answer‑only evaluation substantially overestimates models’ evidence‑based reasoning abilities.

Controlled interventions reveal two fundamental bottlenecks. First, evidence acquisition: models struggle to extract complete structured evidence from scientific figures, accounting for 57.2% of failures. Simple cropping tools yield only a modest +4.5‑point gain, whereas providing gold evidence improves accuracy by up to +37.0 points, highlighting the difficulty of assembling multi‑region evidence. Second, evidence integration: models fail to translate available evidence into correct conclusions for 31.8% of cases, and even with gold evidence the strongest model reaches only 69.1% accuracy on the hardest tasks. These findings demonstrate that current answer‑centric benchmarks dramatically overestimate the evidence‑grounded reasoning capabilities of multimodal research agents. Review

Original Source: https://arxiv.org/abs/2609.11243

[h] Back to Home