NeFut Logo NeFut
Admin Login

[CS.AI] Unveiling Silent Failures in Multimodal Agentic Search

Published at: 2026-07-23 22:00 Last updated: 2026-07-26 07:44
#algorithm #AI #optimization

Abstract

Multimodal agentic search systems increasingly rely on external tools to answer knowledge-intensive visual questions. However, existing evaluations mainly focus on final-answer accuracy and may miss failures in the search trajectory.

In this work, we study such hidden reliability issues as silent failures. We introduce a six-category taxonomy covering modality shortcuts, phantom grounding, wrong-evidence-right-answer cases, over-retrieval laundering, cross-modal contradiction, and provenance hallucination.

Based on this taxonomy, we build a trajectory-level diagnostic pipeline that evaluates both answer correctness and evidence-grounding quality under a unified ReAct-style scaffold.

Experiments on MMSearch-Plus trajectories across four frontier multimodal models show that surface accuracy consistently overestimates true trajectory-level correctness.

We further use cross-judge validation, blank-image stress tests, and tool ablations to show that silent failures are capability-dependent and often shift rather than disappear.

Project Homepage

Blogger's Review: This paper not only uncovers potential silent failures in multimodal search systems but also provides a systematic classification and evaluation method. This offers significant insights for future research, especially in enhancing the reliability and accuracy of models.

Original Source: https://arxiv.org/abs/2607.19793

[h] Back to Home