Deep research agents have become capable of web search, tool use, multimodal evidence analysis and information synthesis. Existing benchmarks, however, mainly test medium‑horizon exploration and rarely assess whether agents can sustain long, dependency‑heavy research processes. We introduce Mr.LHDR (Multimodal real‑world Long‑Horizon Deep Research), a benchmark built from hidden node‑relation graphs covering eight categories. Each question requires on average 12.1 necessary intermediate conclusions with a mean dependency depth of 10.4 before arriving at a short, unique and verifiable answer. At least one non‑text element—such as an image, map, PDF, logo, chart, table or video frame—is included and changes the reasoning state. Mr.LHDR evaluates both final answers and the correctness of intermediate conclusions under annotated dependencies, using Overall Accuracy (OA), Strict Accuracy (SA), Checklist Score (CS) and Dependency‑Aware Checklist Score (DACS). Results show that even the strongest system reaches only 43.1% OA and 34.3% SA, indicating that final‑answer accuracy greatly overestimates complete research success. Removing images reduces DACS by 12.6 points, highlighting the importance of multimodal evidence, while SA consistently declines as reasoning chains grow longer. These findings reveal that sustained, dependency‑consistent evidence integration, rather than isolated fact retrieval, is the key bottleneck for current deep research agents.
Review