LLM agents rely on retrieval tools to access external knowledge, yet visual agentic search is severely limited by standard single‑step retrievers. In existing pipelines the agent must issue a text query for every intermediate step, and performance collapses when visual cues are hard to describe or the retriever fails to surface necessary intermediate evidence among its top results. We hypothesize that offloading multi‑step navigation across the entire embedding space to the retrieval tool can remove this bottleneck. To evaluate systematically we introduce VHOP—a flexible data‑generation framework and benchmark that defines five core difficulty levels, testing both visual matching and search planning. Building on VHOP we develop VHOP‑Router, an end‑to‑end training pipeline that combines supervised fine‑tuning, online imitation learning, and reinforcement learning to turn a standard embedding model into an autoregressive multi‑step retriever. Operating directly in the visual latent space, VHOP‑Router retrieves linked image chains with a single tool call, eliminating the need for the agent to formulate intermediate text queries. Experiments show retrieval success rising from under 5% to 76.3%, task success improving by 52.7%, and average token length dropping from 1886 to 728 (a 61% reduction), whereas merely upgrading the agent yields only a 3.7% gain. Compared with a strong baseline that retrieves the top 50 results per step, VHOP‑Router maintains superior performance while cutting in‑context images by $23\times$ and reducing cumulative API payload by $35\times$. The models also generalize robustly to unseen difficulty levels and realistic test sets. In summary, VHOP and VHOP‑Router provide an efficient and effective solution for visual agentic search that leaves native LLM capabilities untouched.
Review: By performing multi‑step retrieval directly in visual latent space, this approach dramatically boosts efficiency and success rates, offering a promising path forward for visual‑centric agentic search.