This work examines web‑search behavior of conversational LLM agents across four major platforms (ChatGPT, Claude, Grok, DeepSeek). By combining real‑world user interactions (in‑vivo) with controlled experiments using the same models via their APIs (in‑vitro), we evaluate the quality of search decisions, query‑formulation strategies, domain preferences in retrieved results, and how agents transform those results into grounded replies.
We find substantial variation in search invocation frequency among platforms and models, and that more frequent searching does not automatically improve response quality. Agents employ complex querying tactics such as multi‑turn refinement, combination of keywords, and result constraints, while each platform’s built‑in search engine tends to return results from its preferred domains.
Most replies are grounded in retrieved documents, yet some statements lack citations, raising concerns about attribution and reliability. The findings offer guidance for designing future AI agents and web‑search tools optimized for conversational retrieval.
Review