This work investigates silent failures that occur when an agent invokes a tool successfully on the surface, yet the returned information is incomplete or missing and no notification is issued to the user or the agent. Within the ToolUniverse sandbox we selected fifteen scientific tools together with their API specifications and wrapper implementations, and built an audit pipeline to detect such failures.
The audit is organized around seven failure loci: request construction, parameter transmission, API call, wrapper handling, result parsing, post‑processing, and output rendering. By combining LLM‑guided candidate discovery with automated testing and manual validation we identified ninety‑one silent failures. The most frequent issues were missing fields or data and inconsistencies in search, filtering, or ranking criteria. Fifty‑one failures originated in the API layer and twenty‑five in the wrapper layer, indicating that upstream problems can propagate downstream and produce apparently valid scientific outputs.
To address this, we introduce the notion of contextual reliability and propose a set of mechanisms: enforce integrity checks and explicit error flags at the API and wrapper levels; apply schema validation to critical response fields; deploy monitoring dashboards to trace failure propagation; and embed failure‑aware fallback strategies in the agent’s decision‑making process. These measures aim to surface silent failures promptly, alert users, and mitigate downstream impact.
Review