Trajectory evaluation is crucial for improving the reliability of LLM‑based agents, yet repeatedly running it in production is costly. Modern agents emit long traces that include tool calls, observations, retries, and external outputs, but not all raw tokens contribute equally to diagnosis. We introduce LiteTrajEval, a lightweight architecture that performs budget‑bounded trajectory evaluation.
LiteTrajEval first derives compact domain‑specific rule profiles offline. At runtime it preprocesses each trace, marks heuristic failure signals, serializes the trace within a fixed global budget, and invokes a single rubric‑guided LLM judge to produce a structured diagnostic report.
Evaluations on public Magnetic‑One‑style and $\tau$‑bench‑style datasets show that LiteTrajEval improves failure‑localization alignment with human annotations by roughly 20%–35% on Magnetic‑One and up to 23% on $\tau$‑retail, while cutting cost by about 6× and evaluation time by more than 8× compared with AgentRx. The solution has been deployed in our enterprise agentic platform.
Review: LiteTrajEval balances efficiency and interpretability by combining offline rule compression with online budget control, offering a practical cost‑performance trade‑off for large‑scale LLM agent deployments.