NeFut Logo NeFut
中 Admin Login

[CS.AI] Risk-Aware Adaptive Evaluation: Finding High-Impact Failures Under Limited Budgets

Published at: 2026-10-01 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #optimization

Evaluating interactive agents is costly and their behavior stochastic, so reliability must be measured over repeated trials. Conventional benchmarks allocate the budget uniformly, e.g., a read‑only lookup receives as many trials as an irreversible payment action, wasting effort on scenarios that rarely fail. We cast evaluation as a sequential allocation problem with a fixed trial budget: given a set of scenarios with unknown failure behavior, decide which to run and whether to repeat them. To this end we propose a risk‑aware contextual Thompson Sampling policy that combines a pre‑execution scenario context vector, a fixed impact score, and the observed failure outcomes to update the posterior. Offline replay on 70 τ‑bench airline scenarios and 824 recorded trials shows that with the smallest budget—only 50 trials ($6\%$ of the corpus)—the policy recovers $86\%$ of the impact‑weighted failures an oracle could find, versus $25\%$ for uniform allocation. It discovers $3.5\times$ more impact‑weighted failures (215.4 vs. 62.2) with the same number of trials, yields $5\times$ the discovery per dollar, and reduces wasted budget on never‑failing scenarios from $34\%$ to $2.8\%$. A budget sweep reveals the advantage shrinks as the budget approaches the corpus size; paired significance tests indicate scenario context helps mainly at small budgets while posterior‑based exploration aids at moderate budgets. Thus risk‑aware adaptive allocation provides the greatest benefit precisely when evaluation budget is scarce.

Review: The study convincingly demonstrates that integrating contextual information with Bayesian exploration can dramatically improve the efficiency of failure discovery under tight resource constraints, offering a practical approach for safety‑critical system testing.

Original Source: https://arxiv.org/abs/2609.38914

[h] Back to Home