NeFut Logo NeFut
中 Admin Login

[CS.AI] IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking

Published at: 2026-09-25 22:00 Last updated: 2026-09-28 00:49
#AI #Machine Learning #LLM

Banking assistants must leverage account‑specific data to answer queries and often act through tools. Evaluating only the final answer misses critical errors such as asking for information that is already known, relying on stale context, selecting the wrong account, or writing an invalid value after stating the correct one. To address this we introduce IndicBankBench, a benchmark of 799 cases covering five operational domains of Indian retail banking, a capability/refusal domain, and twenty primary axes. Each case is judged at four stages: safety, tool usage, response adequacy, and advisory quality. Tool usage and most safety checks are deterministic; a narrow resolver handles only ambiguous confirmation‑before‑write situations, while a separate LLM judge assesses semantic adequacy. Every case is run three times and evaluated with a strict pass³ metric, requiring success on all trials. Across eleven models, strict reliability ranges from 43.7% to 58.2%, whereas at‑least‑once success spans 60% to 74%. This gap shows that a single successful run can overstate dependable banking behavior. Case‑level diagnostics also separate systems that ask unnecessary questions from those that act but fail to reconcile customer context or fully resolve the request. The cases, mock environment, and evaluation harness are released publicly.

Review

Original Source: https://arxiv.org/abs/2609.29167

[h] Back to Home