Banks require conversational agents that can answer product queries, help with account‑related requests, and operate safely under strict operational and regulatory constraints. General‑purpose language models often fail when tasks need grounded information, correct tool usage, or careful handling of bank‑specific sensitive situations. FiMI Banking is a controlled Indian retail‑banking environment built from vetted banking documents, structured ground truth, synthetic customer profiles, and banking tools. We evaluate two post‑training approaches: preference optimization to improve response‑level safety, and reinforcement learning with verifiable rewards for multi‑turn tool‑use tasks. Preference optimization raises out‑of‑scope refusal from 52% to 80%. Reinforcement learning boosts edge‑case performance from 0.509 to 0.718 and order‑sensitive task performance from 0.590 to 0.679 while reducing generated tokens by 29%. The results demonstrate that preference optimization and verifiable‑reward reinforcement learning address complementary requirements for reliable banking agents.
Review