NeFut Logo NeFut
中 Admin Login

[CS.AI] FinInteract: Benchmarking Clarification and Intent Integration in Ambiguous Financial Question Answering

Published at: 2026-09-24 22:00 Last updated: 2026-09-28 00:49
#AI #Machine Learning #LLM

Large language model (LLM) agents are increasingly answering financial queries by searching regulatory filings. Such queries often hide ambiguity; for example, Meta Platforms' "operating income" is $46.75 B in the consolidated statement but $62.87 B for the Family of Apps segment, and both figures are verifiable. A capable agent should detect this ambiguity and ask for clarification rather than committing to a plausible but unintended reading. Existing financial benchmarks provide a single gold answer per question, making it impossible to distinguish agents that resolve ambiguity from those that merely guess the common interpretation—a problem we call the single‑gold illusion. FinInteract is a bilingual (English/Chinese) benchmark containing 173 instances, each pairing a question with a default and an intended interpretation across a five‑category ambiguity taxonomy. The benchmark grades whether an agent first elicits the correct clarification and then integrates it. Re‑grading identical outputs against the default instead of the intended reading inflates GPT‑4o's accuracy by 3.1×, confirming the illusion. Further analysis shows that when the intended interpretation is supplied, models achieve over 90% accuracy, but when they must elicit it themselves, the best accuracy drops to 28.9%. Performance is relatively balanced for entity‑scope and metric‑definition ambiguities, but weaker for exploratory cases. Conditioning on the ambiguity category improves resolution both at inference time and during training.

Review

Original Source: https://arxiv.org/abs/2609.24002

[h] Back to Home