Financial large language models are increasingly used to summarize reports and disclosures, where numerical hallucination poses real risks. Prior work often blames insufficient numerical reasoning, yet systematic tests under controlled fine‑tuning are lacking. This study compares three variants under a cost‑effective setup: a base instruction‑tuned model, a domain‑adapted model (FT‑A), and a numeracy‑enhanced domain model (FT‑A+B+C). We define a three‑level detectability taxonomy: overt hallucination (fabricated currency amounts), covert‑explicit hallucination (numbers that follow professional conventions), and covert‑implicit hallucination (ungrounded quantitative claims). Results show domain fine‑tuning drastically reduces numerical restraint: the Base model’s hallucination rate is about $5.4\%$, near zero; FT‑A exhibits $82.5\%$ overt hallucination, and FT‑A+B+C reaches $98\%$. Contrary to intuition, numeracy supervision amplifies hallucination across all levels. Analysis identifies template injection—injecting memorized canonical values regardless of input—as the primary mechanism in fine‑tuned models. The findings indicate that numerical hallucination in financial summarization stems from the degradation of numerical restraint during domain adaptation, not from a lack of reasoning ability. We recommend evaluation protocols that cover all detectability levels and deployment practices that incorporate grounding‑aware generation or abstention mechanisms.
Review