In finance, interpreting machine‑learning predictions is critical, yet the numerical outputs of explainable AI are often hard for non‑experts to grasp. Large language models (LLMs) can translate these numbers into natural language, but they may err when inferring numeric changes and feature relationships. This paper proposes an LLM narrative framework for cross‑sectional stock return prediction that combines temporal Shapley additive explanations (SHAP) evidence with historical regime analogs.
Temporal evidence tracks the normalized global SHAP importance of an XGBoost model over six months, yielding a sequence $\{\text{SHAP}_t\}$.
Historical analogs are past periods whose SHAP change patterns resemble the current one; their model performance and subsequent market returns are presented as comparative context.
Within this framework we conduct a progressive reasoning externalization study, sequentially feeding the model raw SHAP sequences, deterministic temporal descriptors, and feature‑relation statements. Each generated claim is verified against provenance‑linked evidence.
Experiments with the Qwen3 family show that externalizing numeric and relational reasoning markedly improves evidence faithfulness as well as temporal and relational accuracy. For Qwen3-32B-Instruct, faithfulness rose from 0.696 to 0.996. Although historical analogs did not boost structured automatic faithfulness, they received higher human‑rated usefulness scores.
Review: Explicitly exposing SHAP‑based evidence and enriching it with analogous historical contexts substantially enhances the trustworthiness and interpretive value of LLM‑generated financial narratives.