NeFut Logo NeFut
中 Admin Login

[CS.AI] Same Text, Different Numbers: The Divergence of LLM-Based Measures

Published at: 2026-09-28 22:00 Last updated: 2026-09-30 01:41
#AI #Machine Learning #LLM

Researchers are increasingly employing generative large language models (LLMs) to turn corporate text into quantitative variables. This study evaluates how invariant such LLM‑based textual measures are to the choice of model, using thirteen constructs such as sentiment, management clarity, uncertainty, answer specificity, and climate and political risk. Seven LLMs from different providers scored earnings‑call transcripts of S&P 500 firms. Cross‑model rank correlations averaged only 0.52, and the portion of variation common across providers accounted for just 34% of total score variance. Disagreement between models did not forecast subsequent analyst or market disagreement, indicating that most divergence stems from model‑specific factors rather than inherent ambiguity in the disclosures. Model choice markedly altered downstream inference, with coefficient magnitudes, signs, and statistical significance shifting across models. Averaging across providers stabilized transcript rankings for most constructs, yet absolute score levels remained sensitive to which models were included. Consequently, LLM‑generated variables should be treated as model‑contingent measurements and validated across providers.

Review

Original Source: https://arxiv.org/abs/2609.31013

[h] Back to Home