NeFut Logo NeFut
中 Admin Login

[CS.AI] The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading?

Published at: 2026-09-28 22:00 Last updated: 2026-09-30 01:41
#Machine Learning #LLM #Artificial Intelligence

We conducted a controlled study on three major LLM families—DeepSeek, GPT, and Gemini—to examine whether adding test‑time reasoning improves economic outcomes in equity trading. The experiment kept the information available at each formation date, prompts, output format, and portfolio construction fixed, varying only the amount of internal reasoning from none to maximal. Evaluation covered a full year of U.S. equities under three input regimes: numeric features, identifiable news, and masked news, generating over 800,000 asset predictions and repeating model outputs to assess stability. Across all families, extra reasoning did not yield a reliable increase in net portfolio returns; DeepSeek even showed a non‑monotonic relationship where more reasoning sometimes reduced performance. Repeated generations produced unstable treatment effects and portfolio selections despite similar overall scores. These findings indicate that while reasoning can alter financial decisions, it does not consistently enhance their economic value, underscoring the need for task‑specific validation before deployment.

Review

Original Source: https://arxiv.org/abs/2609.30705

[h] Back to Home