NeFut Logo NeFut
中 Admin Login

[CS.AI] Propose, Don't Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors

Published at: 2026-09-24 22:00 Last updated: 2026-09-28 00:49
#AI #Machine Learning #LLM

Language‑model agents can now run the entire quantitative factor pipeline: propose factors, back‑test them, select survivors and retire the rest. We ask which steps the agent should retain and answer with a self‑evolution rule: the agent proposes, while a frozen statistical referee—immutable to the agent—judges. The referee scores each submission only after market outcomes are revealed, using a betting scheme, so its false‑discovery guarantee holds at any stopping time for any proposal policy. We cross‑tested three proposers (a script, a bandit algorithm, and a language model) against this frozen referee and three deliberately leaky referees in a synthetic world with planted true factors, a probe‑authoring environment, and a ten‑year walk‑forward on the CSI 500. Findings: the referee determines the number of false admissions—the frozen referee admits 5‑11× fewer sub‑threshold factors than leaky referees under the scripted proposer, and no proposer can close that gap. The proposer determines yield—the language model outperforms the script, matches the bandit, and adds the bandit’s missing capability of writing its own diagnostic probes. The certificate’s price is time: a certified true factor waits about 500 trading days, so the certified portfolio’s Sharpe ratio trails an ungated one.

Review

Original Source: https://arxiv.org/abs/2609.27051

[h] Back to Home