NeFut Logo NeFut
中 Admin Login

[CS.AI] When Should Forecasting Agents Reason? Behavioral Stress Tests for Reliability Routing

Published at: 2026-09-25 22:00 Last updated: 2026-09-28 00:49
#Machine Learning #LLM #Artificial Intelligence

Forecasting agents are increasingly blending language‑model reasoning, retrieval, ensembling, and calibration, yet it remains unclear when each behavior should be trusted. We investigate this on ForecastBench‑style binary forecasting tasks, treating the choice to retrieve, reason, defer to a market prior, or use a historical analog as an observable agent action rather than a hidden implementation detail.

Our central finding is that the optimal mechanism depends on the evidence source: structured analogs dominate for certain data‑generating processes, while market‑or‑crowd priors and conservative baselines perform better for others. Building on this, we introduce ReliabilityRoute, a structural intervention that steers agent behavior using reliability features such as historical coverage, market‑prior availability, source‑prior sharpness, evidence strength, evidence disagreement, and forecasting horizon.

We implement two routing rules. A fixed rule fitted on 2024 data matches a hand‑crafted taxonomy without hard‑coding source names; a walk‑forward self‑adjusting rule refits thresholds from previously resolved vintages and achieves the lowest mean Brier score among deterministic systems across 16 subsequent LLM vintages. Although gains are modest, historical and retrieval baselines remain highly competitive.

The main contribution is a behavioral stress test showing that “more reasoning” is not always better; forecasting agents should first estimate which evidence source deserves control, and routing policies themselves must adapt under auditable constraints. Reproducibility artifacts are available at https://github.com/louiswang524/forcastagent

Review

Original Source: https://arxiv.org/abs/2609.28475

[h] Back to Home