Time series forecasting is increasingly used to guide decisions in transportation, energy, finance, healthcare and infrastructure, yet current evaluation remains overly narrow. Standard benchmarks reward lower held‑out error, while robustness studies usually reduce failure to Gaussian noise, random masking or bounded adversarial perturbations. This hides the real failure modes of deployed systems. Input‑side anomalies are not merely noisy data; they often represent structured events that change temporal dynamics, break cross‑variable dependencies, trigger regime shifts, or propagate from faulty sensors to downstream decisions. Such semantic, causal and system‑level failures cannot be captured by i.i.d. perturbations alone. The rise of foundation models for time series forecasting makes this gap more urgent, as unauditable pre‑training corpora render held‑out generalization increasingly unreliable. We therefore advocate scenario‑grounded stress testing. Each test instance should include historical inputs, future targets, a semantic scenario, an explicit failure operator and a measurable difficulty level. This shift makes evaluation interpretable, attributable and deployment‑relevant, enabling the community to ask not only which model is accurate, but under what conditions it fails and why.
Review