Recent advances in large language model (LLM) agents have created a new asset‑pricing paradigm called Agentic Empirical Asset Pricing (AEAP), where a system autonomously carries out the entire scientific discovery loop. We first formalize AEAP and break it down into core components: data acquisition, feature construction, factor validation, and trade‑signal generation. Existing evaluation practices backtest only the resulting factors or trades, overlooking the autonomous discovery engine that produced them. To address this gap we propose a reference architecture that includes: (1) a standardized factor‑discovery pipeline; (2) rigorous statistical tests for the discovered factors; (3) an out‑of‑sample backtesting protocol for the discovery system itself. Using this architecture we implement SEADS and benchmark it against five re‑implemented baselines on two US equity panels. The results show that no single metric consistently ranks the systems, motivating evaluation across multiple axes simultaneously. A subsequent rolling re‑execution experiment asks whether the discovery process, rather than a static output, is reliable. We also report negative findings and limitations that expose further evaluation pitfalls for future AEAP systems.
Review