NeFut Logo NeFut
Admin Login

[CS.AI] ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies

Published at: 2026-09-07 22:00 Last updated: 2026-09-08 00:37
#AI #Machine Learning #LLM

Large language model (LLM) agents are increasingly proposed for enterprise workflows, yet existing evaluations rarely test whether business‑decision conclusions transfer across competitive market ecologies. We introduce ERPBench, an execution‑instrumented benchmark for enterprise decision agents in a six‑round Enterprise Resource Planning (ERP) simulation that couples pricing, production, procurement, inventory, finance, and shared‑market competition.

ERPBench evaluates the same 100 fixed problems in two matched competitive market ecologies:

Across six model families, this yields 1,200 model‑level trajectories spanning 7,200 decision rounds. Under the observed service configuration, the leading model differs between ecologies: DeepSeek leads in Solo ($252.29M$ mean valuation; mean rank 1.67), whereas Gemini leads in Arena ($263.95M$; rank 1.76). The two ecologies identify the same task‑level winner on only 21 of 100 problems, and Gemini's bottom‑rank rate falls from 22% in Solo to 0% in Arena.

ERPBench supports paired evaluation of whether enterprise‑agent rankings transfer across competitive market ecologies, supplemented by aggregate execution‑intervention analysis. Code and benchmark resources are available at https://github.com/GAIR-NLP/erp-bench.

Review

Original Source: https://arxiv.org/abs/2609.04667

[h] Back to Home