Evaluating LLM agents in enterprise record systems is hampered by the inability to use real production data and the lack of a ground‑truth substitute. To address this, we introduce the Era by Eon Benchmark, tailored for LLM agents that operate with enterprise tools. The benchmark is built around a fully simulated fictional company and includes product simulators, internal databases, benchmark questions, and automatically generated answer keys.
Company attributes—industry, size, business model, application portfolio, and a seed—define each entity. A seeded entity graph supplies shared data to simulators such as Salesforce, Zendesk, Slack, and Gong. A question‑conditioned generator then creates schemas and records for the internal databases using the same graph, ensuring a single consistent enterprise estate.
All expected answers are computed from the final records, guaranteeing exact grading. Internal databases are validated through design checks and answer‑key verification; the entity graph is vetted by a realism scorecard and an adversarial detector. Across 23 generated companies, the average realism score rose from 61.8 to 97.0, with zero synthetic records flagged.
In the simulator‑track comparison, nine models answered the same 33 questions three times each. Accuracy ranged from 42.4% to 76.8%, and after multiple‑comparison correction only three of the 36 pairwise differences remained statistically supported.
Review