State‑of‑the‑art Natural Language to SQL (NL2SQL) models now report execution accuracies above 89% on public benchmarks such as Spider and BIRD. These benchmarks, however, rely on simplified academic schemas and open‑source SQL dialects that do not capture the intricacies of enterprise databases. To address this gap we introduce ESQ‑Bench, an Oracle‑first NL2SQL benchmark that provides systematic complexity tiers (Tier‑1, Tier‑2, Tier‑3) and a silent‑divergence evaluation.
We built and released six populated schemas containing a total of 465 tables and 164 682 rows, with identical seed data on Oracle, PostgreSQL, MySQL, and SQL Server. The benchmark comes with a four‑metric evaluation harness:
- EM (Exact Match)
- EX (Execution Accuracy)
- SR (Survivor Rate)
- SD (Silent Divergence)
A total of 550 gold‑validated question‑SQL pairs are provided (Tier‑1: 95, Tier‑2: 228, Tier‑3: 227).
When using GPT‑4o with schema‑linked prompting, execution match degrades monotonically across tiers: 79.8% (Tier‑1), 60.3% (Tier‑2), and 57.2% (Tier‑3) as of June 2026, compared with 75.6%, 80.4% and 95.8% on an earlier 142‑question pilot slice. EM stays below 7% for all tiers. Among queries that pass EX, silent‑divergence reaches 73%–99%, indicating that models often return semantically wrong results rather than outright failures. Failure analysis shows that wrong‑result semantics dominate at higher tiers.
Claude Sonnet 4.6 with the same schema‑linked prompts achieves 87.4%, 74.9% and 68.7% EX, surpassing GPT‑4o on every tier. Zero‑shot GPT‑4o yields EX of 78.7%, 73.5% and 77.8% on executed queries, flipping the schema‑linked advantage at Tiers 2‑3 due to lower execution rates and survivor bias in the zero‑shot versus schema‑linked comparison.
Locally‑run Llama 3.2 with schema‑linked prompts reaches only 13.3% bank‑wide EX (73 out of 550), underscoring the performance gap between closed‑API models and open‑weight baselines on enterprise Oracle schemas.
Blogger's Review: ESQ‑Bench shines a light on the blind spots of current NL2SQL evaluation practices in real‑world enterprise settings, especially regarding dialect diversity and schema complexity. The steep performance drop on higher‑tier workloads warns that chasing high scores on public benchmarks is no longer sufficient; future work must prioritize robust execution semantics and cross‑dialect generalization.