Building automation systems are increasingly modeled as semantic knowledge graphs (KGs) using ontologies such as Brick and ASHRAE 223P, creating a machine‑readable substrate for AI applications. Translating natural‑language questions into SPARQL (text‑to‑SPARQL) would let operators query these graphs via language agents, yet progress is hampered by the lack of large NL/SPARQL benchmarks.
This paper introduces Build2SPARQL, a large‑scale benchmark generated by a KG‑grounded pipeline. The pipeline first produces and validates SPARQL queries entirely with graph‑traversal code, guaranteeing query correctness independent of downstream models; then only large language models (LLMs) generate the corresponding natural‑language questions. Six query‑pattern families are mined: linear chains, branching, UNION, aggregation, OPTIONAL, and attribute‑filtered, each expressed across five vocabulary registers.
Applied to 201 building KGs (180 Brick, 21 ASHRAE 223P), the pipeline yields 6,136 executable SPARQL queries and 30,680 questions. A two‑rater human validation on 300 sampled questions reports 98.8% semantic fidelity, 98.8% naturalness, and 84.0% operational plausibility. Retrieval‑augmented evaluation on three open‑weight models raises exact‑match accuracy from 0.2%‑20% (zero‑shot) to 56%‑65% (three‑shot retrieved).
Review: Build2SPARQL decouples query generation from natural‑language generation, delivering a sizable, high‑quality benchmark that offers a solid experimental foundation for text‑to‑SPARQL research and demonstrates the promise of retrieval augmentation for boosting zero‑shot model performance.