We introduce GeoOutageBench, a benchmark for evaluating LLM‑driven geospatiotemporal KGQA in multimodal power‑outage and resilience analysis. Unlike existing Web‑KGQA benchmarks, it employs a spatiotemporal KG that fuses visual, textual, and structured data from outage logs, remote‑sensing imagery, weather observations, storm and power events, geographic entities, and domain ontologies. The benchmark defines a competency query taxonomy spanning four difficulty levels: spatiotemporal containment & proximity, spatiotemporal co‑occurrence analysis, multimodal evidence reasoning, and hypothetical evaluation. Across multimodal KG and query classes, GeoOutageBench supports three configurable tasks: (1) assessing LLMs’ ability to resolve ambiguous geospatiotemporal questions via NL‑to‑SPARQL translation; (2) query‑driven measurement of ontology utility; (3) accuracy of answer retrieval in multimodal KGQA. GeoOutageBench offers a design principle and experimental foundation for testing LLM‑KG systems in real‑world infrastructure resilience scenarios. The source code, data, results, and documentation are publicly available at https://github.com/UCF-SAGE/GeoOutageBench.
Review: GeoOutageBench unifies temporal, visual, and ontological information within a single KG, providing a systematic evaluation platform for LLM reasoning in complex infrastructure contexts.