Building automation systems generate massive sensor streams, yet their operational insight is limited by heterogeneous point naming, missing metadata, and fragmented documentation. This systematic review examines 66 peer‑reviewed studies on large language models (LLMs) for HVAC operations published between 2023 and March 2026, classifying each work into five application families and three LLM method families while assessing evidence realism, deployment readiness, and the responsibility boundary between the LLM and physical HVAC decisions. The corpus is heavily weighted toward building energy modelling (BEM) – 32 papers – while load‑forecasting research remains too sparse for sub‑field conclusions. Only four studies provide pilot‑level evidence, and none reports sustained operational deployment; no paper is deemed ready‑now for industry, three are near‑term, and the remaining 63 are research‑only. Nevertheless, bounded human‑in‑the‑loop uses show near‑term promise, such as point‑name normalisation, document‑grounded operator support, BEM workflow assistance, and advisory interfaces that augment physics‑based controllers. Conventional machine learning, model predictive control (MPC), reinforcement learning (RL) and ontology‑based tools continue to dominate high‑frequency control, short‑horizon numerical forecasting, and well‑defined ontology mapping, while autonomous agentic operation and unvalidated occupant proxies remain at the research stage. Current evidence therefore positions LLMs primarily as semantic and workflow layers rather than autonomous HVAC controllers. Future work should prioritise field‑validated benchmarks, orchestration evaluation under operational constraints, and LLM‑MPC/RL architectures with bounded latency and verifiable safety properties.
Review