Source-linked AI summary
Large Language Models for HVAC Operations in Building Energy Systems: A Critical Review of Methods, Applications, and Deployment Readiness
Alexander Neubauer, Tianzhen Hong, Han Li, Mengbo Yu, Amin Darbandi, Yannick Fürst, Martin Kriegel
TL;DR
HVAC systems contain abundant data but fragmented naming, metadata, and documentation limit operational use, creating a need to identify defensible LLM roles. This review codes 66 peer-reviewed studies across application, method, evidence, readiness, and responsibility dimensions, finding no ready-now role and supporting LLMs mainly as bounded semantic and workflow layers.
Problem
Fragmented point naming, missing metadata, and dispersed documentation limit the operational use of building automation data and complicate deployment guidance for LLM-based HVAC roles.
Method
The review codes 66 peer-reviewed LLM-for-HVAC studies across application and method categories, evidence realism, deployment readiness, and responsibility boundaries.
Results
No study is classified as ready-now; 3 are near-term and 63 are research-only, while only four pilot studies and no operational deployments are identified.
Takeaways & Limitations
Current evidence most defensibly supports documentation interfaces, integration accelerators, semantic support layers, and advisory wrappers around established tools and controllers.
Takeaways & Limitations
The corpus is limited to Scopus-indexed peer-reviewed evidence, excluding grey literature and some conference items, so readiness findings do not bound the entire field.
Abstract
from arXiv · showhide
Building automation systems generate rich sensor data yet remain insight-poor because heterogeneous point naming, missing metadata, and fragmented documentation obstruct their operational use. This systematic review analyses and codes 66 peer-reviewed studies on large language models (LLMs) for HVAC operations published between 2023 and March 2026. Each study is classified across five application families and three LLM method families and assessed for evidence realism, deployment readiness, and the responsibility boundary between the LLM and physical HVAC decisions. The corpus is concentrated in building energy modelling (BEM, 32 of 66 papers), while load forecasting remains too sparse for subfield-level conclusions. Only four studies reach pilot-level evidence, and none reports sustained operational deployment. No study was classified as ready-now for industry adoption; three were near-term and 63 research-only. Nevertheless, several bounded, human-in-the-loop uses merit near-term trials, including point-name normalisation, document-grounded operator support, BEM workflow assistance, and advisory interfaces around physics-based controllers. Conventional machine learning (ML), model predictive control (MPC), reinforcement learning (RL) and ontology-based tools remain more adopted for high-frequency control, short-horizon numerical forecasting, and well-posed ontology mapping, while autonomous agentic operation and unvalidated occupant proxies remain research-stage. Current evidence therefore supports LLMs primarily as semantic and workflow layers rather than autonomous HVAC controllers. Future work should prioritise field-validated benchmarks, orchestration evaluation under operational constraints, and LLM-MPC/RL architectures with bounded latency and verifiable safety properties.
Highlights
This review codes 66 LLM-for-HVAC studies across five application and three method categories, assessing evidence, deployment readiness, and responsibility boundaries. It finds no ready-now role while identifying bounded near-term uses.
- 66 studies are coded across 5 HVAC application and 3 LLM method categories.
- No study qualifies as ready-now; 3 are near-term and 63 remain research-only.
- The framework links evidence realism, deployment readiness, and responsibility boundaries between LLM outputs and physical HVAC decisions.
- Point-name cleanup, operator support, and BEM assistance are identified as bounded roles suitable for near-term trials.
- ML, MPC, and RL remain popular AI approaches for HVAC applications.
1. Introduction
HVAC operations are constrained by fragmented information, limited trust, integration difficulties, and governance requirements. The review responds with a structured framework for matching LLM roles to tasks, evidence, safety, and deployment boundaries.
- Building automation systems contain extensive sensor data, but cryptic point names, missing metadata, and inaccessible manuals obstruct diagnosis and control.
- Deployment barriers include trust, integration, and governance, extending beyond numerical optimisation to semantic, technical, and organisational constraints.
- LLMs are positioned primarily as semantic and workflow layers connecting manuals, BAS data, BIM models, ontologies, and operational records.
- The review asks where LLMs can improve operational decision-making without unacceptable safety, governance, or reproducibility risks.
- The framework maps five HVAC application categories against three LLM method categories and evaluates evidence realism, readiness, and responsibility boundaries.
- The practical output distinguishes bounded human-supervised roles from roles requiring further safety mechanisms, field validation, and governance.
2. Background
The background defines the review’s M1–M3 method taxonomy and positions LLMs as semantic infrastructure for fragmented HVAC information. It distinguishes language-based orchestration from conventional control and ontology tools.
- M1–M3 methods: M1 covers prompting and supervised adaptation, including in-context learning, prompt engineering, fine-tuning, LoRA, and PEFT.
- M1–M3 methods: M2 grounds LLM outputs in curated information, retrieval, memory, tools, and multi-step orchestration workflows.
- M1–M3 methods: M3 is defined by non-textual inputs such as floor plans, piping diagrams, heatmaps, sensor plots, or video.
- HVAC operations accumulate heterogeneous BAS, manual, BIM, sensor, ontology, and operator information that no single pipeline integrates automatically.
- The LLM is framed as a semantic and workflow layer that routes fragmented context toward five HVAC application categories, with guardrails before physical actuation.
- The review addresses a gap left by prior reviews that separately emphasised broad coverage, controller performance, governance questions, or other AI modalities.
3. Methodology
The review searches Scopus-indexed peer-reviewed literature, screens papers for substantive HVAC tasks and evaluative evidence, and codes each study by application, method, evidence, readiness, and deployment constraints.
- The methodology combines a search strategy and corpus construction with a taxonomy and coding scheme for application, method, evidence, and deployment constraint.
- The search query combines LLM terms with HVAC or building-energy terms and restricts document types to articles, reviews, and conference papers.
- The corpus is limited to English-language, Scopus-indexed, peer-reviewed publications, excluding much grey literature and rapidly moving engineering practice.
- Eligible borderline studies required a concrete HVAC operational task, system, or dataset together with evaluative evidence.
- Each paper receives primary and, when applicable, secondary application assignments, while quantitative counts use the primary assignment.
- Method coding follows dominant integration logic: agentic workflows and solver feedback are M2, while vision–language input defines M3.
- Evidence and deployment coding records evidence type, readiness, responsibility boundary, and quality criteria, with definitions summarised in Table 2.
4. Application-Level Analysis
The application analysis shows uneven evidence across HVAC use cases: fault diagnosis has comparatively strong empirical grounding, forecasting remains too small for robust conclusions, and BEM is the largest and most methodologically diverse category. Across applications, LLMs are mainly used for bounded assistance, orchestration, and semantic support, while deployment remains constrained by validation, knowledge maintenance, latency, and reliability risks.
- A1 Fault Detection and Diagnosis: A1 fault diagnosis includes field, laboratory, and user-evaluation studies, but no paper reaches pilot or operational evidence.The corpus includes six retrospective field-data studies, two lab or user-evaluation studies, and one simulation-only study.
- A1 Fault Detection and Diagnosis: Knowledge scaffolding improved HVAC diagnosis and document retrieval, but its gains depend on complete, maintained manuals, ontologies, and rule bases.Graph-constrained reasoning achieved 89.1% accuracy while reducing candidate rules by more than 95%; Re2G improved retrieval relevancy by 23.5% and reduced hallucinations by 31.2% versus standard RAG.
- A1 Fault Detection and Diagnosis: Fine-tuned small models matched strong fault-classification performance on well-specified benchmarks, but transfer across equipment classes required separate fine-tuning.Reported F1 scores ranged from 82% to 99%, while comparable performance across VAV and chiller equipment required class-specific fine-tuning.
- A2 Energy Load Forecasting: A2 forecasting contains only three studies, all using retrospective field data, so it cannot support robust subfield comparisons or live-pipeline claims.The studies use LLMs indirectly as code generators, feature adapters, or descriptor layers rather than direct forecasters.
- A3 Building Energy Modelling and Simulation: BEM is the largest application category with 32 papers and the dominant application–method cell is A3×M2 with 15 papers.BEM is also methodologically mixed, with M1 close behind at 14 papers.
- A3 Building Energy Modelling and Simulation: BEM applications span model generation, retrofit advice, model updating, semantic interoperability, and simulation workflows, but many remain advisory or future-facing.Engineer review reduces immediate risk in advisory settings, while model updating and executable workflow automation remain limited by validation, retrieval, tool-use, and orchestration concerns.
4.4. HVAC Control and Optimisation (A4)
A4 is the review’s most safety-critical application category, but its evidence remains dominated by simulation and bounded human-in-the-loop roles. The strongest supported pattern is advisory, explanatory, and design support rather than unconstrained autonomous control.
- Methodological patterns: Advisory, explanatory, and design-support roles are the most credible A4 uses.Examples include RAG-based setpoint advice, explanation layers, Brick-based controller design, and retrieval-augmented policy optimisation.
- Field evidence: 47.9 % energy savings were reported for the LLM-only Office-in-the-Loop condition across two deployment periods spanning approximately 7.5 weeks.Both studies remained human-in-the-loop rather than autonomous safety-critical control.
- Evidence and readiness: 15 primary-coded papers span A4, but 12 are simulation-only, two are near-term Office-in-the-Loop pilots, and one uses retrospective field data.The two pilots report real-office deployment with real occupants and measured outcomes.
- Deployment barriers: Safety remains the binding constraint because prompt engineering alone cannot enforce comfort, setpoint, and equipment-capacity constraints.Physics guardrails are required for high-frequency setpoint generation.
- Deployment barriers: Approximately 2 s per control step suits slow supervisory loops such as chiller sequencing but remains incompatible with sub-second HVAC control cycles.Latency is one of the barriers introduced by agentic, multi-step LLM designs.
- Occupant-facing applications: Digital occupants and comfort proxies remain research-only because their results depend on uncalibrated assumptions about synthetic language-based behaviour.The primary A5 corpus also lacks sustained deployment realism and longitudinal evidence of preference stability.
5. Cross-Cutting Methodological Analysis
The corpus is concentrated in context-grounded BEM workflows, while control and other combinations remain less established because their evaluation requires stronger operational realism. Across categories, evidence is mostly offline or pre-deployment, with only four pilot-coded papers and no operational deployments.
- A1–A5 × M1–M3 distribution: 15 papers form the dominant A3×M2 cell, followed by 14 papers in A3×M1.Context-grounded orchestration for BEM is the largest single cluster in the corpus.
- A1–A5 × M1–M3 distribution: A1×M3, A2×M3, and A5×M3 contain no papers, while A2×M2 contains only one paper.Sparse cells indicate limited evidence rather than established subfields or necessarily promising research gaps.
- Evaluation asymmetries: BEM is easier to demonstrate than control because static benchmarks offer cleaner success criteria and engineers can inspect generated models before use.Control requires closed-loop real-building validation under safety constraints, while operational time series are harder to share and often incomplete.
- Models and temporal trends: GPT-4 and GPT-4o form the largest single model group across all five application categories and all publication years covered.This concentration creates dependence on a small proprietary-model ecosystem; open-weight models are concentrated in constrained fine-tuning or representation-learning studies.
- Models and temporal trends: Context-grounded orchestration increased from 4 of 10 method codes in 2024 (40 %) to 24 of 45 in 2025 (53 %).It overtook prompting and supervised adaptation as the most common method category across the reported period.
- Evidence realism gap: Only four papers were pilot-coded, representing three distinct field deployments, and none reported operational deployment.Only the two Office-in-the-Loop papers were both pilots and near-term; the other two pilots remained research-only.
- Evidence realism gap: 36 of 66 papers met the data-realism criterion, whereas experimental or field realism was met by 13 of 66 papers.Closed-loop validation, reproducibility, and explicit safety or governance reporting were weaker and unevenly reported.
6. Discussion
The discussion frames LLMs as bounded semantic and workflow layers whose deployment value depends on evidence realism, safety boundaries, and human oversight. Current evidence supports near-term trials for infrastructure tooling and advisory roles, while autonomous physical decision-making remains research-stage.
- Deployment framework: The review uses deployment readiness, responsibility boundaries, and evidence realism to distinguish bounded human-supervised roles from roles requiring further validation and governance.The framework is intended to guide where LLMs are technically plausible and operationally defensible.
- Deployment framework: No reviewed study qualifies as ready-now; three are near-term and 63 remain research-only.Ready-now status requires sustained or repeated operational deployment of a reusable, bounded artefact under human oversight.
- Near-term roles: Five defensible near-term workflow roles are point-name cleanup, grounded document support, BEM assistance, MPC/RL advisory interfaces, and technician training.These role types are broader than individual readiness classifications and remain bounded by hard constraints.
- Fit-for-purpose boundary: LLMs are most useful when HVAC work requires interpreting and coordinating heterogeneous information, rather than solving already well-posed numerical prediction or optimisation tasks.The review therefore positions LLMs as semantic or advisory layers alongside specialised numerical and rule-based methods.
- Failure modes: Reasoning brittleness and semantic hallucination expose deployment risks that standard accuracy metrics may miss.Reported issues include run-to-run recommendation variation, incorrect causal reasoning, and syntactically valid code implementing incorrect feature-engineering steps.
7. Conclusion
This review maps the LLM-for-HVAC literature through application, method, evidence, readiness, and responsibility dimensions. It finds that current evidence supports bounded semantic and advisory roles, while field validation, safety architecture, and governance remain the main barriers to deployment.
- Scope and framework: The review maps 66 peer-reviewed studies across five application categories and three method categories, assessing evidence realism, deployment readiness, and responsibility boundaries.The corpus is concentrated in building energy modelling, with context-grounded orchestration as the dominant methodological cluster.
- Evidence and readiness: Only four pilot studies were identified; no role is ready-now, three are near-term, and 63 papers remain research-only.Offline and pre-deployment evaluation dominates the evidence base.
- Conclusion: LLMs are most defensible as documentation interfaces, integration accelerators, semantic support layers, and advisory wrappers around established tools and controllers.Their present value is reducing integration friction, improving access to operational knowledge, and supporting human decision-making rather than replacing it.
CRediT authorship contribution statement
The authorship statement assigns conceptualization, methodology, investigation, analysis, data curation, visualization, writing, supervision, resources, project administration, and funding roles across the listed contributors.
- Contributions: Alexander Neubauer led conceptualization, methodology, investigation, formal analysis, data curation, visualization, and drafting.He also contributed to review and editing.
- Contributions: Tianzhen Hong contributed conceptualization, review and editing, supervision, and resources.Han Li contributed conceptualization and review and editing.
- Contributions: Han Li, Mengbo Yu, Amin Darbandi, Yannick Fürst, and Martin Kriegel contributed through the roles listed in the statement.Martin Kriegel’s listed roles include project administration, funding acquisition, supervision, and resources.
Declaration of Competing Interest
The authors disclose editorial conflicts involving two authors and request independent handling of the manuscript’s editorial and peer-review process.
- Competing interests: Tianzhen Hong is the journal’s Executive Editor and Han Li is Guest Editor for the special issue receiving the manuscript.The authors request that neither participate in editorial decision-making or peer review.
- Competing interests: Two studies included in the review corpus were co-authored by members of the author group.
Declaration of Generative AI and AI-assisted technologies in the writing process
The authors used ChatGPT, Claude, and NotebookLM for language editing, readability improvement, and structured extraction, then reviewed and edited the resulting content.
- The authors used ChatGPT, Claude, and NotebookLM to assist with language editing, readability improvement, and structured literature extraction.They reviewed and edited the content after using these tools.
Appendix A. Corpus Summary Tables
The appendix provides per-application-category summaries and selected study tables covering the reviewed corpus, with complete paper-level coding supplied separately.
- The appendix tables summarise system scope, LLM methodology, key results, primary limitations, and deployment-readiness classifications.
- Tables A1 and A2 provide primary-coded sets for fault detection and diagnosis and energy load forecasting.
- Tables A3 and A4 cover selected studies on building energy modelling, simulation, interoperability, and HVAC control and optimisation.The A4 table excludes the advisory-MPC wrapper of Liang et al., which is characterised elsewhere.
- Table A5 covers primary-coded studies on thermal comfort and occupant interaction, while the appendix also identifies studies such as the Building Genome Project and two real buildings in Shenzhen.
Appendix B. Supplementary Reference Tables
The supplementary tables distil deployment archetypes, stakeholder-specific guardrails, and reported cost and latency figures from the reviewed corpus.
- Table A6 presents fit-for-purpose deployment archetypes distilled from the corpus, with use horizons representing role-level practitioner guidance rather than per-study readiness codes.Empty cost and latency cells indicate that the source papers reported no direct figure.
- Table A7 organises defensible LLM roles, main risks, and required guardrails by stakeholder group.
- Table A8 reports cost and latency figures from the reviewed corpus, leaving cells empty when source papers provided no direct value.