Paper deep dive
ADAPT: Physics-Aware Diffusion-based World Models for Adaptive Predictive Transferable HVAC Control
Xu Yang, Kailai Sun, Dianyu Zhong, Qianchuan Zhao
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/21/2026, 3:41:36 AM
Summary
The paper introduces ADAPT, a physics-aware conditional diffusion indoor environmental world model (IEWM) designed for robust and transferable HVAC control. ADAPT addresses challenges such as delayed thermodynamic responses, partial observability, and out-of-distribution (OOD) generalization across seasons and climate regions. It utilizes a diffusion model to predict a short-horizon thermal baseline under a held-action protocol, capturing latent thermal inertia. A key innovation is a learnable multi-zone heat-balance regularizer that constrains generated trajectories to satisfy physical thermodynamics without requiring explicit building geometry or manually calibrated parameters. Integrated with downstream reinforcement learning, ADAPT demonstrates significant improvements in energy efficiency and occupant comfort compared to state-of-the-art baselines, maintaining robustness under OOD scenarios.
Entities (10)
Relation Signals (9)
ADAPT → incorporates → Multi-zone Heat-Balance Regularizer
confidence 95% · a learnable multi-zone heat-balance regularizer constrains generated trajectories to satisfy transferable building thermodynamics
ADAPT → integrateswith → Reinforcement Learning (RL)
confidence 95% · The learned IEWM is subsequently integrated into a downstream reinforcement learning controller
ADAPT → reduces → occupant discomfort
confidence 95% · and occupant discomfort by 30.2% compared with state-of-the-art baselines
ADAPT → reduces → HVAC energy consumption
confidence 95% · ADAPT reduces HVAC energy consumption by 7.3% ... compared with state-of-the-art baselines
ADAPT → uses → Diffusion Model
confidence 95% · we propose ADAPT, a physics-aware conditional diffusion indoor environmental world model
Multi-zone Heat-Balance Regularizer → constrains → generated trajectories
confidence 90% · constrains generated trajectories to satisfy transferable building thermodynamics
ADAPT → evaluatedon → Sinergym
confidence 90% · Extensive experiments on SemibuildingSim and Sinergym demonstrate that ADAPT reduces HVAC energy consumption
ADAPT → evaluatedon → SemiBuildingSim
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Buildings account for roughly one-third of global energy consumption and CO$_2$ emissions. Optimizing indoor climate systems plays a critical role for urban climate mitigation aligned with UN Sustainable Development Goals 11 and 13. However, indoor delayed thermodynamic responses and partial observability severely hinder existing methods, which are primarily limited by implicit thermal inertia, occupancy dynamic prediction, and cumulative prediction errors, especially for out-of-distribution environments. In practice, these challenges are further exacerbated by the high cost and privacy burden of dense indoor sensing, forcing operators to collect only limited data in a single operating regime while expecting controllers to generalize reliably across unseen seasons and climate regions. To address this problem, we propose ADAPT, a physics-aware conditional diffusion indoor environmental world model for HVAC control. The model predicts a short-horizon held-action thermal baseline to capture the latent thermal inertia of the buildings. The diffusion backbone utilizes the robustness of generative models, while a learnable multi-zone heat-balance regularizer constrains generated trajectories to satisfy transferable building thermodynamics without requiring known building geometry or manually calibrated thermal parameters. A credit assignment is then design for the downstream reinforcement learning. Extensive experiments on SemibuildingSim and Sinergym demonstrate that ADAPT reduces HVAC energy consumption by 7.3\% and occupant discomfort by 30.2\% compared with state-of-the-art baselines under IID control. Under OOD control scenarios spanning unseen seasons and climate regions, ADAPT maintains robust performance with only marginal degradation relative to its IID performance, substantially outperforming existing methods in transfer robustness.
Tags
Links
- Source: https://arxiv.org/abs/2608.19804v1
- Canonical: https://arxiv.org/abs/2608.19804v1
Trouble viewing inline? Open PDF directly →
Full Text
84,881 characters extracted from source content.
Expand or collapse full text
ADAPT: A Diffusion-based Adaptive Physics-aware Indoor Environmental World Model for Transferable HVAC Control Xu Yang Department of Automation Tsinghua University, Beijing, China yangxu24@mails.tsinghua.edu.cn Kailai Sun ∗ SMART, Singapore Massachusetts Institute of Technology Cambridge, MA, USA skl24@mit.edu Dianyu Zhong Department of Automation Tsinghua University, Beijing, China Qianchuan Zhao ∗ Department of Automation Tsinghua University, Beijing, China zhaoqc@tsinghua.edu.cn Abstract Buildings account for about one-third of global energy consump- tion and CO 2 emissions. The optimization of indoor climate sys- tems plays a critical role in urban climate mitigation, aligning with UN Sustainable Development Goals 11 and 13. However, build- ing thermal dynamics are delayed and partially observable, while seasonal weather, solar radiation, and internal heat gains induce substantial differences across operating conditions. Existing meth- ods struggle to accurately predict indoor thermal dynamics and generalize robustly across different conditions. To address these challenges, we propose ADAPT, a novel physics-aware conditional diffusion indoor environmental world model (IEWM) for accurate and robust indoor climate control. ADAPT leverages the expres- sive generative capability of diffusion models to predict a thermal baseline that explicitly captures latent thermal inertia. We further incorporate a multi-zone physical heat-balance regularization to constrain transferable thermal dynamics, enabling physically con- sistent thermal baseline generation without requiring manually calibrated thermal parameters. The learned IEWM is subsequently integrated into a downstream reinforcement learning controller for delayed credit assignment. Experiments on SemiBuildingSim and Sinergym environments demonstrate that introducing IEWM into indoor climate systems improves downstream control perfor- mance, reducing HVAC energy consumption by 7.3% and occupant discomfort by 30.2% compared with the strongest baseline. More im- portantly, under cross-season and cross-climate out-of-distribution transfer, ADAPTmaintains robust control performance with only marginal degradation relative to IID, substantially outperforming sequential and data-driven world models. This interdisciplinary study offers an energy-efficient, transferable AI solution for con- trol science, building science, and energy sustainability. The code: https://github.com/xuyangthu88/ADAPT. Keywords World models, diffusion models, HVAC control, physics-informed machine learning, building energy, reinforcement learning ∗ Corresponding author. 1 Introduction The United Nations reports that 55% of the world’s population lived in urban areas in 2018, and this share is projected to reach 68% by 2050 [8,11]. Cities concentrate energy use, emissions, infrastructure, and human exposure to climate risk [28,40]. Buildings are a major part of this challenge. The UNEP and GlobalABC 2025 report further states that buildings consumed 32% of global energy and contributed 34% of global CO 2 emissions [5,50]. Building energy efficiency is important for climate action and urban sustainability [4, 7]. Heating, ventilation, and air-conditioning (HVAC) systems con- tribute the largest share of building operational energy consump- tion [39]. At the same time, climate change is making heating and cooling demand increasingly dynamic, raising the requirements for closed-loop efficient HVAC control [51]. Recent studies sug- gest that accurately capturing weather-driven building thermal dynamics requires finer-grained temporal modeling, as evolving climate patterns introduce larger fluctuations to building thermal loads in many regions [46,66,67]. On the other hand, the HVAC system plays a key role as an interface between climate mitiga- tion and human well-being. Indoor thermal comfort, as defined in ASHRAE Standard 55, is a basic condition for health, produc- tivity, and occupant satisfaction [15,47]. These global demands make occupant-centric HVAC control important at the intersec- tion of energy systems and occupancy [48]. To support building decarbonization, intelligent HVAC control should capture critical weather-driven thermal dynamics and remain robust under chang- ing weather and seasonal conditions, for energy efficiency and occupant comfort [21, 24, 55]. A practical challenge in HVAC control is the delayed and par- tially observed (PO) thermal response of buildings. Due to thermal inertia, heat is continuously stored and released by walls, floors, furniture, air volumes, and HVAC equipment, causing the effect of a control action to emerge only after multiple control intervals [53]. Meanwhile, occupancy, outdoor weather, and solar radiation jointly drive indoor thermal dynamics, while heat storage is not directly measured by sensors, rendering the latent thermal state only par- tially observable. Model Predictive Control (MPC) has long been the classic frame- work for HVAC control [1,14,27]. However, MPC is sensitive to model mismatch, disturbance forecasts, and solver overhead in real deployments. More recently, Deep Reinforcement Learning (DRL) 1 Conference’27, August 2027, San Jose, California, USAYang et al. has emerged as a promising alternative by directly learning con- trol policies from interaction data [32,44,52,55,56]. Nevertheless, most DRL methods remain fundamentally reactive and struggle to assign credit under delayed response systems [41]. Although recur- rent architectures and Transformer-based policies partially mitigate delayed thermal responses and the PO problem by exploiting his- torical observations [36], they still implicitly infer latent thermal dynamics from past trajectories, resulting in limited generalization across different conditions. These limitations naturally motivate learning an explicit predic- tive model of indoor thermal dynamics. Recently, world models have demonstrated remarkable success in long-horizon decision mak- ing by providing predictive representations for planning and value estimation in robotics and embodied AI [6,13,18,19,23,58,64]. Bringing this to HVAC control is appealing, as world model explic- itly models delayed thermal dynamics and provides informative future trajectories for downstream controllers. However, HVAC systems pose domain-specific challenges when applying world models in these systems. First, HVAC systems are partially observable and undergo substantial thermal distribution shifts across seasons, weather conditions, and climate regions, mak- ing accurate prediction considerably challenging. However, there is lack of world models that explicitly model indoor thermal dynamics for HVAC control. Second, sequential and vanilla data-driven world models tend to exploit specific statistical correlations of the train- ing dataset rather than physically invariant thermal mechanisms. As weather, solar radiation, and climate shift out of the training distribution, their predictions become inaccurate and physically implausible, directly degrading downstream control performance. To address the above challenges, we propose ADAPT (A Diffusion- based Adaptive Physics-aware indoor environmental world model for Transferable HVAC control), a physics-aware conditional dif- fusion IEWM for robust HVAC control. Specifically, we introduce a generative conditional diffusion model to capture the distribu- tion of future thermal baselines. To improve out-of-distribution generalization, we further introduce a physics-aware regulariza- tion that encourages generated baselines to satisfy fundamental thermal dynamics while preserving the ability of diffusion models. The regularization is based on a differentiable multi-zone physical heat-balance equation that models inter-zone exchange, outdoor- envelope exchange, solar radiation and internal heat gains. The IEWM is general and transferable, without requiring knowledge of the building geometry or hand-calibrated parameters. The predicted thermal baselines are used to support delayed credit assignment in downstream RL control. Our contributions are summarized as follows: • We introduce ADAPT, a physics-aware conditional diffusion IEWM for HVAC control. The IEWM predicts the thermal base- line that captures delayed indoor thermal dynamics and support downstream RL credit assignment. •We develop a physics-aware regularization based on a multi- zone heat-balance equation, encouraging diffusion-generated trajectories to satisfy fundamental thermal dynamics without requiring building-specific thermal parameters. •We demonstrate that ADAPT consistently improves in-distribution control performance on SemiBuildingSim and Sinergym, reduc- ing HVAC energy consumption by 7.3% and occupant discomfort by 30.2% compared with the strongest baseline. • Extensive transfer experiments across seasons (summer↔winter) and climate regions (Stockholm↔Arizona) demonstrate that ADAPTachieves substantially stronger robustness than baselines under OOD scenarios, leading to more generalized downstream HVAC control. 2 Related Work 2.1 Reinforcement Learning and Transfer for HVAC Control RL has become a leading data-driven alternative to model predictive control for HVAC operation, learning policies directly from interac- tion data without a hand-calibrated dynamics model [24,32,44,54– 56]. EnergyPlus-driven benchmarks such as Sinergym have further consolidated reproducible evaluation across zone temperature con- trol and demand response [9]. Yet HVAC control remains struc- turally hard for model-free RL: hidden thermal storage introduces delayed action effects under partial observability, and standard temporal-difference learning routes credit through this delay only implicitly. Recent surveys emphasize this gap and the resulting brittleness across seasons and buildings [35,38]. Transfer-oriented HVAC research has approached the problem by policy transfer across buildings [62] and online fine-tuning under distribution shift [12], but both still require target-domain interaction and offer limited guarantees when weather and internal loads move outside the training domain [2]. ADAPT takes a complementary route: rather than adapting the policy after deployment, it learns a gener- ative indoor environmental world model whose predicted thermal baseline transfers across seasons and climate regions and is fed to a branch-wise action-value network [49], so the controller receives an explicit forecast of delayed thermal drift instead of an implicit latent state. 2.2 World Models for Decision Making Learned dynamics have long been used for sample-efficient policy optimization and value expansion through PETS, MBPO, MBVE, and their extensions [16,18,26,58,63], and Dreamer, TD-MPC, and TD-MPC2 establish scalable latent world models for long-horizon planning [6,19,20,22,23]. Most of these methods target generic control domains where the observed state is largely sufficient and action-sequence search is the dominant challenge. HVAC control differs on both axes: observation is only a partial view of hidden thermal state, and the deployment distribution shifts across sea- sons and climate zones, so the world model is stressed more by prediction than by planning. ADAPT reflects this distinction by using a generative indoor environmental world model to forecast the future indoor trajectory under an imagined control assumption, and by providing the resulting thermal baseline to a downstream RL controller for baseline-conditioned action evaluation and delayed credit assignment. 2.3 Diffusion Models for Time Series and Control Diffusion models generate data by reversing a noise process [17,37, 45] and have been extended to probabilistic time-series forecasting 2 ADAPT: A Diffusion-based Adaptive Physics-aware Indoor Environmental World Model for Transferable HVAC ControlConference’27, August 2027, San Jose, California, USA and imputation [30,42]. In sequential decision making, Diffuser recasts planning as trajectory denoising and diffusion policies gen- erate control sequences [10,25]; in energy, recent work applies diffusion to probabilistic building load forecasting [29,57]. These works establish diffusion as a powerful tool for uncertain temporal prediction, but they do not target the two properties HVAC control needs most: physically consistent multi-step indoor trajectories under distribution shift, and a control interface that turns those trajectories into decision-relevant thermal baselines. ADAPTcloses this gap with a conditional diffusion IEWM whose multi-step roll- outs are constrained by a learnable heat-balance regularizer, so generated trajectories remain on physically plausible manifolds across seasons, weather, and climate regions rather than reproduc- ing source-domain correlations. 3 Method We propose the ADAPT framework. The key idea is to learn a trans- ferable indoor environmental world model (IEWM) that captures delayed indoor thermal dynamics, enabling HVAC controllers to make efficient decisions across various seasons and climate regions. 3.1 Problem Formulation We model the optimal HVAC control problem as a Partially Ob- servable Markov Decision Process (POMDP), defined by the tuple M= (S,O,A,푃,푅,훾). In building environment, the true under- lying physical stateSincludes unmeasurable thermal storage in structural mass (e.g., concrete walls and windows). However, at time step푡, the agent only receives partial sensory observations 표 푡 ∈ O: 표 푡 =[T 푡 , H 푡 , O 표푐 푡 , T 표푢푡 푡 , S 푡 , Z 푡 ],(1) where T 푡 ∈ R 푁 is the vector of zone temperatures,퐻 푡 is relative humidity,푂 표푐 푡 denotes measured or estimated occupancy pattern, 푇 표푢푡 푡 denotes outdoor temperature,푆 푡 denotes solar radiation inten- sity, and 푍 푡 collects auxiliary sensor variables. The controller maps observations to a multidimensional discrete action vector푎 푡 ∈ A, representing discretized HVAC control de- mands such as fan modes or temperature setpoints. The underlying thermal state evolves according to the thermodynamic transition 푃(푠 푡+1 |푠 푡 ,푎 푡 ). Due to the thermal inertia of building structures, the effect of an action at time푡may only become clearly visible after several control intervals. The objective is to optimize a policy휋 that maximizes the discounted return: max 휋 E 휋 Í ∞ 푡=0 훾 푡 푟 푡 . 3.2 Generative IEWM Modeling HVAC systems are partially observable and exhibit substantial dis- tribution shifts across seasons, weather conditions, and climate regions, posing a greater challenge to world-model prediction than action-planning. To capture the resulting uncertainty and complex thermal dynamics, we formulate indoor environmental modeling as a conditional trajectory generation problem and learn a conditional diffusion IEWM, which models the distribution of a future thermal baseline conditioned on historical trajectories. Specifically, the conditioning context consists of historical trajec- tories together with a held-action protocol that specifies the future control sequence, and is defined as c 푡 =(표 푡−푁 :푡 ,푎 푡−푁 :푡−1 ,푎 푡 :푡+퐻−1 = 푎 푡−1 ),(2) By assuming held-action, the prediction trajectory ˆ 푦 푡 serves as a counterfactual baseline that captures the building’s latent thermal inertia when the previous action푎 푡−1 is maintained. This baseline provides a reference for evaluating the long-term impact of the cur- rent control action, thereby improving delayed credit assignment in reinforcement learning. The IEWM predicts a future thermal baseline under the conditioning context: ˆ y 푡 = ˆ 표 푡+1:푡+퐻 ∼ 푝 휃 (y 푡 | c 푡 ),(3) To align the training objective, we enforce a held-action protocol (푎 푡 :푡+퐻−1 = 푎 푡−1 ) during offline data collection, utilizing the ground truth 표 푡+1:푡+퐻 to provide supervision. Given the conditioning context c 푡 , the IEWM models the condi- tional distribution of future indoor thermal baseline using a denois- ing diffusion probabilistic model (DDPM). The diffusion process consists of a forward noising process and a learned reverse denois- ing process. Forward noising process. Given a normalized future trajec- tory y 0 푡 , the forward diffusion process gradually perturbs it with Gaussian noise. Following DDPM, the noisy trajectory at diffusion step 푘 is sampled directly from 푞(y 푘 푡 | y 0 푡 )=N( √ ̄ 훼 푘 y 0 푡 , (1− ̄ 훼 푘 )I),(4) where ̄ 훼 푘 is a predefined variance schedule. Reverse denoising process. The reverse process reconstructs clean trajectories conditioned on the HVAC context c 푡 . Following the state diffusion formulation of GlobeDiff [61], the reverse Markov chain is defined as 푝 휃 ( ˆ y 0:퐾 푡 | c 푡 )= 푝( ˆ y 퐾 푡 ) 퐾 Ö 푘=1 푝 휃 ( ˆ y 푘−1 푡 | ˆ y 푘 푡 , c 푡 ),(5) where the terminal distribution is푝( ˆ y 퐾 푡 )=N(0,I). Each reverse transition is parameterized as 푝 휃 ( ˆ y 푘−1 푡 | ˆ y 푘 푡 , c 푡 )=N( ˆ 흁 휃 ( ˆ y 푘 푡 ,푘, c 푡 ), 훽 푘 I),(6) where the denoising network predicts the injected noise. The re- verse mean is computed as ˆ 흁 휃 = 1 √ 훼 푘 ˆ y 푘 푡 − 훽 푘 √ 1− ̄ 훼 푘 ˆ 휖 휃 ( ˆ y 푘 푡 ,푘, c 푡 ) .(7) The denoising network ˆ 휖 휃 takes as input the noisy trajectory, the diffusion step, and the conditioning context, enabling trajectory generation consistent with the historical observations, weather conditions, and actions. Training objective. For each training sample, we randomly select a diffusion step, inject Gaussian noise according to Eq. 4, and optimize the denoising network to predict the injected noise: L diff = E h 휖− ˆ 휖 휃 ( √ ̄ 훼 푘 y 0 푡 + √︁ 1− ̄ 훼 푘 휖, 푘, c 푡 ) 1 i .(8) Inference. During inference, the model starts from Gaussian noise and iteratively performs reverse denoising conditioned on c 푡 : ˆ y 푘−1 푡 = ˆ 흁 휃 ( ˆ y 푘 푡 ,푘, c 푡 )+ √︁ 훽 푘 휖 푘 , 휖 푘 ∼N(0, I).(9) The final sample ˆ y 0 푡 is taken as the predicted thermal baseline, denoted by ˆ y 푡 , which is used for evaluating delayed action effects. 3 Conference’27, August 2027, San Jose, California, USAYang et al. Physics-aware Diffusion IEWM Conditions Conditional diffusion model 흐 휽 Thermal baseline ෝ 풚 풕 = ෝ 풐 풕+ퟏ ,⋯, ෝ 풐 풕+푯 Gaussian noise 풚 풕 푲 ~푵(ퟎ;푰) Denoising 풐 풕−푯:풕 풂 풕−푯:풕−ퟏ Physics-aware Diffusion IEWM Training Physics-aware Regularization Physics-aware Regularization 푪 풊 푻 풊,풌+ퟏ −푻 풊,풌 /∆풕=푸 풊,풌 풛풐풏풆 +푸 풊,풌 풐풖풕 +푸 풊,풌 풓풂풅 +푸 풊,풌 풊풏풕 +푸 풊,풌 풉풗풂풄 +풅 풊 푳=푳 풅풊풇 +흀 풕풓풖풆 푳 풕풓풖풆 + 흀 풑풓풆풅 푳 풑풓풆풅 +흀 풅 풅 ퟐ RL Policy Training Dataset Control Action Thermal Baseline Online RL Policy Learning IEWM Transfer Value Network Advantage Network Thermal baseline ෝ 풚 풕 풐 풕−푯:풕 풂 풕−푯:풕−ퟏ IEWM Conditional Sampling Only 풐 풕 푄 푡+퐻 푸(풐 풕 ,풂 풕 , ෝ 풚 풕 ) V(풐 풕 ) 푡 푨(풐 풕 ,풂 풕 , ෝ 풚 풕 ) Cross season transfer Cross climate-region transfer Figure 1: Overall framework of ADAPT. The training process of ADAPT is divided into two phases. Phase I: Physics-aware Diffusion IEWM Training. Phase I: Online RL Policy Learning. IEWM is frozen during phase I. 3.3 Physics-aware Thermal Regularization The conditional diffusion IEWM can effectively model offline trajec- tories. However, vanilla data-driven learning may exploit domain- specific statistical correlations rather than invariant thermal dy- namics. Such overfitting often generalizes poorly across seasons and climate regions. To improve out-of-distribution robustness, we introduce a physics-aware regularization based on a differentiable multi-zone physical heat-balance equation, which encourages gen- erated trajectories to satisfy fundamental thermal dynamics while preserving the flexibility of the diffusion model. For each zone푖, the predicted temperature is encouraged to satisfy the following heat-balance equation: 퐶 푖 ˆ 푇 푖,푘+1 − ˆ 푇 푖,푘 Δ푡 = ˆ 푄 푧표푛푒 푖,푘 + ˆ 푄 표푢푡 푖,푘 + ˆ 푄 푟푎푑 푖,푘 + ˆ 푄 푖푛푡 푖,푘 + ˆ 푄 ℎ푣푎푐 푖,푘 +푑 푖 .(10) Rather than serving as an exact physical constraint, Eq.(10)acts as a soft structural prior that regularizes diffusion predictions. All physical coefficients are jointly learnt with the diffusion model, allowing the regularizer to adapt to different buildings while main- taining physically meaningful heat-transfer mechanisms. The heat-balance equation consists of five interpretable com- ponents. The inter-zone term models conductive heat exchange among neighboring zones using a symmetric nonnegative coupling matrix: ˆ 푄 푧표푛푒 푖,푘 = 푁 ∑︁ 푗=1 푔 푖푗 ( ˆ 푇 푗,푘 − ˆ 푇 푖,푘 ), 푔 푖푗 =푔 푗푖 ≥ 0.(11) The outdoor-envelope term captures heat transfer between in- door and outdoor environments: ˆ 푄 표푢푡 푖,푘 = 푘 푖 (푇 표푢푡 푘 − ˆ 푇 푖,푘 ), 푘 푖 ≥ 0.(12) The radiation term models weather-dependent heat gains: ˆ 푄 푟푎푑 푖,푘 = s ⊤ 푖 r 푘 ,s 푖 ≥ 0,(13) The internal-gain term represents heat generated by occupants: ˆ 푄 푖푛푡 푖,푘 =훾 푖 ˆ 표 푖푛푡 푖,푘 , 훾 푖 ≥ 0.(14) Finally, the HVAC term models action-dependent heat exchange induced by the air-conditioning system, where ˆ 푢 ℎ푣푎푐 푖,푘 = 휙(푎 푖,푘 ) is learned: ˆ 푄 ℎ푣푎푐 푖,푘 = 휂 푖 ˆ 푢 ℎ푣푎푐 푖,푘 (푇 푠푢푝 푖,푘 − ˆ 푇 푖,푘 ), 휂 푖 ≥ 0.(15) A residual bias term푑 푖 captures varying unmodeled disturbances. We further penalize its magnitude to discourage the model from explaining missing physics through an unconstrained bias. Collec- tively, these structured components encourage the diffusion IEWM to capture transferable thermal dynamics instead of relying on specific data, thereby improving robustness in domain transfer. Joint Training with Data and Physics. We jointly optimize the conditional diffusion IEWM and the physics-aware regular- ization through two complementary thermal residual objectives with distinct roles. The first residual is evaluated on ground-truth trajectories to identify the physical coefficients of the heat-balance model, independent of the diffusion model. Based on these learned coefficients, the second residual is evaluated on diffusion-generated trajectories to regularize the conditional diffusion IEWM, encour- aging generated predictions to satisfy physical thermal dynamics. For an observed trajectory, we define the ground-truth thermal residual as R 푖,푘 =퐶 푖 푇 푖,푘+1 −푇 푖,푘 Δ푡 −(푄 푧표푛푒 푖,푘 +푄 표푢푡 푖,푘 +푄 푟푎푑 푖,푘 +푄 푖푛푡 푖,푘 +푄 ℎ푣푎푐 푖,푘 +푑 푖 ). (16) The corresponding parameter-identification loss is L true = 1 퐻푁 푡+퐻−1 ∑︁ 푘=푡 푁 ∑︁ 푖=1 휌(R 푖,푘 (y 푡 )),(17) where휌(·)denotes the Huber loss. Using the identified physical coefficients, we further define the thermal residual for diffusion- generated trajectories: ˆ R 푖,푘 =퐶 푖 ˆ 푇 푖,푘+1 − ˆ 푇 푖,푘 Δ푡 −( ˆ 푄 푧표푛푒 푖,푘 + ˆ 푄 표푢푡 푖,푘 + ˆ 푄 푟푎푑 푖,푘 + ˆ 푄 푖푛푡 푖,푘 + ˆ 푄 ℎ푣푎푐 푖,푘 +푑 푖 ). (18) The corresponding prediction regularization loss is L pred = 1 퐻푁 푡+퐻−1 ∑︁ 푘=푡 푁 ∑︁ 푖=1 휌( ˆ R 푖,푘 ( ˆ y 푡 )).(19) This objective directly regularizes the conditional diffusion IEWM. It encourages generated trajectories to satisfy the learned multi- zone heat-balance dynamics, thereby reducing physically implausi- ble predictions and discouraging season-specific statistical overfit- ting under distribution shifts. The overall training objective is L=L diff + 휆 true L true + 휆 pred L pred + 휆 푑 ∥d∥ 2 2 .(20) whereL diff is the standard diffusion denoising objective,L true identifies the parameters of the heat-balance model from observed trajectories,L pred regularizes the conditional diffusion IEWM with physical dynamics. The bias penalty further discourages the regu- larizer from explaining missing physical effects through the uncon- strained residual term, resulting in the IEWM that is both physically consistent and robust to out-of-distribution operating conditions. 4 ADAPT: A Diffusion-based Adaptive Physics-aware Indoor Environmental World Model for Transferable HVAC ControlConference’27, August 2027, San Jose, California, USA 3.4 Diffusion IEWM for RL learning To leverage the proposed IEWM for downstream HVAC control, we integrate it into an online RL framework. After offline train- ing, the IEWM is frozen and queried twice at each decision step. Before action selection, it predicts a thermal baseline, estimating how the indoor environment would evolve if the previous HVAC action were maintained. This baseline provides an explicit forecast of delayed thermal dynamics for action evaluation. After an action is selected, the same IEWM predicts a committed-action rollout by repeatedly applying the selected action, providing imagined future transitions for delayed credit assignment. The former im- proves action evaluation, whereas the latter improves temporal credit assignment. For multidimensional HVAC control, we follow the branch-wise action-value architecture in [49], where each branch independently selects one actuator value 푎 푡,푑 . The thermal baseline ˆ 푦 푡 is incorpo- rated into the branch-wise action-value function as 푄 푑 (표 푡 ,푎 푡,푑 , ˆ y 푡 ;Φ)=푉 휓 (표 푡 )+퐴 휂 푑 (표 푡 ,푎 푡,푑 , ˆ y 푡 )− max 푎 ′ ∈A 푑 퐴 휂 푑 (표 푡 ,푎 ′ , ˆ y 푡 ), (21) The thermal baseline is injected only into the advantage stream, enabling action preferences to be evaluated with explicit knowledge of future thermal dynamics while keeping state-value estimation based solely on current observation. After executing the selected action, the environment provides the real one-step transition(푥 푡+1 ,푟 푡 ). Starting from this, the IEWM predicts the future trajectory assuming the selected action is con- tinuously executed ˆ y ′ 푡 =[ ˆ 표 (푎 푡 ) 푡+2 , . . ., ˆ 표 (푎 푡 ) 푡+퐻+1 ] ∼ 푝 휃 (y | 표 푡−퐿+1:푡+1 ,푎 푡−퐿+1:푡 , ̃ 푎 푡+1:푡+퐻 = 푎 푡 ). (22) This committed rollout estimates the delayed thermal effect of executing the chosen action. Future rewards are evaluated using the environment reward function without learning an additional reward model. For horizon ℎ, the expanded return is 퐺 (ℎ) 푡 (푎 푡 )= 푟 푡 + ℎ−1 ∑︁ 푘=1 훾 푘 ˆ 푟 (푎 푡 ) 푡+푘 +훾 ℎ 푉 ̄ 휓 ( ̃ 푥 (푎 푡 ) 푡+ℎ ),(23) where ˆ 푟 (푎 푡 ) 푡+푘 denotes the imagined future reward computed from the committed rollout, and푉 ̄ 휓 is the value function of the target network. To balance short and long-term delayed effects, the multi- step returns are aggregated into a TD-휆 target: 휏 휆 푡 =(1− 휆) 퐻−1 ∑︁ ℎ=1 휆 ℎ−1 퐺 (ℎ) 푡 (푎 푡 )+ 휆 퐻−1 퐺 (퐻) 푡 (푎 푡 ).(24) This target propagates delayed comfort and energy feedback through the IEWM prediction, enabling more effective credit assignment under long thermal inertia. The branch-wise Q-network is optimized by Eq. 25. Full pseu- docode is provided in Algorithm A. L RL (Φ)= E D rl " 1 퐷 퐷 ∑︁ 푑=1 휏 휆 푡 −푄 푑 (표 푡 ,푎 푡,푑 , ˆ y 푡 ;Φ) 2 # .(25) 4 Experiments and Results Our empirical evaluation is designed to answer three research ques- tions (RQs): •RQ1: World-model benefit. To what extent does IEWM im- prove HVAC control performance compared to existing HVAC control methods? •RQ2: Physics-aware modeling. How effectively does the pro- posed physics-aware regularization improve prediction perfor- mance and thermal physical consistency over vanilla data-driven IEWMs? •RQ3: Transferable world model for control. How effectively does the proposed physics-aware IEWM improve the HVAC con- trol transferability under seasonal and climate-region transfer? 4.1 Experimental Setup and Evaluation Metrics Environments. We introduce SemiBuildingSim, a high-fidelity HVAC simulation benchmark derived from a real 7-zone commer- cial office building in Hebei, China [59,65]. The simulator is cali- brated using over 9,000 time-aligned operational records, including weather, occupancy, and HVAC control logs, thereby capturing realistic occupant-driven disturbances, and long-horizon thermal inertia. We evaluate seasonal transfer by training and testing across the building’s summer and winter operating conditions. The con- trol interval is 5 minutes. To further evaluate cross-climate robustness, we employ Siner- gym [9], an EnergyPlus-based open-source benchmark, using the 2ZoneDataCenterHVACenvironment. We perform climate-region transfer between Stockholm, Sweden, characterized by a cold con- tinental climate, and Arizona, USA, characterized by a hot desert climate, creating a challenging out-of-distribution evaluation across different climatic conditions. The control interval is 15 minutes. Prediction metrics. We report mean absolute error (MAE), root mean squared error (RMSE), coefficient of variation of the RMSE (CVRMSE) for key temperature and humidity dimensions, and occupancy exact-match rate (Occ EMR) when applicable. Room- temperature and return-temperature CVRMSE are the primary evaluation metrics, as they are the key thermal variables governed by the multi-zone RC heat-balance equations and directly reflect both physical consistency and control-oriented prediction quality. HVAC control Metrics. HVAC control is a multi-objective op- timization problem that requires balancing thermal regulation, en- ergy efficiency, and control stability. Accordingly, we adopt evalua- tion metrics that are tailored to the objectives of each benchmark. Thermal performance is benchmark-dependent: for SemiBuildingSim, we report occupant-centric comfort metrics, whereas for the Siner- gym, we report the average temperature violation. •Occupant-centric thermal comfort: We evaluate thermal com- fort using the Predicted Percentage of Dissatisfied (PPD) and the absolute Predicted Mean Vote (|PMV|) according to ASHRAE Standard 55. Comfort metrics are computed primarily during occupied periods. •Energy consumption: We report the total HVAC electricity consumption (kWh), which quantifies the energy efficiency and operational cost of the control policy. •Actuator action fluctuation: We evaluate control smoothness usingAF= 1 푇 Í 푇 푡=1 Í 퐷 푑=1 (푎 푡,푑 − 푎 푡−1,푑 ) 2 , where lower values in- dicate smoother actions, reducing unnecessary high-frequency control oscillations and improving operational stability. 5 Conference’27, August 2027, San Jose, California, USAYang et al. •Average temperature violation: Defined as the time-averaged absolute temperature violation outside the prescribed target tem- perature range, where zero indicates that the indoor temperature always remains within the acceptable bounds. 4.2 Main Results Answer for RQ1: To answer RQ1, we first evaluate ADAPT on SemiBuildingSim against MPC [22], model-free RL (A2C [33], DQN [34], PPO [43], BDQ [49]), history-aware RL (TransformerRL [36]), and model-based RL (MBVE [16], MBPO [26], DreamerV3 [20]). All controllers are trained and evaluated under the in-distribution (ID) setting. As shown in Tab. 4, ADAPT achieves the best balance be- tween energy efficiency, occupant comfort, and control smoothness, which demonstrates that introducing the diffusion IEWM provides a clear benefit for HVAC control. Table 1: Performance comparisons on SemiBuildingSim Sum- mer (mean± std over three seeds, ID). Lower is better. AlgorithmEnergy (kWh)↓ Abs PMV↓ PPD (%)↓ Action Fluctuation↓ MPC182.34± 1.59 0.60± 0.04 17.45± 1.1916.57± 0.93 A2C157.81± 7.59 0.50± 0.06 14.05± 1.5612.89± 3.35 PPO154.83± 3.10 0.41± 0.01 12.37± 0.339.75± 0.85 DQN155.77± 2.03 0.49± 0.06 13.19± 1.889.86± 1.52 BDQ154.57± 5.23 0.37± 0.01 10.88± 0.7810.20± 1.17 TransformerRL 149.34± 2.93 0.37± 0.03 10.05± 0.509.42± 1.05 MBVE150.73± 5.88 0.36± 0.01 10.27± 0.149.61± 1.86 MBPO150.99± 3.12 0.38± 0.04 10.47± 0.4912.47± 1.85 DreamerV3149.43± 1.02 0.36± 0.01 9.88± 1.538.37± 0.15 ADAPT (Ours) 138.49± 1.71 0.22± 0.01 7.01± 0.056.73± 0.31 The improvement stems from the thermal baseline predicted by the IEWM, which summarizes the future evolution of the in- door thermal dynamics while maintaining the current action. By explicitly modeling delayed thermal responses and building ther- mal inertia, the controller can optimize against anticipated thermal dynamics rather than relying solely on instantaneous observations, thereby avoiding unnecessary heating and cooling actions. To further understand this benefit, Fig. 2 compares the training curves of different methods. ADAPT converges substantially faster and reaches a higher return, indicating that the predicted thermal baseline significantly improves RL sample efficiency. The Sinergym results (Fig. 3) further demonstrate that the bene- fits of the IEWM generalize beyond SemiBuildingSim. ADAPT si- multaneously reduces temperature violation while increasing en- ergy savings. The improvement is particularly pronounced during the summer months in Stockholm, when longer daylight hours and stronger solar heat gains produce larger and more dynamic cooling loads. By predicting a thermal baseline that captures delayed indoor thermal dynamics, the IEWM enables proactive cooling decisions, suppressing unnecessary HVAC actuation while maintaining tem- peratures within the desired operating range. Additional results are shown in C. Answer for RQ2: To answer RQ2, we compare four IEWM designs: the full ADAPT (Ours), w/o Phys., which removes the proposed physics-aware regularization, MambaFormer [60], and a VAE-based [31] world model. Table 2 shows that the proposed physics-aware IEWM consistently achieves the best multi-step pre- diction performance under seasonal transfer. The improvement 0.00.20.40.60.81.0 Environment Steps 1e6 25000 22500 20000 17500 15000 12500 10000 7500 Episodic Return ADAPT(Ours) PPO TransformerRL DQN BDQ A2C Steps: 1.0× Value: 1.0× Steps: 2.14× Value: 0.57× Steps: 3.73× Value: 0.61× Steps: 4.88× Value: 0.44× Steps: 1.28× Value: 0.62× Steps: 3.90× Value: 0.49× Figure 2: Training curves on SemiBuildingSim Summer. Stars mark the first point where each method reaches 95% of its final episodic return. The labels report relative convergence steps (lower is better) and relative covergence return (higher is better), both normalized to ADAPT (Ours)= 1.0×. Jan Feb Mar Apr May Jun Jul Aug Sep Oct Nov Dec -6% 2% 10% 18% 26% (a) Energy Saving vs RBC Jan Feb Mar Apr May Jun Jul Aug Sep Oct Nov Dec -0.157 0.249 0.656 1.062 1.469 (b) Temperature Violation RBC A2C BDQ DQN PPO TransformerRL MBPO MBVE DreamerV3 ADAPT(Ours) Figure 3: Comparison of ADAPT with baselines in Sinergym Stockholm environment (ID). (a) energy savings relative to rule-based controller (RBC), (b) temperature violation. is particularly pronounced for room temperature and return tem- perature, while the purely data-driven diffusion model and other sequence modeling baselines degrade substantially under OOD conditions. Table 2: OOD prediction performance under seasonal transfer on SemiBuildingSim. Except Occ EMR, lower is better. CategoryMetricADAPT (Ours) w/o Phys. MambaFormer VAE Summer→ Winter (OOD) Overall MAE (↓)0.0350.0400.1190.131 RMSE (↓)0.0740.0830.2070.222 Occ EMR (%) (↑)98.47598.08677.02182.855 CVRMSE (%) Return Temp4.6338.0699.51713.367 Room Temp11.56016.13443.21044.983 Winter→ Summer (OOD) Overall MAE (↓)0.0490.0590.1810.204 RMSE (↓)0.0860.1050.3100.348 Occ EMR (%) (↑)77.48674.89276.97775.538 CVRMSE (%) Return Temp6.5319.23718.10521.279 Room Temp6.8479.56034.03138.174 The gains are most evident on room temperature and return temperature because these variables are directly governed by the multi-zone RC heat-balance equations and therefore provide the 6 ADAPT: A Diffusion-based Adaptive Physics-aware Indoor Environmental World Model for Transferable HVAC ControlConference’27, August 2027, San Jose, California, USA most stringent test of physical consistency. By explicitly regu- larizing the denoising process with thermal dynamics, the pro- posed physics-aware objective prevents the diffusion model from exploiting season-specific statistical correlations and instead en- courages predictions that remain consistent with the underlying heat-transfer process. Consequently, the learned world model main- tains substantially better thermal-state prediction when weather conditions and building operating regimes change. Figure 4 further demonstrates that the benefits of the proposed physics-aware IEWM generalize across climate regions. Under Stockholm and Arizona bi-transfer, ADAPT consistently achieves the best overall prediction performance across all reported met- rics, whereas purely data-driven world models exhibit substantially larger degradation after climate transfer. These results demonstrate that incorporating physics-aware thermal constraints significantly improves the cross-region robustness and generalization ability of the learned indoor environment world model. Table 3: Ablation study on the physics loss weight휆 pred and the bias penalty 휆 푑 . Lower is better. AblationValueMAE↓ RMSE↓ CVRMSE (%)↓ Return Temp↓ Room Temp↓ 휆 pred 00.0590.10513.6719.2379.560 0.2 (Ours) 0.049 0.08610.5136.5316.847 0.50.0520.09311.3687.0296.988 0.80.0560.10112.5016.3148.321 1.00.0620.10613.4756.2989.313 휆 푑 00.0610.10813.7836.6709.557 0.1 (Ours) 0.049 0.08610.5136.5316.847 0.30.0510.08810.8686.1436.935 0.50.0540.09611.6226.1937.177 Answer for RQ3: To answer RQ3, we evaluate the downstream control performance of the IEWMs introduced in RQ2. To isolate the contribution of world-model design, all compared methods use the same RL controller described in Section 3.4, while only the underlying IEWM is varied. Figure 5 shows that under seasonal transfer, ADAPTconsistently achieves the best trade-off between energy efficiency, occupant comfort, and control smoothness. The same trend is observed un- der climate-region transfer (Fig. 6). Across both transfer directions between Stockholm and Arizona, ADAPT consistently delivers higher energy savings with lower temperature violation than se- quential and purely data-driven world models. The consistent ad- vantage across different transfer scenarios demonstrates that im- proved world-model transferability leads to more robust and reliable downstream RL control. Additional results are shown in C.2. The transferability originates from the proposed physics-aware regularization. By modeling climate-invariant thermal dynamics in- stead of environment-specific statistical correlations, ADAPT main- tains accurate thermal prediction under OOD scenarios. Conse- quently, the downstream controller receives reliable future thermal baseline even in unseen environments, leading to more robust pol- icy optimization and consistently stronger transfer performance. Ablation on Physics-Loss Weights. Tables 3 investigates the influence of the two physics-aware regularization weights under the SemiBuildingSim Winter→Summer transfer. Removing either term consistently deteriorates OOD prediction performance, demonstrat- ing that both physical consistency and residual-bias regulariza- tion are essential for learning transferable thermal dynamics. Con- versely, excessively large weights also reduce prediction accuracy, indicating that overly strong physical constraints may compromise the expressive capacity of the diffusion model. We therefore use 휆 pred = 0.2 and 휆 푑 = 0.1 throughout the paper. 5 Discussion As global urbanization and climate change reshape building energy demand, effective modeling and control of indoor thermal dynam- ics is important for building decarbonization, occupant comfort, and peak-load management, advancing the UN’s SDGs 11 and 13. Our key contribution lies in introducing physics-aware generative indoor environmental world models as a transferable predictive foundation for HVAC control. Instead of learning a reactive con- troller tied to one season and one climate region, our framework utilizes world model to predict a thermal baseline that transfers across seasons and climate zones, supporting energy-efficient build- ing operation, comfort improvement, and cross-region deployment for energy system and building science . A central challenge in HVAC control is the combination of de- layed thermal dynamics, partial observability, and substantial ther- mal distribution shifts across seasons, weather conditions, and climate regions. Since the effect of an HVAC action emerges grad- ually, effective control requires a world model that explicitly pre- dicts future indoor thermal evolution for action evaluation and delayed credit assignment. We address this by introducing a physics- aware generative world model to guide RL learning. Besides, unlike existing deterministic and vanilla data-driven world models that primarily learn specific statistical correlations, our physics-aware diffusion IEWM learns transferable thermal dynamics through gen- erative trajectory modeling regularized by a multi-zone physical heat-balance equation. Consequently, the proposed IEWM learns from one operating condition and is transferable across seasons, weather conditions, and climate regions. In summary, our framework can offer building operators, HVAC engineers, and energy scientists a powerful generative tool for predicting delayed indoor thermal responses, transferring HVAC control across seasons and climate regions, and jointly improving energy consumption and occupant comfort across offices, residen- tial buildings, and high-density AI infrastructure such as data-center cooling. This work established a decarbonized, physically consis- tent, and transferable HVAC control process aligned with SDGs 11 and 13, advancing the development of AI for building science, energy systems, and climate mitigation. 5.1 Limitations and Ethical Considerations While the proposed physics-aware regularization significantly im- proves the transferability of the IEWM, it currently constrains only indoor temperature dynamics through the multi-zone heat-balance equation. More complex processes, such as humidity, are not explic- itly modeled. Incorporating richer thermo-hygrometric dynamics into the world model is an important direction for future work. All experiments are conducted using publicly available build- ing simulators and weather data. No private, household-level, or personally identifiable information is collected, used, or inferred during model development or evaluation. 6 Conclusion This paper presents ADAPT, a physics-aware diffusion IEWM for robust HVAC control under seasonal and climate-region transfer. 7 Conference’27, August 2027, San Jose, California, USAYang et al. MAERMSEE-HumW-HumE-ZoneTW-ZoneT 0 0.5 1 Normalized error (1 = worst in metric) 0.0020.002 0.011 0.012 0.0200.020 0.029 0.028 2.39 2.72 5.97 6.32 1.94 2.11 4.39 4.67 0.019 0.017 0.074 0.063 0.512 0.588 0.591 0.620 Stockholm → Arizona — IID MAERMSEE-HumW-HumE-ZoneTW-ZoneT 0 0.5 1 0.120 0.169 0.312 0.290 0.311 0.412 0.605 0.483 14.6 18.0 41.8 37.1 20.6 23.3 30.3 36.9 2.04 3.05 3.66 2.86 1.63 3.43 5.34 6.60 Stockholm → Arizona — OOD MAERMSEE-HumW-HumE-ZoneTW-ZoneT 0 0.5 1 Normalized error (1 = worst in metric) 0.0030.003 0.016 0.015 0.022 0.021 0.037 0.034 2.18 2.20 6.83 6.41 3.19 3.26 7.56 7.10 0.024 0.021 0.081 0.064 0.567 0.200 0.569 0.546 Arizona → Stockholm — IID MAERMSEE-HumW-HumE-ZoneTW-ZoneT 0 0.5 1 0.067 0.074 0.191 0.224 0.174 0.226 0.387 0.354 16.5 22.1 30.1 45.6 16.0 20.0 25.9 42.5 1.30 1.89 2.90 2.22 0.708 1.55 3.79 6.86 Arizona → Stockholm — OOD IID — source regionOOD — target region (cross-region transfer) ADAPT(Ours)w/o Phys.MambaFormerVAEbest in metric Figure 4: Cross-region transfer prediction performance on Sinergym 2ZoneDataCenter (Stockholm↔Arizona). CVRMSE are annotated above each bar. Red outlines indicate the best value per metric (lower is better). ADAPT (Ours) w/o Phys. Mamba WM VAE WM 140 150 160 Energy (kWh) ADAPT (Ours) w/o Phys. Mamba WM VAE WM 0.3 0.4 0.5 Abs PMV ADAPT (Ours) w/o Phys. Mamba WM VAE WM 5 10 15 PPD (%) ADAPT (Ours) w/o Phys. Mamba WM VAE WM 5.0 7.5 10.0 Action Fluctuation ADAPT (Ours)w/o Phys.Mamba WMVAE WM Figure 5: Comparison of different IEWMs downstream control performance on SemibuildingSim Winter→Summer transfer. Lower is better. 678910 Energy Saving (%)↑ 0.4 0.6 0.8 1.0 Avg. Temperature Violation ( ∘ C) ↓ ADAPT(Ours) better Stockholm → Arizona 81012 Energy Saving (%)↑ 0.0 0.2 0.4 0.6 0.8 Avg. Temperature Violation ( ∘ C) ↓ ADAPT(Ours) better Arizona → Stockholm ADAPT (Ours)w/o Phys.Mamba WMVAE WMIIDOOD (transfer) Figure 6: Comparison of different IEWMs downstream con- trol performance on Sinergym cross climate-region transfer. Experiments demonstrate that ADAPT consistently improves world- model prediction, downstream control performance, and transfer robustness over existing baselines. By integrating physics-aware thermal dynamics into generative world modeling, ADAPT provides a practical and transferable foundation for energy-efficient building control. We hope this work promotes further research on physics- guided world models for AI-driven building energy management, contributing to more sustainable and intelligent buildings. 7 GenAI Disclosure The authors declare that AI tools were used only for language polishing and grammar checking. All scientific contributions, in- cluding ideas, experiments, and analyses, are the authors’ own. All intellectual content, data analysis, interpretations, and conclusions were conceived, written, and verified by the authors. 8 ADAPT: A Diffusion-based Adaptive Physics-aware Indoor Environmental World Model for Transferable HVAC ControlConference’27, August 2027, San Jose, California, USA References [1]Abdul Afram and Farrokh Janabi-Sharifi. 2014. Theory and applications of HVAC control systems–A review of model predictive control (MPC). Building and environment 72 (2014), 343–355. [2]Khalil Al Sayed, Abhinandana Boodi, Roozbeh Sadeghian Broujeny, and Karim Beddiar. 2024. Reinforcement learning for HVAC control in intelligent buildings: A technical and conceptual review. Journal of Building Engineering 95 (2024), 110085. [3]American Society of Heating, Refrigerating and Air-Conditioning En- gineers. 2023.ANSI/ASHRAE Standard 55-2023: Thermal Environ- mental Conditions for Human Occupancy.ASHRAE, Atlanta, GA. https://w.ashrae.org/technical-resources/bookstore/standard-55-thermal- environmental-conditions-for-human-occupancy [4]Yu Qian Ang, Zachary Michael Berzolla, Samuel Letellier-Duchesne, and Christoph F Reinhart. 2023. Carbon reduction technology pathways for existing buildings in eight cities. Nature communications 14, 1 (2023), 1689. [5] Hassam Ayaz, Mohammed Faizal, and Abdelmalek Bouazza. 2024. Energy, eco- nomic, and carbon emission analysis of a residential building with an energy pile system. Renewable Energy 220 (2024), 119712. [6]Adrien Bardes, Quentin Garrido, Jean Ponce, Xinlei Chen, Michael Rabbat, Yann LeCun, Mido Assran, and Nicolas Ballas. [n. d.]. Revisiting Feature Prediction for Learning Visual Representations from Video. Transactions on Machine Learning Research ([n. d.]). [7]Chenhang Bian, Ka Lung Cheung, Xi Chen, and Chi Chung Lee. 2025. Integrating microclimate modelling with building energy simulation and solar photovoltaic potential estimation: The parametric analysis and optimization of urban design. Applied Energy 380 (2025), 125062. [8]John Bongaarts. 2020. United Nations Department of Economic and Social Affairs, Population Division World Family Planning 2020: Highlights, United Nations Publications, 2020. 46 p. Popul Dev Rev 46, 4 (2020), 857–858. [9]Alejandro Campoy-Nieves, Antonio Manjavacas, Javier Jiménez-Raboso, Miguel Molina-Solana, and Juan Gómez-Romero. 2025. Sinergym–A virtual testbed for building energy optimization with Reinforcement Learning. Energy and Buildings 327 (2025), 115075. [10]Cheng Chi, Zhenjia Xu, Siyuan Feng, Eric Cousineau, Yilun Du, Benjamin Burch- fiel, Russ Tedrake, and Shuran Song. 2025. Diffusion policy: Visuomotor policy learning via action diffusion. The International Journal of Robotics Research 44, 10-11 (2025), 1684–1704. [11] Conference of the Parties serving as the meeting of the Parties to the Paris Agreement (CMA). 2023. Outcome of the First Global Stocktake. UNFCCC Decision 1/CMA.5. FCCC/PA/CMA/2023/L.17. [12]Davide Coraci, Silvio Brandi, Tianzhen Hong, and Alfonso Capozzoli. 2023. Online transfer learning strategy for enhancing the scalability and deployment of deep reinforcement learning control in smart buildings. Applied Energy 333 (2023), 120598. [13]Jingtao Ding, Yunke Zhang, Yu Shang, Yuheng Zhang, Zefang Zong, Jie Feng, Yuan Yuan, Hongyuan Su, Nian Li, Nicholas Sukiennik, et al.2025. Understanding world or predicting future? a comprehensive survey of world models. Comput. Surveys 58, 3 (2025), 1–38. [14] Ján Drgoňa, Javier Arroyo, Iago Cupeiro Figueroa, David Blum, Krzysztof Arendt, Donghun Kim, Enric Perarnau Ollé, Juraj Oravec, Michael Wetter, Draguna L Vrabie, et al.2020. All you need to know about model predictive control for buildings. Annual reviews in control 50 (2020), 190–232. [15] Poul O Fanger. 1970. Thermal comfort. Analysis and applications in environmen- tal engineering. (1970). [16]Vladimir Feinberg, Alvin Wan, Ion Stoica, Michael I Jordan, Joseph E Gonzalez, and Sergey Levine. 2018. Model-based value estimation for efficient model-free reinforcement learning. arXiv preprint arXiv:1803.00101 (2018). [17]Hongfan Gao, Wangmeng Shen, Xiangfei Qiu, Ronghui Xu, Bin Yang, and Jilin Hu. 2025. SSD-TS: Exploring the potential of linear state space models for diffusion models in time series imputation. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2. 649–660. [18] Ignat Georgiev, Varun Giridhar, Nicklas Hansen, and Animesh Garg. [n. d.]. PWM: Policy Learning with Multi-Task World Models. In The Thirteenth International Conference on Learning Representations. [19]Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap. 2023. Mas- tering diverse domains through world models. arXiv preprint arXiv:2301.04104 (2023). [20]Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap. 2025. Mas- tering diverse control tasks through world models. Nature 640, 8059 (2025), 647–653. [21]Ahmed M Hanafi, Mohamed Ahmed Moawed, and Osama Ezzat Abdellatif. 2024. Advancing sustainable energy management: a comprehensive review of artificial intelligence techniques in building. Engineering Research Journal (Shoubra) 53, 2 (2024), 26–46. [22]Nicklas Hansen, Hao Su, and Xiaolong Wang. [n. d.]. TD-MPC2: Scalable, Robust World Models for Continuous Control. In The Twelfth International Conference on Learning Representations. [23]N Hansen, X Wang, and H Su. 2022. Temporal Difference Learning for Model Predictive Control. In International Conference on Machine Learning, PMLR. [24] Thi Ngoc Yen Huynh, Anh Tuan Nguyen, Yonghan Ahn, Bee Lan Oo, and Ben- son TH Lim. 2025. Multi objectives reinforcement learning for smart buildings: A systematic review of algorithms, applications and future perspectives. Energy and Buildings 345 (2025), 116045. [25]Michael Janner, Yilun Du, Joshua Tenenbaum, and Sergey Levine. 2022. Planning with Diffusion for Flexible Behavior Synthesis. In International Conference on Machine Learning. PMLR, 9902–9915. [26]Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine. 2019. When to trust your model: Model-based policy optimization. In Advances in neural information processing systems, Vol. 32. [27]Michaela Killian and Martin Kozek. 2016. Ten questions concerning model predictive control for energy efficient buildings. Building and Environment 105 (2016), 403–412. [28]Hoesung Lee, Katherine Calvin, Dipak Dasgupta, Gerhard Krinner, Aditi Mukherji, Peter Thorne, Christopher Trisos, José Romero, Paulina Aldunce, Ko Barrett, et al.2023. Climate change 2023: synthesis report. Contribution of working groups I, I and I to the sixth assessment report of the intergovernmental panel on climate change. [29] Rui Liang, Yang Deng, and Dan Wang. 2024. Probabilistic Building Load Forecast- ing via Conditional Diffusion Model. In Proceedings of the 15th ACM International Conference on Future and Sustainable Energy Systems. 490–491. [30]Yuansan Liu, Sudanthi Wijewickrema, Dongting Hu, Christofer Bester, Stephen O’Leary, and James Bailey. 2025. Stochastic diffusion: A diffusion based model for stochastic time series forecasting. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2. 1939–1950. [31]Jie Lu, Chaobo Zhang, Bozheng Li, Yang Zhao, Ruchi Choudhary, and Max Langtry. 2025. Self-attention variational autoencoder-based method for incom- plete model parameter imputation of digital twin building energy systems. Energy and Buildings 328 (2025), 115162. [32]Antonio Manjavacas, Alejandro Campoy-Nieves, Javier Jiménez-Raboso, Miguel Molina-Solana, and Juan Gómez-Romero. 2024. An experimental evaluation of Deep Reinforcement Learning algorithms for HVAC control. arXiv e-prints (2024), arXiv–2401. [33] Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Tim- othy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. 2016. Asyn- chronous methods for deep reinforcement learning. In International conference on machine learning. PmLR, 1928–1937. [34] Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al.2015. Human-level control through deep reinforcement learning. nature 518, 7540 (2015), 529–533. [35]Zoltan Nagy, Gregor Henze, Sourav Dey, Javier Arroyo, Lieve Helsen, Xiangyu Zhang, Bingqing Chen, Kadir Amasyali, Kuldeep Kurte, Ahmed Zamzam, et al. 2023. Ten questions concerning reinforcement learning for building energy management. Building and Environment 241 (2023), 110435. [36] Tianwei Ni, Michel Ma, Benjamin Eysenbach, and Pierre-Luc Bacon. 2023. When do transformers shine in rl? decoupling memory from credit assignment. Ad- vances in Neural Information Processing Systems 36 (2023), 50429–50452. [37]Alexander Quinn Nichol and Prafulla Dhariwal. 2021. Improved denoising diffu- sion probabilistic models. In International conference on machine learning. PMLR, 8162–8171. [38]Kingsley Nweye, Bo Liu, Peter Stone, and Zoltan Nagy. 2022. Real-world chal- lenges for multi-agent reinforcement learning in grid-interactive buildings. En- ergy and AI 10 (2022), 100202. [39]Luis Pérez-Lombard, José Ortiz, and Christine Pout. 2008. A review on buildings energy consumption information. Energy and buildings 40, 3 (2008), 394–398. [40]UN Environment Programme. 2022. "2022 GLOBAL STATUS REPORT FOR BUILDINGS AND CONSTRUCTION". https://globalabc.org. [41]Aditya Ramesh, Kenny John Young, Louis Kirsch, and Jürgen Schmidhuber. [n. d.]. Sequence Compression Speeds Up Credit Assignment in Reinforcement Learning. In Forty-first International Conference on Machine Learning. [42]Kashif Rasul, Calvin Seward, Ingmar Schuster, and Roland Vollgraf. 2021. Au- toregressive denoising diffusion models for multivariate probabilistic time series forecasting. In International conference on machine learning. PMLR, 8857–8868. [43]John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347 (2017). [44]Alberto Silvestri, Davide Coraci, Silvio Brandi, Alfonso Capozzoli, and Arno Schlueter. 2025. Practical deployment of reinforcement learning for building controls using an imitation learning approach. Energy and Buildings 335 (2025), 115511. [45]Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Ste- fano Ermon, and Ben Poole. 2021. Score-Based Generative Modeling through Stochastic Differential Equations. In International Conference on Learning Repre- sentations. 9 Conference’27, August 2027, San Jose, California, USAYang et al. [46]Iain Staffell, Stefan Pfenninger, and Nathan Johnson. 2023. A global model of hourly space heating and cooling demand at multiple spatial scales. Nature Energy 8, 12 (2023), 1328–1344. [47] Kailai Sun, Irfan Qaisar, Muhammad Arslan Khan, Tian Xing, and Qianchuan Zhao. 2023. Building occupancy number prediction: A Transformer approach. Building and environment 244 (2023), 110807. [48] Kailai Sun, Qianchuan Zhao, and Jianhong Zou. 2020. A review of building occupancy measurement systems. Energy and Buildings 216 (2020), 109965. [49] Arash Tavakoli, Fabio Pardo, and Petar Kormushev. 2018. Action branching architectures for deep reinforcement learning. In Proceedings of the aaai conference on artificial intelligence, Vol. 32. [50]United Nations Environment Programme and Global Alliance for Build- ings and Construction. 2025. Global Status Report for Buildings and Construction 2024/2025.Technical Report. United Nations Environment Programme.https://w.unep.org/resources/report/global-status-report- buildings-and-construction-20242025 [51]Rik van Heerden, Oreane Y Edelenbosch, Vassilis Daioglou, Thomas Le Gallic, Luiz Bernardo Baptista, Alice Di Bella, Francesco Pietro Colelli, Johannes Em- merling, Panagiotis Fragkos, Robin Hasse, et al.2025. Demand-side strategies enable rapid and deep cuts in buildings and transport emissions to 2050. Nature Energy 10, 3 (2025), 380–394. [52] José R Vázquez-Canteli and Zoltán Nagy. 2019. Reinforcement learning for demand response: A review of algorithms and modeling techniques. Applied energy 235 (2019), 1072–1089. [53]Stijn Verbeke and Amaryllis Audenaert. 2018. Thermal inertia in buildings: A review of impacts across climate and building use. Renewable and sustainable energy reviews 82 (2018), 2300–2318. [54] Jingwei Wang, Qianyue Hao, Wenzhen Huang, Xiaochen Fan, Zhentao Tang, Bin Wang, Jianye Hao, and Yong Li. 2024. Dyps: Dynamic parameter sharing in multi-agent reinforcement learning for spatio-temporal resource allocation. In Proceedings of the 30th ACM SIGKDD conference on knowledge discovery and data mining. 3128–3139. [55]Liangxu Wang and Tianzhen Hong. 2020. Reinforcement learning for building controls: The opportunities and challenges. Applied Energy 269 (2020), 115036. doi:10.1016/j.apenergy.2020.115036 [56]Xiangwei Wang, Peng Wang, Renke Huang, Xiuli Zhu, Javier Arroyo, and Ning Li. 2025. Safe deep reinforcement learning for building energy management. Applied Energy 377 (2025), 124328. [57] Zhixian Wang, Qingsong Wen, Chaoli Zhang, Liang Sun, and Yi Wang. 2024. Dif- fLoad: Uncertainty quantification in electrical load forecasting with the diffusion model. IEEE Transactions on Power Systems 40, 2 (2024), 1777–1789. [58]Jialong Wu, Shaofeng Yin, Ningya Feng, and Mingsheng Long. [n. d.]. RLVR- World: Training World Models with Reinforcement Learning. In The Thirty-ninth Annual Conference on Neural Information Processing Systems. [59]Tian Xing, Hu Yan, Kailai Sun, Yifan Wang, Xuetao Wang, and Qianchuan Zhao. 2022. Honeycomb: An open-source distributed system for smart buildings. Pat- terns 3, 11 (2022). [60]Xiongxiao Xu, Yueqing Liang, Baixiang Huang, Zhiling Lan, and Kai Shu. 2024. Integrating mamba and transformer for long-short range time series forecasting. arXiv preprint arXiv:2404.14757 2, 7 (2024). [61] Yiqin Yang, Xu Yang, Yuhua Jiang, Ni Mu, Hao Hu, Runpeng Xie, Ziyou Zhang, Siyuan Li, Yuan-Hua Ni, Qianchuan Zhao, et al.2026. GlobeDiff: State Diffu- sion Process for Partial Observability in Multi-Agent Systems. arXiv preprint arXiv:2602.15776 (2026). [62]Xinbo Zhao, Yingxue Zhang, Xin Zhang, Yu Yang, Yiqun Xie, Yanhua Li, and Jun Luo. 2024. Urban-focused multi-task offline reinforcement learning with contrastive data sharing. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 4512–4523. [63]Yi Zhao, Aidan Scannell, Yuxin Hou, Tianyu Cui, Le Chen, Dieter Büchler, Arno Solin, Juho Kannala, and Joni Pajarinen. 2025. Generalist World Model Pre- Training for Efficient Reinforcement Learning. In ICLR 2025 Workshop on World Models: Understanding, Modelling and Scaling. [64]Zhida Zhao, Talas Fu, Yifan Wang, Lijun Wang, and Huchuan Lu. 2026. From fore- casting to planning: Policy world model for collaborative state-action prediction. Advances in Neural Information Processing Systems 38 (2026), 134585–134611. [65]Dianyu Zhong, Tian Xing, Kailai Sun, Ziyou Zhang, Qianchuan Zhao, and Jian Kang. 2025. Topology-aware hypergraph reinforcement learning for indoor occupant-centric HVAC control. Energy and Buildings (2025), 116219. [66] Mengting Zhu, Mengqi Zhao, Rongqi Zhu, Jiyong Eom, Yuyu Zhou, Sha Yu, Fengqiao Mei, and Yang Ou. 2026. Warming-driven shifts in global building energy use reshape climate mitigation planning. Nature Communications (2026). [67] Xingchen Zou, Weilin Ruan, Siru Zhong, Yuehong Hu, and Yuxuan Liang. 2025. Fine-grained Urban Heat Island Effect Forecasting: A Context-aware Thermody- namic Modeling Framework. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2. 4226–4237. 10 ADAPT: A Diffusion-based Adaptive Physics-aware Indoor Environmental World Model for Transferable HVAC ControlConference’27, August 2027, San Jose, California, USA A Training and Inference Algorithm Algorithm 1 summarizes the two-stage training and online-use procedure of ADAPT. Phase I trains the physics-aware diffusion IEWM by jointly minimizing the diffusion denoising loss (Eq. 8), the true-trajectory thermal residual (Eq. 17), the predicted-trajectory thermal residual (Eq. 19), and the residual-bias regularizer (Eq. 20). Phase I freezes the world model and runs baseline-conditioned RL: at each step, the frozen IEWM produces (i) the held-action thermal baseline ˆ y 푡 that conditions the branch-wise action-value network (Eq. 21) and drives greedy action selection (Eq.??), and (i) the action-committed rollout ˆ y ′ 푡 that supplies delayed states for the held-action TD-휆target (Eqs. 23–24). The policy parametersΦ are updated by minimizingL RL (Eq. 25), followed by a soft update of the target network. Algorithm 1 Training and online use of ADAPT Require:Offline trajectoriesD off , RL replay bufferD rl , horizon 퐻 , diffusion steps 퐾 , reward evaluator 푔, target update rate 휅 Phase I: Physics-aware diffusion IEWM training 1: for each IEWM update do 2:Sample(표 푡−푁 :푡+퐻 ,푎 푡−푁 :푡−1 ) ∼D off 3:Set ̃ 푎 푡+ℎ ← 푎 푡−1 for ℎ= 0, . . .,퐻 − 1 and build c 푡 by Eq. 2 4:Set the clean target y 0 푡 ← [표 푡+1 , . . .,표 푡+퐻 ] 5:Sample 푘 ∼ Uniform(1, . . .,퐾) and 휖 ∼N(0, I) 6:Construct y 푘 푡 ← √ ̄ 훼 푘 y 0 푡 + √ 1− ̄ 훼 푘 휖 by Eq. 4 7:Predict ˆ 휖 휃 (y 푘 푡 ,푘, c 푡 ) and computeL diff by Eq. 8 8:Generate ˆ y 푡 and extract ˆ T 푡+1:푡+퐻 9:ComputeL true andL pred by Eqs. 17 and 19 10:Update 휃 and thermal coefficients by minimizing Eq. 20 11: end for 12: Freeze 휃 Phase I: Online RL Policy Learning 13: for each environment step 푡 do 14:Generate ˆ y 푡 from Eq. 9 15: for each branch 푑= 1, . . .,퐷 do 16:Select 푎 푡,푑 with 푄 푑 (표 푡 ,푎, ˆ y 푡 ;Φ) from Eqs. 21 17: end for 18:Execute 푎 푡 ; observe 표 푡+1 and 푟 푡 19:Generate committed rollout ˆ y ′ 푡 using Eq. 22 20:Store(표 푡−퐿:푡 ,푎 푡−퐿:푡 ,푟 푡 ,표 푡+1 , ˆ y 푡 , ˆ y ′ 푡 ) inD rl 21: if it is time to updateΦ then 22:Sample a minibatch fromD rl 23:Compute ˆ 푟 (푎 푡 ) 푡+푘 for 푘= 1, . . .,퐻 − 1 24:Compute 퐺 (ℎ) 푡 (푎 푡 ) and푦 휆 푡 by Eqs. 23 and 24 25:UpdateΦ by minimizingL RL (Φ) in Eq. 25 26:Soft-update target parameters ̄ Φ← 휅Φ+(1−휅) ̄ Φ 27: end if 28: end for 11 Conference’27, August 2027, San Jose, California, USAYang et al. B Experimental and Environmental Details B.1 SemiBuildingSim Environment B.1.1 Environment Construction and Case Study Modeling. To com- prehensively evaluate the proposed control strategies, we utilize SemiBuildingSim, a high-fidelity simulator designed to replicate the complex thermal and energy dynamics of a real-world building. The simulation environment is grounded in historical data collected from a seven-zone office building in Hebei Province, China [59]. The building operates primarily during standard working hours (9:00- 19:00 on weekdays), with reduced occupancy around lunchtime. An average of 12.4 people occupied the office (August 7 to August 22), with typical weekday usage and no regular weekend activity. As shown in Fig. 7, the building comprises several individual rooms and a hall. One fan coil unit (FCU) is installed in each thermal zone. FCUs 1–4 each serve an individual cellular office, whereas FCUs 5–7 serve zones that allow inter-zone heat exchange. The FCUs receive chilled or heated water from a central refrigeration station comprising a heat pump and a circulating water pump. The pump distributes water to all FCUs through a closed-loop pipe net- work, while the heat pump provides the thermal source for cooling and heating. An overview of the HVAC system configuration and representative devices is shown in Fig. 8. Figure 7: Layout of the office building. Source from [65]. The simulator comprises three rigorously coupled components: indoor zone models, parameterized HVAC equipment models, and a dynamic cooling water pipe network [59]. The indoor thermal dynamics are captured using a lumped capacitance approach that accounts for building envelope conduction (via the OTTV method), solar radiation, instantaneous occupancy loads, inter-zone heat transfer, and the sensible cooling/heating capacities delivered by the terminal HVAC units. The hydraulic pipe network utilizes a directed graph topology, solving the complete set of nonlinear hydraulic equations at each time step using the Newton-Raphson method to ensure mass continuity and system pressure-flow balance. To support simulator calibration and monitor indoor thermal conditions, a comprehensive sensing infrastructure is deployed. At the zone level, each thermal zone is equipped with indoor air temperature sensors (Fig. 8(f)), and occupancy is inferred from ceiling-mounted video cameras (Fig. 8(c, g)). At the terminal side, each FCU is instrumented with a water flow meter, supply and return water temperature sensors, and pressure sensors (Fig. 8(b)). In the central plant, the variable frequency drive (VFD) circulating water pumps (1 Hz control resolution) are equipped with pressure and temperature sensors (Fig. 8(d)), and their electrical power is measured via electrical meters (Fig. 8(e)). All sensors operated con- tinuously during HVAC runtime. B.1.2 Observation and Action Spaces. The environment operates as a partially observable Markov decision process (POMDP), where the agent receives states and issues control commands at a fixed discrete timestep ofΔ푡= 5 minutes. Observation Space: The state vector s 푡 captures the current thermodynamic and operational context of the building. This in- cludes the indoor air temperatures of the 7 zones푇 in,푖 , the instanta- neous occupant counts푛 p,푖 , outdoor weather conditions, and the current operational states of the HVAC system components. Action Space: The control variables actuate the terminal distri- bution systems across the building. The action space is formulated as a MultiDiscrete space, allowing for the independent control of the 7 FCUs installed in the respective thermal zones. For each FCU, the controller selects from 4 discrete fan operational modes (e.g., off, low, medium, high). Consequently, the joint action at time step 푡 is defined as a 푡 ∈ 0, 1, 2, 3 7 . B.1.3 Performance Metrics and Reward Function. To comprehen- sively assess the proposed building control strategies, we adopt a dual-objective evaluation framework that quantifies both occu- pant thermal comfort and HVAC energy use. These metrics define the reinforcement learning (RL) reward signal and are used for continuous controller performance evaluation. Thermal Discomfort and Energy Metrics. Thermal comfort is eval- uated using the Predicted Mean Vote (PMV) and Predicted Per- centage of Dissatisfied (PPD) indices [15]. For the specific office conditions, the calculations assume an air velocity of푣=0.15 m/s, relative humidity of RH=40%, clothing insulation of퐼 푐푙 =0.63 clo, and a metabolic rate of푀=1.1 met. Due to the well-insulated en- velope, the mean radiant temperature푇 푟 is approximated to equal the zone air temperature푇 푎 . A PMV value of 0 indicates neutral thermal sensation, with an acceptable range of[−0.5,+0.5]corresponding to a PPD below 10% [3]. The system-level discomfort at time step푡is computed as the occupancy-weighted mean PPD across all active zones: PPD mean,푡 = 1 푁 푡 Í 퐽 푖=1 푛 푖,푡 · PPD 푖,푡 , 푁 푡 > 0, 0,푁 푡 = 0, (26) where퐽=7 is the number of zones,푛 푖,푡 is the occupancy count in zone 푖, and 푁 푡 = Í 퐽 푖=1 푛 푖,푡 is the total number of occupants. Energy consumption is measured as the total HVAC electrical power푃 푡 (in kWh) at each time step푡, returned by the simulator. Cumulative energy over an episode is computed as 퐸= Í 푡 푃 푡 ·Δ푡 . Composite Reward Function. The objective of the RL agent is to maximize the expected cumulative reward. At each time step푡, the composite reward푟 푡 is formulated to balance thermal comfort, energy conservation, and system operational stability: 푟 푡 =−푟 comfort − 훼 · 푟 energy − 훽· 푅 smooth,푡 (27) 12 ADAPT: A Diffusion-based Adaptive Physics-aware Indoor Environmental World Model for Transferable HVAC ControlConference’27, August 2027, San Jose, California, USA Figure 8: HVAC system configuration. (a) FCUs of zones 5 to 7. (b) Sensors of FCU 1. (c) Video sampling of zone 7. (d) Sensors of water pumps. (e) Electrical cabinet. (f) Temperature sensor. (g) Video camera of zone 2. Source from [65]. where푟 comfort penalizes thermal discomfort based on the calculated PPD mean,푡 , and푟 energy penalizes the total electrical power consump- tion 푃 푡 . To optimize control performance and ensure physically plau- sible control, we introduce a smoothness penalty푅 smooth,푡 . This term explicitly discourages high-frequency oscillation and abrupt switching between FCU modes, thereby reducing mechanical wear and extending equipment longevity. The weighting coefficients are empirically established: the energy penalty coefficient is set to훼=20. For the smoothness regularization, the coefficient is 훽=5 when employing a Log smoothness reward, and훽=3 when utilizing a linear smoothness reward. B.2 Sinergym 2ZoneDataCenter Environment B.2.1 Environment Construction. To benchmark our proposed con- trol strategy against baselines, we utilize the2ZoneDataCenter environment provided by the open-source Sinergym framework [? ]. Sinergym serves as a Python-based virtual testbed that integrates deep reinforcement learning algorithms with the high-fidelity En- ergyPlus building simulation engine. The selected building model represents a single-story data center with a total surface area of 491.3m 2 . The facility is divided into two asymmetrical thermal zones: the West Zone and the East Zone. Unlike typical commercial buildings, the primary internal heat load in this environment is generated continuously by the hosted server racks (ITE objects), and the building has no windows or internal mass. To manage this intense thermal output, each zone is equipped with an HVAC system consisting of an air economizer, direct and indirect evaporative coolers, a single-speed direct expansion (DX) cooling coil, a chilled water coil, and a Variable Air Volume (VAV) system without reheat air terminal units. B.2.2 Observation and Action Spaces. The environment is formu- lated as a discrete-time Markov Decision Process (MDP), where the agent interacts with the EnergyPlus backend via Sinergym’s Gymnasium interface. Observation Space: The state vector s 푡 encompasses both ex- ogenous disturbances and the internal thermodynamic state of the data center. Key observations include outdoor weather variables (e.g., outdoor air temperature, relative humidity, solar radiation), the current indoor air temperatures of the West and East zones, and the instantaneous power consumption of the HVAC components and IT equipment. Action Space: The control variables manipulate the thermal setpoints of the HVAC system to regulate the cooling capacity delivered to the servers. The action space dictates the continuous cooling and heating temperature setpoints for both the West and East zones. By dynamically adjusting these setpoints, the controller coordinates the mechanical cooling and economizers. B.2.3 Reward Function and Comfort Constraints. In data center operations, maintaining a strict thermal environment is critical for equipment reliability, while minimizing energy consumption is essential for operational sustainability. Following Sinergym’s official linear reward architecture, we define a dual-objective reward function that balances energy efficiency and thermal safety. For our specific case study, the operational comfort tempera- ture range for the server zones is strictly defined as[푇 low ,푇 high ]= 13 Conference’27, August 2027, San Jose, California, USAYang et al. [20 ◦ C,26 ◦ C]. The temperature deviation penaltyΔ푇 푖,푡 at time step 푡is calculated as the magnitude of the violation from this safe operating band for each zone: Δ푇 푖,푡 = max(0,푇 푖,푡 − 26)+ max(0, 20−푇 푖,푡 ),(28) where푇 푖,푡 is the indoor air temperature of zone 푖 ∈ West, East. The step reward푟 푡 computes a weighted sum of the energy penalty and the temperature violation penalty. To ensure both ob- jectives are scaled appropriately, the terms are normalized relative to their maximum operational limits: 푟 푡 =−훼 · 휔 · 퐸 푡 퐸 max − 훽·(1− 휔)· Í 푖 Δ푇 푖,푡 푇 max ,(29) where퐸 푡 is the total electrical power consumption of the HVAC sys- tem at time step푡, and퐸 max and푇 max are normalization constants established by the environment’s empirical bounds. The weighting coefficient휔 ∈ [0,1]dictates the trade-off between energy con- servation and thermal compliance. By imposing a hard penalty on temperatures outside the 20 ◦ Cto 26 ◦ Cthreshold, the agent is incen- tivized to leverage free cooling dynamically without jeopardizing server integrity. We set 휔= 0.5, 훼= 1, 훽= 5× 10 −3 . 14 ADAPT: A Diffusion-based Adaptive Physics-aware Indoor Environmental World Model for Transferable HVAC ControlConference’27, August 2027, San Jose, California, USA C Additional Experiments C.1 Sinergym ID Control Comparison We provide additional ID control results on two representative scenarios: the Winter setting of SemiBuildingSim and the Arizona setting of Sinergym. Tab.??and Fig. 9 show that ADAPT consis- tently achieves higher energy savings while maintaining occupant comfort within the desired temperature range. These results further demonstrate that the proposed physics-aware diffusion IEWM accu- rately captures building thermal dynamics under different seasons and climate conditions, providing reliable thermal baselines that consistently improve downstream HVAC control. Table 4: Performance comparisons on SemiBuildingSim Sum- mer (mean± std over three seeds, ID). Lower is better. AlgorithmEnergy (kWh)↓ Abs PMV↓ PPD (%)↓ Action Fluctuation↓ MPC251.00± 9.79 0.56± 0.06 16.20± 1.089.94± 1.13 A2C247.28± 3.21 0.48± 0.03 13.05± 1.466.02± 1.03 PPO236.52± 2.90 0.46± 0.04 12.62± 0.754.94± 1.46 DQN248.41± 3.89 0.50± 0.05 13.02± 0.776.59± 0.55 BDQ241.87± 2.73 0.45± 0.05 11.40± 1.285.40± 0.46 TransformerRL 233.97± 4.36 0.42± 0.03 11.72± 1.624.66± 0.65 MBVE236.25± 3.03 0.42± 0.02 11.10± 0.185.08± 0.71 MBPO234.66± 2.49 0.41± 0.01 11.34± 0.124.23± 0.77 DreamerV3228.84± 2.43 0.40± 0.02 10.65± 0.355.29± 0.88 ADAPT (Ours) 219.48± 1.50 0.28± 0.01 7.56± 0.053.62± 0.48 Jan Feb Mar Apr May Jun Jul Aug Sep Oct Nov Dec -4% 4% 12% 20% 28% (a) Energy Saving vs RBC Jan Feb Mar Apr May Jun Jul Aug Sep Oct Nov Dec -0.214 0.339 0.892 1.444 1.997 (b) Temperature Violation RBC A2C BDQ DQN PPO TransformerRL MBPO MBVE DreamerV3 ADAPT(Ours) Figure 9: Comparison of ADAPT with baseline controllers in Sinergym under the ID control setting for the Arizona environment. Subfigure (a) reports energy saving relative to RBC, and subfigure (b) reports temperature violation. Higher energy saving and lower temperature violation indicate bet- ter overall control performance. C.2 SemibuildSim OOD Control Comparison Figure 10 shows that under Summer→Winter seasonal transfer, ADAPTconsistently achieves the best balance between HVAC en- ergy efficiency and occupant comfort. The proposed physics-aware diffusion IEWM accurately captures transferable building thermal dynamics across seasons, enabling robust downstream control with lower energy consumption while maintaining indoor temperatures within the desired comfort range. ADAPT (Ours) w/o Phys. Mamba WM VAE WM 220 230 Energy (kWh) ADAPT (Ours) w/o Phys. Mamba WM VAE WM 0.30 0.35 0.40 Abs PMV ADAPT (Ours) w/o Phys. Mamba WM VAE WM 8 9 10 PPD (%) ADAPT (Ours) w/o Phys. Mamba WM VAE WM 4 6 Action Fluctuation ADAPT (Ours)w/o Phys.Mamba WMVAE WM Figure 10: Comparison of different IEWMs downstream con- trol performance on SemibuildingSim Winter→Summer transfer. Lower is better. 15 Conference’27, August 2027, San Jose, California, USAYang et al. D Implementation Details D.1 Implementation of World Models This section describes the implementation details of the world models used in our experiments. The proposed method is ADAPT , the two neural world-model baselines are the MambaFormer and the VAE. All models are implemented in PyTorch and optimized with Adam. Dataset collection and input layout. We construct transition datasets with a held-action protocol. Each training sample contains a his- tory window as the conditioning variable and a flattened future observation sequence as the prediction target, 푐 푡 =[o 푡−푁+1:푡 , a 푡−푁+1:푡−1 ,푎 푡 :푡+퐻−1 = 푎 푡−1 ],y= o 푡+1:푡+퐻 . (30) The previous action푎 푡−1 is held fixed over the forecast horizon, matching the dataset generation procedure used by all world models. In SemiBuildingSim, the control interval is 5 minutes, the context length is푁=6, and the forecast horizon is퐻=3. In Sinergym, the control interval is 15 minutes, the context length is푁=4, and the forecast horizon is 퐻= 2. Normalization and training protocol. Continuous inputs and tar- gets are normalized before training. Unless otherwise specified, models use a random seed of 3407, test split 0.2, batch size 64, learn- ing rate 10 −4 , weight decay 10 −5 , maximum 200 epochs, gradient clipping with norm 1.0, and early stopping on validation RMSE with patience 12. ADAPT diffusion world model. ADAPT formulates world-model prediction as conditional denoising. A Temporal U-Net denoiser 푓 휃 is trained inside a Gaussian diffusion model with 20 denoising steps, dimension multiplier(8), hidden dimension 256, exponential moving average decay 0.995, and gradient accumulation every 2 mini-batches. The model uses an L1 denoising loss. Given the conditioning vector푐 푡 , the reverse process iteratively samples ˆ y∼ 푝 휃 (y | c),(31) and the resulting denormalized sequence is evaluated as the multi- step world-model forecast. At evaluation time we use the EMA denoiser and enable clipped denoising for stable conditional sam- ples. Thermal physics regularization. ADAPT augments the denoising objective with a differentiable multi-zone RC residual. For room푖, the residual is based on 퐶 푖 푑푇 푖 푑푡 =−[퐾T] 푖 +푘 out 푖 푇 out +푏 hvac 푖 푢 푖 +푏 occ 푖 푂 푖 +푑 푖 ,(32) where퐾= 퐿 + diag(푘 out )and퐿is a learned symmetric graph Laplacian. The HVAC driver uses the held FCU fan mode and supply- water temperature, 푢 푖 =−휙 푖 (fan 푖 )(푇 푖 −푇 sup 푖 ),(33) where휙 푖 (·)is a learned monotone map over the discrete fan modes. The physical penalty combines a Huber penalty on the predicted trajectory residual, a Huber penalty on the true trajectory residual for RC parameter calibration, and a small bias regularizer: L phys = 휆 pred 휌 훽 (푟 pred )+ 휆 true 휌 훽 (푟 true )+ 휆 d ∥푑∥ 2 2 .(34) Table 5: Physical-loss hyper-parameters of ADAPT. ParameterValue Predicted RC residual weight 휆 pred 0.2 True RC residual weight 휆 true 0.5 Bias penalty weight 휆 bias 1×10 −1 Time stepΔ푡1.0 Huber threshold 훽0.25 Condition history[o 푡−푁+1:푡 , a 푡−푁+1:푡−1 ] Physical/min-max normalization Future target y= o 푡+1:푡+퐻 Forward diffusion: add Gaussian noise Temporal U-Net denoiser Reverse diffusion, 20 steps Forecast ˆ y Residual loss on denormalized temperatures Figure 11: ADAPT architecture for conditional diffusion world-model prediction. Tab. 5 reports the physical-loss parameters used by ADAPT . Evaluation metrics. For continuous state variables, we report Mean Absolute Error (MAE), Root Mean Squared Error (RMSE), coefficient of variation of RMSE (CVRMSE), and푅 2 where applicable. For SemiBuildingSim occupancy-related targets, we additionally report Occupancy Exact Match Rate (OCC EMR) after converting the normalized predictions back to the original scale. ADAPT architecture. Figure 11 illustrates the ADAPT training and sampling pipeline. The conditioning history is repeated across the diffusion horizon, the Temporal U-Net denoises the future target vector, and the heat-balance equation regularizes the denormalized temperature trajectory during training. D.2 MPC Implementation Details We use a sampling-based model predictive control (MPC) approach for online decision making. At each environment step, the controller plans over a finite horizon using the learned indoor environment world model (IEWM), evaluates imagined trajectories using a combi- nation of a reward evaluator and a learned terminal value function, and executes only the first action of the best-scoring sequence. The procedure is repeated after the next observation is received. 16 ADAPT: A Diffusion-based Adaptive Physics-aware Indoor Environmental World Model for Transferable HVAC ControlConference’27, August 2027, San Jose, California, USA Planning objective. Let표 푡 denote the current observation and let a 푡 :푡+퐻−1 = (푎 푡 , . . .,푎 푡+퐻−1 )be a candidate action sequence of horizon퐻. The IEWM is used to recursively predict future states. We specifically utilize its single-step forward prediction, ˆ 표 푡+푘+1 = 푓 휃 ( ˆ 표 푡+푘 ,푎 푡+푘 ), ˆ 표 푡 =표 푡 ,(35) where all planning rollouts are performed in the prediction space of the IEWM. To account for the long-term return beyond the fi- nite planning horizon, we bootstrap using a learned terminal value function푉 휓 (·). The quality of a candidate sequence is computed by accumulating discounted rewards and adding the temporal differ- ence (TD) value estimate at the terminal state, 퐽(a 푡 :푡+퐻−1 )= 퐻−1 ∑︁ 푘=0 훾 푘 푔( ˆ 표 푡+푘 ,푎 푡+푘 )+훾 퐻 푉 휓 ( ˆ 표 푡+퐻 ),(36) where 푔(·) is the reward evaluator. Iterative action-sequence planning. Following the trajectory op- timization scheme used in TD-MPC2 [22], we employ an itera- tive sampling procedure. At each planning iteration푗, the MPC maintains a factorized Gaussian proposal distribution over action sequences: 푞 (푗) (a)= 퐻−1 Ö 푘=0 N(푎 푡+푘 ;휇 (푗) 푘 ,(휎 (푗) 푘 ) 2 퐼).(37) For the initial iteration (푗=1), the mean sequence휇 (1) is initialized to the center of the valid action space. At each iteration, we sample 푁 candidate action sequences from this proposal distribution. Each sampled sequence is rolled out with the IEWM and scored by Eq.(36). We then select the top퐾sequences to update the dis- tribution parameters휇 (푗+1) and휎 (푗+1) for the next iteration using the Cross-Entropy Method (CEM). After푀planning iterations, the controller selects the highest-scoring sequence from the final elite set, a ★ 푡 :푡+퐻−1 = arg max a (푖) 푡 :푡+퐻−1 퐽(a (푖) 푡 :푡+퐻−1 ),(38) and executes only its first action푎 ★ 푡 . This receding-horizon proce- dure replans after every real environment transition. Hyperparameters. We use the MPC hyperparameters detailed in Tab. 6. Table 6: MPC hyperparameters. HyperparameterValue Planning horizon 퐻6 Number of candidate sequences 푁16 Number of elite samples 퐾4 Planning iterations 푀3 Discount factor훾0.98 D.3 Online RL Policy Learning All methods are trained for a total of 10 6 timesteps. For the discrete action space, DQN utilizes aDiscreterepresentation, while all other algorithms employ aMultiDiscreteaction space, modeled as factorized categorical distributions over action branches. To ensure a fair comparison, model-based extensions (MBVE, MBPO) are implemented using the same PPO backbone. Detailed hyper- parameters for our proposed DECARB, the model-free baselines, and the model-based baselines are summarized in Tab. 7, Tab. 8, and Tab. 9, respectively. D.4 Computational Cost All experiments are conducted on an NVIDIA A100 GPU. The offline training of the indoor environmental world model requires about 12 hours. Subsequently, the downstream online RL training takes an average of 10 hours in SemiBuildingSim and 16 hours in the Sinergym environment. During inference, one time step diffusion costs about 150ms, ensuring the real-time control of HVAC systems. 17 Conference’27, August 2027, San Jose, California, USAYang et al. Table 7: Hyper-parameters for the proposed ADAPT (퐻 푓표푟푒 ,퐻,푇 on SemiBuildingSim). Hyper-parameterValueHyper-parameterValue OptimizerAdamLearning rate2×10 −3 Discount factor훾0.99 Replay buffer size5×10 5 Batch size64Exploration 휀1.0→ 0.04 Target update interval1000 stepsTD(휆) parameter0.8 Bootstrap horizon 퐻3 Forecaster horizon 퐻 fore 3 Network architecture (512, 512)Forecaster History step푇6 Table 8: Hyper-parameters for Model-Free Baselines. ParameterDQNBDQA2C PPO TransformerRL Learning rate2×10 −3 2×10 −3 8×10 −4 10 −3 10 −3 Discount factor훾0.990.990.980.980.98 Batch size64644012001200 Network hidden dim256512256256256 GAE 휆–0.90.80.8 PPO Clip 휀–0.20.2 Entropy coef.–00.010.01 Table 9: Hyper-parameters for Model-Based Baselines (MBVE and MBPO built on the PPO backbone). ParameterMBVEMBPODreamerV3 Rollout / Imagination horizon888 Transitions / Starts per iter.2561024256 Blend / Return 휆1.0–0.95 Warmup (iters)333 18