Enhancing Energy Management in Multi-Zone Buildings Using the On-Policy Reinforcement Learning Algorithm SARSA

Authors

  • Mohamed Abdalla Abd El-Hameed Attia Computers and Control Department, Faculty of Engineering, Tanta University, Tanta, Egypt
  • Mohamed Abdelaal Ahmed Abdelaal Computers and Control Department, Faculty of Engineering, Tanta University, Tanta, Egypt
  • Elsayed Abd Elhameed Sallam Computers and Control Department, Faculty of Engineering, Tanta University, Tanta, Egypt

DOI:

https://doi.org/10.54536/ajsts.v5i2.7966

Keywords:

Buildings Energy Management, Energy Management, HVAC Control, Multi-Zone Buildings, Reinforcement Learning, Sarsa, Smart Buildings

Abstract

THVAC systems in commercial buildings, particularly open-plan offices, account for substantial energy use. Thermal coupling between adjacent zones complicates control, while conventional rule-based and model-based approaches may not adapt effectively to changing occupancy and environmental conditions. Deep reinforcement learning methods can also impose high computational demands, limiting their practical deployment. This study developed and evaluated a lightweight, computationally inexpensive HVAC control policy for multi-zone open-plan offices using the on-policy reinforcement learning algorithm SARSA. A tabular SARSA controller was implemented in an EnergyPlus–Python co-simulation of a six-zone office model. The agent was trained using realistic occupancy schedules, TMY3 weather data, and 5-minute control intervals. A multi-objective reward function balanced energy efficiency, thermal comfort, and switching frequency. Performance was benchmarked against rule-based, model predictive control (MPC), and deep Q-network (DQN) controllers across several U.S. climate regions. The SARSA controller reduced annual HVAC energy use by 39.39%, from 18,200 to 11,030 kWh. Thermal comfort violations were limited to 2.85%, with an average temperature deviation of 1.2 °F and a 15-minute recovery time. Switching frequency decreased to 4.1 ON/OFF cycles per zone per day, approximately half that of rule-based control. Across different climates, energy savings ranged from 32.6% to 40.7%. These findings show that tabular SARSA is an effective and computationally efficient alternative to deep reinforcement learning for real-time HVAC control in multi-zone smart buildings.

Downloads

Download data is not yet available.

References

Afram, A., & Janabi-Sharifi, F. (2014). Theory and applications of HVAC control systems–a review of model predictive control (MPC). Building and Environment, 72, 343-355.

Afroz, Z., Shafiullah, G., Urmee, T., & Higgins, G. (2018). Modeling techniques used in building HVAC control systems: a review. Renewable and sustainable energy reviews, 83, 64-84.

Ahmad, M. W., Mourshed, M., Yuce, B., & Rezgui, Y. (2016). Computational intelligence techniques for HVAC systems: a review. Building Simulation,

Asim, N., Badiei, M., Mohammad, M., Razali, H., Rajabi, A., Chin Haw, L., & Jameelah Ghazali, M. (2022). Sustainability of heating, ventilation and air-conditioning (HVAC) systems in buildings—an overview. International journal of environmental research and public health, 19(2), 1016.

Azuatalam, D., Lee, W.-L., De Nijs, F., & Liebman, A. (2020). Reinforcement learning for whole-building HVAC control and demand response. Energy and AI, 2, 100020.

Balaji, B., Xu, J., Nwokafor, A., Gupta, R., & Agarwal, Y. (2013). Sentinel: occupancy based HVAC actuation using existing WiFi infrastructure within commercial buildings. Proceedings of the 11th ACM conference on embedded networked sensor systems,

Bengea, S. C., Kelman, A. D., Borrelli, F., Taylor, R., & Narayanan, S. (2014). Implementation of model predictive control for an HVAC system in a mid-size commercial building. HVAC&R Research, 20(1), 121-135.

Bianchini, G., Casini, M., Pepe, D., Vicino, A., & Zanvettor, G. G. (2019). An integrated model predictive control approach for optimal HVAC and energy storage operation in large-scale buildings. Applied Energy, 240, 327-340.

Cao, Y., Wei, W., Wang, J., Mei, S., Shafie-khah, M., & Catalao, J. P. (2018). Capacity planning of energy hub in multi-carrier energy networks: a data-driven robust stochastic programming approach. IEEE Transactions on Sustainable Energy, 11(1), 3-14.

Ding, X., Du, W., & Cerpa, A. E. (2020). Mb2c: Model-based deep reinforcement learning for multi-zone building control. Proceedings of the 7th ACM international conference on systems for energy-efficient buildings, cities, and transportation

DOE, U. (2010). EnergyPlus Documentation. US Department of Energy.

Freire, R. Z., Oliveira, G. H., & Mendes, N. (2008). Predictive controllers for thermal comfort optimization and energy savings. Energy and Buildings, 40(7), 1353-1365.

Fu, Q., Hu, L., Wu, H., Hu, F., Hu, W., & Chen, J. (2018). A Sarsa-based adaptive controller for building energy conservation. Journal of Computational Methods in Science and Engineering, 18(2), 329-338.

Gao, G., Li, J., & Wen, Y. (2020). DeepComfort: energy efficient thermal comfort control in buildings via reinforcement learning. IEEE Internet of Things Journal, 7(9), 8472-8484.

Ghahramani, A., Tang, C., & Becerik-Gerber, B. (2015). An online learning approach for quantifying personalized thermal comfort via adaptive stochastic modeling. Building and Environment, 92, 86-96.

Guo, S., Zheng, S., Hu, Y., Hong, J., Wu, X., & Tang, M. (2019). Embodied energy use in the global construction industry. Applied Energy, 256, 113838.

Homod, R. Z., Mahlia, T., & Mohamed, H. A. (2009). Pid-cascade for hvac system control. International Conference on Control, Instrumentation and Mechatronic Engineering (CIM09)

Jiang, Z., Risbeck, M. J., Ramamurti, V., Murugesan, S., Amores, J., Zhang, C., Lee, Y. M., & Drees, K. H. (2021). Building HVAC control with reinforcement learning for reduction of energy cost and demand charge. Energy and Buildings, 239, 110833.

Kouvaritakis, B., & Cannon, M. (2016). Model predictive control. Switzerland: Springer International Publishing, 38(13-56), 7.

Lazaridis, C. R., Michailidis, I., Karatzinis, G., Michailidis, P., & Kosmatopoulos, E. (2024). Evaluating reinforcement learning algorithms in residential energy saving and comfort management. Energies, 17(3), 581.

Li, F., & Du, Y. (2023). Intelligent multi-zone residential HVAC control strategy based on deep reinforcement learning. In Deep Learning for Power System Applications: Case Studies Linking Artificial Intelligence and Power Systems (pp. 71-96). Springer.

Michailidis, P., Michailidis, I., & Kosmatopoulos, E. (2025). Reinforcement learning for optimizing renewable energy utilization in buildings: a review on applications and innovations. Energies, 18(7), 1724.

Nguyen, D. H. (2024). The role of reinforcement learning control for optimizing building energy management systems.

Niemann, P., & Schmitz, G. (2020). Impacts of occupancy on energy demand and thermal comfort for a large-sized administration building. Building and Environment, 182, 107027.

Ojadi, J. O., Odionu, C. S., Onukwulu, E. C., & Owulade, O. A. (2024). AI-enabled smart grid systems for energy efficiency and carbon footprint reduction in urban energy networks. International Journal of Multidisciplinary Research and Growth Evaluation, 5(1), 1549-1566.

Pokharel, B. (2025). A comprehensive study of modern control strategies: From classical PID to optimal, robust, adaptive and intelligent methods. International Research Journal of Modernization in Engineering Technology and Science, 7(9), 3035–3049. https://doi.org/10.56726/IRJMETS83160

Purwanto, Y. K., & Kang, D.-K. (2024). Multi-agent deep reinforcement learning for fighting game: a comparative study of PPO and A2C. International Journal of Internet, Broadcasting and Communication, 16(3), 192-198.

Sierla, S., Ihasalo, H., & Vyatkin, V. (2022). A review of reinforcement learning applications to control of heating, ventilation and air conditioning systems. Energies, 15(10), 3526.

Sivamayil, K., Rajasekar, E., Aljafari, B., Nikolovski, S., Vairavasundaram, S., & Vairavasundaram, I. (2023). A systematic study on reinforcement learning based applications. Energies, 16(3), 1512.

Standard, A. (1992). Thermal environmental conditions for human occupancy. ANSI/ASHRAE, 55, 5.

Sutton, R. S., & Barto, A. G. (1998). Reinforcement learning: An introduction (Vol. 1). MIT press Cambridge.

Tariq, S., Ali, U., Kim, S., & Yoo, C. (2025). Multi-agent distributed reinforcement learning for energy-efficient thermal comfort control in multi-zone buildings with diverse occupancy patterns. Energy, 137082.

Watkins, C. J., & Dayan, P. (1992). Q-learning. Machine learning, 8(3), 279-292.

Wei, T., Wang, Y., & Zhu, Q. (2017). Deep reinforcement learning for hvac control in smart buildings. Design Automation Conference (DAC)

Wei, X., Kusiak, A., Li, M., Tang, F., & Zeng, Y. (2015). Multi-objective optimization of the HVAC (heating, ventilation, and air conditioning) system performance. Energy, 83, 294-306.

Zare, M., Kebria, P. M., Khosravi, A., & Nahavandi, S. (2024). A survey of imitation learning: algorithms recent developments, and challenges. IEEE Transactions on Cybernetics.

Downloads

Published

2026-09-14

How to Cite

El-Hameed Attia, M. A. A. ., Abdelaal, M. A. A. ., & Sallam, E. A. E. . (2026). Enhancing Energy Management in Multi-Zone Buildings Using the On-Policy Reinforcement Learning Algorithm SARSA. American Journal of Smart Technology and Solutions, 5(2), 106-115. https://doi.org/10.54536/ajsts.v5i2.7966

Similar Articles

31-40 of 50

You may also start an advanced similarity search for this article.