Next Article in Journal
Probabilistic Voltage Stability Screening Under Stochastic Load Allocation at Weak Buses Using Stability Index
Next Article in Special Issue
Type-2 Fuzzy C-Means-Based Clustering-Decomposed Coordination of Directional Overcurrent Relays
Previous Article in Journal
Progress in the Energy Transition Process in EU Countries—A Sustainable Multi-Criteria Assessment
Previous Article in Special Issue
A Review of Power Grid Frameworks for Planning Under Uncertainty
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Reinforcement Learning Methods for the Stochastic Optimal Control of an Industrial Power-to-Heat System

1
Institute of Mathematics, Brandenburg University of Technology Cottbus-Senftenberg, 03013 Cottbus, Germany
2
Department Simulation and Virtual Design, Institute of Low-Carbon Industrial Processes, German Aerospace Center, Weinbergstraße 10, 03050 Cottbus, Germany
*
Author to whom correspondence should be addressed.
Energies 2026, 19(4), 1046; https://doi.org/10.3390/en19041046
Submission received: 18 December 2025 / Revised: 5 February 2026 / Accepted: 8 February 2026 / Published: 17 February 2026
(This article belongs to the Special Issue Optimization and Machine Learning Approaches for Power Systems)

Abstract

The optimal control of sustainable energy supply systems, including renewable energies and energy storage, takes a central role in the decarbonization of industrial systems. However, the use of fluctuating renewable energies leads to fluctuations in energy generation and requires a suitable control strategy for the complex systems in order to ensure energy supply. In this paper, we consider an electrified power-to-heat system which is designed to supply heat in the form of superheated steam for industrial processes. The system consists of a high-temperature heat pump for heat supply, a wind turbine for power generation, a sensible thermal energy storage for storing excess heat, and a steam generator for providing steam. If the system’s energy demand cannot be covered by electricity from the wind turbine, additional electricity must be purchased from the power grid. For this system, we investigate the cost-optimal operation, aiming to minimize the electricity cost from the grid by a suitable system control depending on the available wind power and the amount of stored thermal energy. This is a decision-making problem under uncertainty regarding the future prices for electricity from the grid and the future generation of wind power. The resulting stochastic optimal control problem is treated as finite-horizon Markov decision process for a multi-dimensional controlled state process. We first consider the classical backward recursion technique for solving the associated dynamic programming equation for the value function and compute the optimal decision rule. Since that approach suffers from the curse of dimensionality, we also apply reinforcement learning techniques, namely Q-learning, that are able to provide a good approximate solution to the optimization problem within reasonable time.
Keywords: stochastic optimal control; Markov decision process; dynamic programming; Q-learning; power-to-heat system; renewable energy; cost-optimal energy management stochastic optimal control; Markov decision process; dynamic programming; Q-learning; power-to-heat system; renewable energy; cost-optimal energy management

Share and Cite

MDPI and ACS Style

Pilling, E.; Bähr, M.; Wunderlich, R. Reinforcement Learning Methods for the Stochastic Optimal Control of an Industrial Power-to-Heat System. Energies 2026, 19, 1046. https://doi.org/10.3390/en19041046

AMA Style

Pilling E, Bähr M, Wunderlich R. Reinforcement Learning Methods for the Stochastic Optimal Control of an Industrial Power-to-Heat System. Energies. 2026; 19(4):1046. https://doi.org/10.3390/en19041046

Chicago/Turabian Style

Pilling, Eric, Martin Bähr, and Ralf Wunderlich. 2026. "Reinforcement Learning Methods for the Stochastic Optimal Control of an Industrial Power-to-Heat System" Energies 19, no. 4: 1046. https://doi.org/10.3390/en19041046

APA Style

Pilling, E., Bähr, M., & Wunderlich, R. (2026). Reinforcement Learning Methods for the Stochastic Optimal Control of an Industrial Power-to-Heat System. Energies, 19(4), 1046. https://doi.org/10.3390/en19041046

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop