Next Article in Journal
Experimental Implementation of a Second-Order Adaptive Fuzzy Logic Controller for PMSG-Based Wind Energy Conversion Systems
Previous Article in Journal
Distributed Quantum-Assisted Multi-SAPF Architecture Based on Deterministic Current Control and Asynchronous QUBO–QAOA–VQE Supervisory Optimization
Previous Article in Special Issue
Design of Parallel Hybrid Active Power Filter with Adaptive DC-Link Voltage Control Based on Artificial Neural Network
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Stable Offline Reinforcement Learning for Switched Reluctance Motor Drives via Multi-Demonstrator Policy Distillation

by
Franklin Sánchez
1,*,
María Isabel Milanés-Montero
2 and
Enrique Romero-Cadaval
2
1
Departamento de Eléctrica, Electrónica y Telecomunicaciones, Universidad de las Fuerzas Armadas ESPE, Sangolquí 171103, Ecuador
2
Power Electrical and Electronic Systems (PE&ES), School of Industrial Engineering, University of Extremadura, 06006 Badajoz, Spain
*
Author to whom correspondence should be addressed.
Electronics 2026, 15(18), 4289; https://doi.org/10.3390/electronics15184289 (registering DOI)
Submission received: 14 August 2026 / Revised: 2 September 2026 / Accepted: 8 September 2026 / Published: 19 September 2026
(This article belongs to the Special Issue Power Quality and Power Electronics Systems in Electromobility)

Abstract

Finite-control-set model predictive control provides excellent torque–speed regulation for switched reluctance motor drives but requires an online combinatorial search at every control instant, making low-cost embedded implementation challenging. This article investigates whether offline reinforcement learning can distill policies from multiple classical controllers into a single feedforward policy requiring neither online optimization nor controller gain tuning. A replay buffer is populated with trajectories generated by three demonstrators—hysteresis current control, proportional–integral control with pulse-width modulation, and finite-control-set model predictive control—using a finite-element model of a four-phase 8/6 switched reluctance machine parameterized from measurements of the physical drive. An implicit Q-learning agent then learns a control policy without evaluating actions outside the offline dataset. The central finding is that demonstration diversity governs the stability of offline reinforcement learning on this problem: policies trained from a single demonstrator experience early mode collapse in all fifteen runs, whereas two- or three-demonstrator datasets converge stably in all fifteen. Behavior cloning trained on the identical buffer, split, architecture, and deployed controller provides the reference point for interpreting this result. It matches the offline RL policy on torque quality and improves on its speed regulation, exhibiting none of the seed-to-seed fragility seen at no load while requiring roughly 8% more switching transitions. The stability requirement therefore appears to be a property of the advantage-weighted offline RL objective rather than the control task, and the measured benefit of that objective on this problem is confined to switching effort. We report this rather than claim a broader advantage. The characterization of the distilled controller shows that it generalizes to operating points that are not included in the training dataset, gains nothing systematic beyond approximately 60% of the replay buffer, remains insensitive to ±20% perturbations of all reward weights, and degrades gracefully under measurement noise while the current mask enforces the peak-current constraint throughout. A deployment analysis shows that the 18,432 multiply–accumulate policy meets a 50μs control period in its existing form at a measured cost of about 2% in torque ripple. All the results are simulation-based on a finite-element model parameterized from a physical machine.
Keywords: switched reluctance motor; offline reinforcement learning; implicit Q-learning; model predictive control; policy distillation; torque ripple; training stability; electromobility switched reluctance motor; offline reinforcement learning; implicit Q-learning; model predictive control; policy distillation; torque ripple; training stability; electromobility

Share and Cite

MDPI and ACS Style

Sánchez, F.; Milanés-Montero, M.I.; Romero-Cadaval, E. Stable Offline Reinforcement Learning for Switched Reluctance Motor Drives via Multi-Demonstrator Policy Distillation. Electronics 2026, 15, 4289. https://doi.org/10.3390/electronics15184289

AMA Style

Sánchez F, Milanés-Montero MI, Romero-Cadaval E. Stable Offline Reinforcement Learning for Switched Reluctance Motor Drives via Multi-Demonstrator Policy Distillation. Electronics. 2026; 15(18):4289. https://doi.org/10.3390/electronics15184289

Chicago/Turabian Style

Sánchez, Franklin, María Isabel Milanés-Montero, and Enrique Romero-Cadaval. 2026. "Stable Offline Reinforcement Learning for Switched Reluctance Motor Drives via Multi-Demonstrator Policy Distillation" Electronics 15, no. 18: 4289. https://doi.org/10.3390/electronics15184289

APA Style

Sánchez, F., Milanés-Montero, M. I., & Romero-Cadaval, E. (2026). Stable Offline Reinforcement Learning for Switched Reluctance Motor Drives via Multi-Demonstrator Policy Distillation. Electronics, 15(18), 4289. https://doi.org/10.3390/electronics15184289

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop