An Optimized Active Compensation Control Framework for High-Speed Railway Pantograph via Imitation-Guided Deep Reinforcement Learning
Abstract
1. Introduction
- A CPO-LQR baseline controller is developed by optimizing the -weight matrix via the CPO algorithm, achieving strong PCCF suppression; its control outputs are further refined using an offline control law based on the dual pantograph–catenary model structure, resulting in expert actions that outperform the baseline and provide theoretically optimal control performance.
- To enhance practical applicability, a compensation control strategy is proposed by superimposing the offline expert force as well as the compensation force onto real-time CPO-LQR output , enabling an active controller based on a single pantograph–catenary model structure and yielding superior suppression of extreme oscillations.
- The compensatory action is trained using the BC-SAC algorithm, which embeds behavioral cloning loss into SAC to allow partial expert imitation while preserving the environment-driven interaction capabilities; an attenuation mechanism further balances exploration and imitation, leading to better performance than pure expert-based compensation.
2. Preliminaries and Problem Formulation
2.1. Mathematical Modeling of the Pantograph–Catenary Coupling System
- Lateral vibration effects of the catenary are omitted. Although studies [31,32] have employed aerodynamic simulations and finite-element analyses to examine catenary lateral dynamics under abnormal conditions, our work focuses on vertical current collection and control algorithm design, since under normal operating conditions lateral contact-wire vibration is not a primary driver of current collection performance.
- The contact wire and messenger wire are modeled as Euler–Bernoulli beams with constant stiffness and tension.
- Adjacent anchor sections of the catenary are treated as independent subsystems.
- The mid-span of each dropper is represented by a spring element, as illustrated in Figure 2.
2.2. Markov Decision Process Modeling
2.2.1. State Space
2.2.2. Action Space
2.3. Definition of the Objective Function and the Reward Function
2.3.1. Specification of Performance Metrics
2.3.2. Objective Function
2.3.3. Reward Function
3. CPO-LQR-BC-SAC Based Active Control Strategy
3.1. Overall Framework and Design Rationale
- Baseline Controller Construction—Part I:Use the CPO algorithm to offline-tune the LQR weight matrix, obtaining high-performance baseline control actions.
- Expert Action Refinement—Part II:Based on the dual pantograph–catenary model, design an offline control law (CPO-LQR-CtrlLaw) to perform a secondary optimization of the baseline actions, producing ideal control actions that surpass the baseline performance—these are treated as expert actions for the agent’s policy to imitate.
- Compensator Deployment—Part III:Given that standalone imitation learning or reinforcement learning controls cannot fully outperform the baseline controller, and that the dual-model approach faces practical deployment constraints, we conducted experiments that show superimposing real-time CPO-LQR outputs with the secondarily optimized control actions more effectively suppress extreme oscillations. Accordingly, a DRL compensator integrating imitation learning was constructed to replace the secondary optimization control actions. In the online test, the trained BC-SAC compensatory controller works in parallel with the baseline LQR; their outputs are weighted and combined to act on the pantograph–catenary system, enhancing fluctuation suppression and contact force dynamic performance.
3.2. Design of the Primary Controller Based on CPO-LQR
3.2.1. LQR Controller Baseline
3.2.2. Optimization Method for the -Weight Matrix
3.2.3. The Offline Control Law
3.3. Design of the Compensatory Controller Based on BC-SAC
3.3.1. Soft Actor-Critic Algorithm
3.3.2. Actor Network Integrated with Behavior Cloning
3.4. Update Process of the Proposed Framework
| Algorithm 1. Pseudocode of the active pantograph compensation control algorithm based on the CPO-LQR-BC-SAC framework. |
| The CPO-LQR-BC-SAC Framework |
| Input: |
| Environment of the pantograph–catenary coupling system. |
| -weight matrix and fixed -weight matrix. |
| tuned by the offline control law. |
| . |
| Replay Buffer . |
| Behavior-cloning weight decay schedule . |
| Output: |
| Trained compensatory policy . |
| Procedure: |
| 1. Initialization: |
| Randomly initialize Actor parameters . |
| . |
| Initialize temperature coefficient . |
| Initialize the replay buffer as empty. |
| Initialize behavior-cloning weight . |
| 2. for episode = 1 to do |
| . |
| 4. for step = 1 to M do |
| . |
| . |
| . |
| . |
| , and done. |
| , done) into D. |
| . |
| 12. if size(D) ≥ batch size, then: |
| 13. Sample size per batch B from D. |
| 14. Update Critic networks using the entropy-augmented Bellman targets based on (21) and (22). |
| 15. Compute current behavior-cloning weight based on (27). |
| 16. Update Actor network based on (25) and (26). |
| 17. Soft-update target Critic network parameters based on (23). |
| 18. Update temperature parameter based on (24). |
| 19. if done break |
| 20. end for |
| 21. end for |
| . |
4. Experimental Validation and Result Analysis
4.1. Hyperparameter Settings
4.2. Control Performance of the Baseline Controller
4.2.1. Training and Testing Results
4.2.2. Comparative Validation with the Offline Control Law
4.3. Control Performance of the Proposed Framework
4.3.1. Training and Testing Results
4.3.2. Comparative Validation Based on Expert Action Compensation
4.4. Performance Evaluation
4.4.1. Speed Range Performance Evaluation and Comparative Validation
4.4.2. Robustness Validation Across Different Pantograph Types
5. Conclusions
- (1)
- The baseline CPO-LQR controller is constructed and optimized using the CPO algorithm, yielding a high-performance control policy that effectively reduces PCCF fluctuations across varying speeds.
- (2)
- An offline secondary-tuned control law based on a dual-model structure further refines the control actions and provides expert demonstrations that enhance oscillation suppression.
- (3)
- A practical compensation strategy is developed by integrating real-time CPO-LQR outputs with expert action corrections within a unified single-model framework.
- (4)
- The CPO-LQR-BC-SAC learning framework is trained through a hybrid of behavior cloning and SAC, enabling it to imitate expert actions while maintaining exploratory capabilities.
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Wu, G.; Dong, K.; Xu, Z. Pantograph–catenary electrical contact system of high-speed railways: Recent progress, challenges, and outlooks. Railway Eng. Sci. 2022, 30, 437–467. [Google Scholar] [CrossRef] [Scilit]
- Daocharoenporn, S.; Mongkolwongrojn, M.; Kulkarni, S.; Shabana, A.A. Prediction of the pantograph/catenary wear using nonlinear multibody system dynamic algorithms. J. Tribol. 2019, 141, 051603. [Google Scholar] [CrossRef] [Scilit]
- Karaduman, G.; Akin, E. A deep learning based method for detecting wear on the current collector strips’ surfaces of the pantograph in railways. IEEE Access 2020, 8, 183799–183812. [Google Scholar] [CrossRef] [Scilit]
- Wu, Q.; Gu, X.P.; Ma, Z.; Wang, A. A study on the vibration characteristics and damage mechanism of pantograph strips in a railway electrification system. Machines 2022, 10, 710. [Google Scholar] [CrossRef] [Scilit]
- Mariscotti, A. The electrical behaviour of railway pantograph arcs. Energies 2023, 16, 1465. [Google Scholar] [CrossRef] [Scilit]
- Liu, Y.; Quan, W.; Lu, X.; Liu, X.; Gao, S.; Zhao, H.; Yu, L.; Zheng, J. A novel arcing detection model of pantograph–catenary for high-speed train in complex scenes. IEEE Trans. Instrum. Meas. 2023, 72, 5012013. [Google Scholar] [CrossRef] [Scilit]
- Yu, X.; Wang, Z.; Song, M.; Song, L.; Yang, J.; Su, Y. Simulation study on arc temperature of urban rail DC pantograph–catenary and arc ablation of contact line. Machines 2024, 12, 514. [Google Scholar] [CrossRef] [Scilit]
- Al-Awad, N.A.; Abboud, I.K.; Al-Rawi, M.F. Genetic algorithm–PID controller for model order reduction pantograph–catenary system. Appl. Comput. Sci. 2021, 17, 28–39. [Google Scholar] [CrossRef] [Scilit]
- Farhan, M.F.; Shukor, N.S.A.; Ahmad, M.A.; Suid, M.H.; Ghazali, M.R.; Jusof, M.F. A simplified fuzzy logic controller design based safe experimentation dynamics for pantograph–catenary system. Indones. J. Electr. Eng. Comput. Sci. 2019, 14, 903–911. [Google Scholar] [CrossRef] [Scilit]
- Song, Y.; Liu, Z.; Ouyang, H.; Wang, H.; Lu, X. Sliding mode control with PD sliding surface for high-speed railway pantograph-catenary contact force under strong stochastic wind field. Shock Vib. 2017, 2017, 4895321. [Google Scholar] [CrossRef] [Scilit]
- Wang, B.; Wen, S.; Shen, Y. Active LQR control of a fractional-order pantograph–catenary system based on feedback linearization. Math. Probl. Eng. 2022, 2022, 2213697. [Google Scholar] [CrossRef] [Scilit]
- Jin, X.; Lv, H.; Tao, Y.; Lu, J.; Lv, J.; Opinat Ikiela, N.V. Deep reinforcement learning-based active disturbance rejection control for trajectory tracking of autonomous ground electric vehicles. Machines 2025, 13, 523. [Google Scholar] [CrossRef] [Scilit]
- Lin, Y.; Liu, X.; Zheng, Z. Discretionary lane-change decision and control via parameterized soft actor–critic for hybrid action space. Machines 2024, 12, 213. [Google Scholar] [CrossRef] [Scilit]
- Gao, H.; Jiang, S.; Li, Z.; Wang, R.; Liu, Y.; Liu, J. A two-stage multi-agent deep reinforcement learning method for urban distribution network reconfiguration considering switch contribution. IEEE Trans. Power Syst. 2024, 39, 7064–7076. [Google Scholar] [CrossRef] [Scilit]
- Liu, W.; Feng, Q.; Xiao, S.; Li, H. Automatic tracking control strategy of autonomous trains considering speed restrictions: Using the improved offline deep reinforcement learning method. IEEE Access 2024, 12, 75426–75441. [Google Scholar] [CrossRef] [Scilit]
- Liu, W.; Feng, Q.; Li, H. Optimizing passengers’ experience: A goal-oriented reinforcement learning speed control approach for urban railway trains. Proc. Inst. Mech. Eng. Part F J. Rail Rapid Transit 2024, 238, 1283–1295. [Google Scholar] [CrossRef] [Scilit]
- Wang, H.; Han, Z.; Liu, Z.; Wu, Y. Deep reinforcement learning based active pantograph control strategy in high-speed railway. IEEE Trans. Veh. Technol. 2022, 72, 227–238. [Google Scholar] [CrossRef] [Scilit]
- Wang, H.; Liu, Z.; Han, Z.; Wu, Y.; Liu, D. Rapid adaptation for active pantograph control in high-speed railway via deep meta reinforcement learning. IEEE Trans. Cybern. 2023, 54, 2811–2823. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Wang, Y.; Chen, X.; Wang, Y.; Chang, Z. An Improved Deep Deterministic Policy Gradient Pantograph Active Control Strategy for High-Speed Railways. Electronics 2024, 13, 3545. [Google Scholar] [CrossRef] [Scilit]
- Sharma, R.; Mahajan, P.; Garg, R. Deep-reinforcement-learning-based controller design for pantograph and catenary system. Sādhanā 2025, 50, 46. [Google Scholar] [CrossRef] [Scilit]
- Ambrósio, J.; Pombo, J.; Pereira, M. Optimization of high-speed railway pantographs for improving pantograph–catenary contact. Theor. Appl. Mech. Lett. 2013, 3, 013006. [Google Scholar] [CrossRef] [Scilit]
- Bruni, S.; Ambrósio, J.; Carnicero, A.; Cho, Y.H.; Finner, L.; Ikeda, M.; Kwon, S.Y.; Massat, J.-P.; Stichel, S.; Tur, M.; et al. The results of the pantograph–catenary interaction benchmark. Veh. Syst. Dyn. 2015, 53, 412–435. [Google Scholar] [CrossRef] [Scilit]
- Zhu, M.; Zhang, S.Y.; Jiang, J.Z.; Macdonald, J.; Neild, S.; Antunes, P.; Pombo, J.; Cullingford, S.; Askill, M.; Fielder, S. Enhancing pantograph–catenary dynamic performance using an inertia-integrated damping system. Veh. Syst. Dyn. 2022, 60, 1909–1932. [Google Scholar] [CrossRef] [Scilit]
- Yu, W.; Li, J.; Yuan, J.; Ji, X. LQR controller design of active suspension based on genetic algorithm. In Proceedings of the 2021 IEEE 5th ITNEC, Xi’an, China, 25–27 June 2021; pp. 1056–1060. [Google Scholar] [CrossRef] [Scilit]
- Tang, L.; Luo Ren, N.; Funkhouser, S. Semi-active suspension control with PSO-tuned LQR controller based on MR damper. Int. J. Automot. Mech. Eng. 2023, 20, 10512–10522. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Li, H.; Meng, H.; Wang, Y. Dynamic characteristics of an underframe semi-active inerter-based suspended device for high-speed train based on LQR control. Bull. Pol. Acad. Sci. Tech. Sci. 2022, 70, 141722. [Google Scholar] [CrossRef] [Scilit]
- Alfi, S.; Bruni, S.; Goodall, R.M.; Ward, C.P. Secondary yaw control to improve curving vs. stability trade-off for a railway vehicle. Veh. Syst. Dyn. 2022, 61, 1367–1386. [Google Scholar] [CrossRef] [Scilit]
- EN50318; Railway Applications—Current Collection Systems—Validation of Simulation of the Dynamic Interaction Between Pantograph and Overhead Contact Line. European Committee for Electrotechnical Standardization: Brussels, Belgium, 2018; pp. 1–20.
- GB/T 32591; Railway Applications—Current Collection Systems—Validation of Simulation of the Dynamic Interaction Between the Pantograph and the Overhead Contact Line. General Administration of Quality Supervision, Inspection and Quarantine of the People’s Republic of China, and the Standardization Administration of China: Beijing, China, 2016; pp. 1–10.
- Guo, J.; Yang, S.; Gao, G. Research on active control of the pantograph–catenary system with varying stiffness (in Chinese). J. Vib. Shock. 2005, 24, 15–144. [Google Scholar] [CrossRef]
- Song, Y.; Liu, Z.; Wang, H.; Lu, X.; Zhang, J. Nonlinear analysis of wind-induced vibration of high-speed railway catenary and its influence on pantograph–catenary interaction. Veh. Syst. Dyn. 2016, 54, 723–747. [Google Scholar] [CrossRef] [Scilit]
- Song, Y.; Zhang, M.; Øiseth, O.; Rønnquist, A. Wind deflection analysis of railway catenary under crosswind based on nonlinear finite element model and wind tunnel test. Mech. Mach. Theory 2022, 168, 104608. [Google Scholar] [CrossRef] [Scilit]
- Song, Y.; Ouyang, H.; Liu, Z.; Mei, G.; Wang, H.; Lu, X. Active control of contact force for high-speed railway pantograph–catenary based on multi-body pantograph model. Mech. Mach. Theory 2017, 115, 35–59. [Google Scholar] [CrossRef] [Scilit]
- Zhou, H.; Liu, Z.; Xiong, J.; Duan, F. Characteristic analysis of pantograph–catenary detachment arc based on double-pantograph catenary dynamics in electrified railways. IET Electr. Syst. Transp. 2022, 12, 238–250. [Google Scholar] [CrossRef] [Scilit]
- Abdel-Basset, M.; Mohamed, R.; Abouhawwash, M. Crested porcupine optimizer: A new nature-inspired metaheuristic. Knowl.- Based Syst. 2024, 284, 111257. [Google Scholar] [CrossRef] [Scilit]
- Haarnoja, T.; Zhou, A.; Hartikainen, K.; Tucker, G.; Ha, S.; Tan, J.; Kumar, V.; Zhu, H.; Gupta, A.; Abbeel, P.; et al. Soft actor-critic algorithms and applications. arXiv 2018, arXiv:1812.05905. [Google Scholar] [CrossRef] [Scilit]
- Torabi, F.; Warnell, G.; Stone, P. Behavioral cloning from observation. arXiv 2018, arXiv:1805.01954. [Google Scholar] [CrossRef] [Scilit]





















| Description of Parameters | Symbols | Values |
|---|---|---|
| Population size | ||
| Maximum number of iterations | ||
| Cycle reduction period | ||
| Minimum population size | ||
| Initial visual deterrence factor | ||
| Acoustic deterrence factor | ||
| Differential evolution mutation factor | ||
| Crossover probability | ||
| Lévy flight distribution exponent | ||
| Upper bounds of the -weight matrix coefficients | ||
| Lower bounds of the -weight matrix coefficients |
| Description of Parameters | Symbols | Values |
|---|---|---|
| Discount factor | ||
| Soft update rate | 1 × 10−2 | |
| Learning rate | 2 × 10−4 | |
| Number of network layers | ||
| Number of neurons | ||
| Activation function | ||
| Optimizer | ||
| Maximum number of episodes | 2 × 103 | |
| Maximum iterations of each step | 1.5 × 102 | |
| Compensatory controller weight factor | ||
| Simple size per batch | ||
| Replay buffer |
| Speed Conditions | -Weight Matrices |
|---|---|
| 320 km/h | |
| 340 km/h | |
| 360 km/h | |
| 380 km/h |
| Speeds | Schemes | Pct. Decline | Pct. Decline | ||
|---|---|---|---|---|---|
| 320 km/h | Passive Control | 108.47 | - | 92.86 N | - |
| LQR | 48.38 | 55.40% | 41.12 N | 55.71% | |
| SAC | 37.20 | 65.71% | 27.76 N | 70.10% | |
| CPO-LQR | 34.38 | 68.30% | 28.86 N | 68.93% | |
| CPO-LQR-CtrlLaw | 30.07 | 72.28% | 25.52 N | 72.52% | |
| ECCPO-LQR | 32.24 | 70.28% | 27.62 N | 70.25% | |
| CPO-LQR-SAC | 24.33 | 77.58% | 19.94 N | 78.52% | |
| CPO-LQR-BC-SAC | 22.77 | 79.01% | 18.68 N | 79.87% | |
| 340 km/h | Passive Control | 137.79 | - | 118.98 N | - |
| LQR | 52.62 | 61.81% | 44.62 N | 62.51% | |
| SAC | 45.33 | 67.10% | 32.18 N | 72.95% | |
| CPO-LQR | 39.48 | 71.35% | 32.86 N | 72.38% | |
| CPO-LQR-CtrlLaw | 31.43 | 77.19% | 27.10 N | 77.22% | |
| ECCPO-LQR | 34.01 | 75.32% | 30.24 N | 74.58% | |
| CPO-LQR-SAC | 28.79 | 79.11% | 23.98 N | 79.85% | |
| CPO-LQR-BC-SAC | 27.08 | 80.35% | 22.44 N | 81.15% | |
| 360 km/h | Passive Control | 137.72 | - | 118.68 N | - |
| LQR | 54.54 | 60.40% | 48.70 N | 58.96% | |
| SAC | 51.46 | 62.63% | 39.26 N | 66.92% | |
| CPO-LQR | 38.25 | 72.22% | 32.90 N | 72.27% | |
| CPO-LQR-CtrlLaw | 32.23 | 76.60% | 27.12 N | 77.15% | |
| ECCPO-LQR | 34.40 | 75.02% | 30.12 N | 74.61% | |
| CPO-LQR-SAC | 33.80 | 75.46% | 28.06 N | 76.35% | |
| CPO-LQR-BC-SAC | 30.98 | 77.50% | 26.52 N | 77.66% | |
| 380 km/h | Passive Control | 144.75 | - | 126.06 N | - |
| LQR | 52.44 | 63.78% | 44.68 N | 64.56% | |
| SAC | 52.05 | 64.04% | 42.04 N | 66.66% | |
| CPO-LQR | 41.65 | 71.23% | 35.92 N | 71.51% | |
| CPO-LQR-CtrlLaw | 32.71 | 77.40% | 27.24 N | 78.39% | |
| ECCPO-LQR | 36.18 | 75.01% | 30.86 N | 75.53% | |
| CPO-LQR-SAC | 34.10 | 76.44% | 28.70 N | 77.23% | |
| CPO-LQR-BC-SAC | 31.68 | 78.11% | 26.48 N | 78.99% |
| Speeds | Methods | Pct. Decline | |
|---|---|---|---|
| 320 km/h | VFPID [19] | 51.86 | 32.88% |
| PH∞ [18] | 35.64 | 7.38% | |
| PPO [18] | 34.51 | 10.31% | |
| CB-DMRL [18] | 32.82 | 14.71% | |
| IDDPG [19] | 42.98 | 45.12% | |
| CPO-LQR-BC-SAC (ours) | 22.77 | 79.01% | |
| 340 km/h | VFPID [19] | - | - |
| PH∞ [18] | 33.81 | 11.16% | |
| PPO [18] | 32.94 | 13.47% | |
| CB-DMRL [18] | 31.62 | 16.93% | |
| IDDPG [19] | - | - | |
| CPO-LQR-BC-SAC (ours) | 27.08 | 80.35% | |
| 360 km/h | VFPID [19] | 52.25 | 35.59% |
| PH∞ [18] | 42.83 | 12.41% | |
| PPO [18] | 41.10 | 15.93% | |
| CB-DMRL [18] | 38.52 | 21.22% | |
| IDDPG [19] | 43.99 | 45.76% | |
| CPO-LQR-BC-SAC (ours) | 30.98 | 77.50% | |
| 380 km/h | VFPID [19] | - | - |
| PH∞ [18] | 64.13 | 15.21% | |
| PPO [18] | 57.94 | 23.40% | |
| CB-DMRL [18] | 48.65 | 35.69% | |
| IDDPG [19] | - | - | |
| CPO-LQR-BC-SAC (ours) | 31.68 | 78.11% |
| Parameters | DSA350S | SSS400+ |
|---|---|---|
| (kg) | 6.4 | 6.1 |
| (kg) | 7 | 10.2 |
| (kg) | 12 | 10.3 |
| (N·m−1) | 2650 | 10,400 |
| (N·m−1) | 10,000 | 10,600 |
| (N·m−1) | 0 | 0 |
| (N·s·m−1) | 100 | 10 |
| (N·s·m−1) | 100 | 0 |
| (N·s·m−1) | 70 | 120 |
| Types | Methods | Pct. Decline | |
|---|---|---|---|
| DSA380 | Passive Control | 108.47 | - |
| CPO-LQR | 34.38 | 68.30% | |
| CPO-LQR-BC-SAC (ours) | 22.77 | 79.01% | |
| DSA350S | Passive Control | 52.07 | - |
| CPO-LQR | 25.70 | 50.65% | |
| CPO-LQR-BC-SAC (ours) | 24.41 | 53.13% | |
| SSS400+ | Passive Control | 108.56 | - |
| CPO-LQR | 40.25 | 62.92% | |
| CPO-LQR-BC-SAC (ours) | 38.27 | 64.74% |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
Share and Cite
Han, Z.; Feng, Q.; Liu, W.; Liu, Y.; Yang, H.; Li, H.; Xu, M.; Xiao, S. An Optimized Active Compensation Control Framework for High-Speed Railway Pantograph via Imitation-Guided Deep Reinforcement Learning. Machines 2025, 13, 769. https://doi.org/10.3390/machines13090769
Han Z, Feng Q, Liu W, Liu Y, Yang H, Li H, Xu M, Xiao S. An Optimized Active Compensation Control Framework for High-Speed Railway Pantograph via Imitation-Guided Deep Reinforcement Learning. Machines. 2025; 13(9):769. https://doi.org/10.3390/machines13090769
Chicago/Turabian StyleHan, Zhun, Qingsheng Feng, Wangyang Liu, Yuqi Liu, Hangtao Yang, Hong Li, Mingxia Xu, and Shuai Xiao. 2025. "An Optimized Active Compensation Control Framework for High-Speed Railway Pantograph via Imitation-Guided Deep Reinforcement Learning" Machines 13, no. 9: 769. https://doi.org/10.3390/machines13090769
APA StyleHan, Z., Feng, Q., Liu, W., Liu, Y., Yang, H., Li, H., Xu, M., & Xiao, S. (2025). An Optimized Active Compensation Control Framework for High-Speed Railway Pantograph via Imitation-Guided Deep Reinforcement Learning. Machines, 13(9), 769. https://doi.org/10.3390/machines13090769

