Reinforcement Learning for Real-Time Control Using Quanser Platforms: A Structured Narrative Review
Abstract
1. Introduction
- Synthesise the experimental findings reported in the selected literature according to the RL algorithm, control objective, hardware configuration, and reported performance.
- Critically assess recurring limitations related to hardware–software integration, computational latency, sample efficiency, safety, reproducibility, and simulation-to-hardware transfer.
- Identify research opportunities for improving Quanser-based RL evaluation through enhanced sensing, modular multi-agent configurations, standardised machine learning interfaces, safety-aware control mechanisms, and higher-fidelity digital twins.
- A comprehensive overview and structured classification of published research employing Quanser systems for reinforcement learning (RL)-based real-time control.
- A comparative synthesis of the RL algorithms, control objectives, hardware platforms, and experimentally observed outcomes reported in the reviewed primary studies.
- A critical assessment of recurring challenges related to hardware–software integration, computational latency, sample efficiency, safety constraints, reproducibility, and simulation-to-hardware transfer.
- A platform-level comparison of the sensing, actuation, software-interface, and multi-agent capabilities of the Quanser Aero, Aero 2, 3-DOF Helicopter, and Autonomous Vehicles Research Studio (AVRS) systems.
- A research roadmap for improving the suitability of Quanser platforms as reproducible RL benchmarks through advanced sensing, modular architectures, standardised software interfaces, safety-aware RL, and higher-fidelity digital twins.

- Note: All Figures containing product images were obtained from the official Quanser website and are reproduced solely for scholarly identification and discussion of the reviewed experimental platforms. In accordance with the Quanser Terms and Conditions (https://www.quanser.com/terms-conditions/, accesssed on 8 September 2026), which permit website content to be used as a resource when appropriate copyright notices, references, and credits are provided, the original source is acknowledged in each applicable figure caption. All product names, images, trademarks, and logos remain the property of Quanser Inc., and their inclusion in this review does not imply sponsorship or endorsement.
2. Overview of Quanser Platforms
2.1. Quanser Aero
2.2. Quanser Aero 2
2.3. Quanser 3-DOF Helicopter Model
2.4. Quanser Autonomous Vehicle Research Studio (AVRS)
2.5. Comparison of Hardware, Sensors, and APIs
3. Reinforcement Learning for Control: Background
3.1. Analytical Taxonomy of RL-Based Control Studies
3.2. Evolution of RL Algorithms for Control
3.3. Simulation-Based, Sim2real and Hardware-Based Training
4. Review of RL-Based Research Contributions Using Quanser Products
4.1. Study-Level Evidence Base
4.2. Factors Limiting Published RL Experiments on AVRS
4.3. Cross-Comparison of RL Algorithms
5. Capabilities of Quanser Platforms for RL
6. Limitations and Challenges
6.1. Simulator-to-Reality Gap: Evidence from Quanser Studies
6.2. Limited Scalability for Multi-Agent Learning
6.3. Hardware and Safety Constraints
6.4. Academic Scope and Engineering Transferability
6.5. Computational Challenges in Real-Time Deployment
6.6. Constructive Solutions for the Limitations
- Bridging the Simulator-to-Reality Gap: Methods including domain randomisation, system identification, and physics-informed digital twins can mitigate the discrepancies between simulated and physical settings.
- Improving Multi-Agent Scalability: Implementing lightweight communication protocols, modular testbed expansions, and edge-computing technologies can improve the scalability of AVRS for extensive multi-agent reinforcement learning studies.
- Reducing Hardware Stress: Hybrid training methodologies, wherein policies are initially established in simulation and subsequently refined on hardware, might mitigate wear and tear. Safety-conscious reinforcement learning and confined exploration frameworks can enhance hardware protection during the learning process.
- Addressing Computational Demands: Utilising optimised neural architectures, hardware accelerators (such as GPUs and TPUs), and real-time inference frameworks helps mitigate latency challenges in embedded reinforcement learning deployment.
6.7. Recommended Experimental Protocol for Reproducible and Safe RL Validation
- Hardware characterisation and baseline control: Calibrate the platform and verify encoder offsets, actuator directions, command limits, sampling rate, communication delay, and emergency-stop operation. Use safe, low-amplitude excitation to estimate friction, dead zones, thrust gains, coupling, and delays. Implement PID, LQR, or MPC as both a performance baseline and, where appropriate, a backup controller.
- Environment and reward specification: Report the observation vector, action limits, reward terms, termination conditions, reference signals, control frequency, filtering, normalisation, and safety constraints. Simulated observations should reproduce the noise, resolution, delay, filtering, and state availability of the physical system. Individual reward components should be provided to allow reconstruction.
- Simulation training and tuning: Tune hyperparameters in simulation using a documented search procedure and multiple random seeds. Report interactions, episodes, wall-clock time, convergence criteria, and variability. For Sim2Real transfer, vary model parameters, sensor noise, latency, friction, actuator gains, and payload during training.
- Safety-gated hardware deployment: Before deployment, test the policy under uncertainty, disturbances, noise, delay, saturation, and unseen references. Begin hardware trials with reduced state and action limits. Action clipping, rate limits, software end stops, watchdog timers, automatic termination, and a backup controller should operate independently of the RL policy.
- Evaluation and reproducibility reporting: Evaluate using predefined, unseen trajectories and disturbances. Compare against at least one conventional controller under identical conditions. Report repeated-trial statistics, tracking error, overshoot, settling time, control effort, constraint violations, safety interventions, inference time, hardware interaction time, and failures. Where possible, release code, configurations, model parameters, random seeds, trained policies, and hardware/software versions.
7. Future Recommendations and Directions
7.1. Platform-Specific Digital Twins and Benchmarking
7.2. Sensing, Communication, and AVRS Multi-Agent Reproducibility
7.3. Open Software, Safety, and Computational Infrastructure
8. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Dyvik, M.; Fjereide, D.E.; Rotondo, D. Modeling and identification of the Quanser Aero using a detailed description of friction and centripetal forces. In Proceedings of the Scandinavian Simulation Society Conference, Vasteras, Sweden, 25–28 September 2023; pp. 246–253. [Google Scholar]
- Rotondo, D.; Sanchez, H.S. Experiences with using Kahoot! in control theoretical courses. In Proceedings of the European Control Conference (ECC), Stockholm, Sweden, 25–28 June 2024; pp. 2678–2684. [Google Scholar]
- Iza, J.; Paredes, E.; Herrera, M.; Benítez, D.; Pérez-Pérez, N.; Camacho, O. Real-time experimental benchmarking of control strategies for a coupled 2-DOF helicopter. Eng 2026, 7, 170. [Google Scholar] [CrossRef] [Scilit]
- Pereda Perez, G. Modeling and Control Using Feedback Linearization of a Quanser Aero 2 Device. Master’s Thesis, Universitat Politecnica de Catalunya, Barcelona, Spain, 2024. [Google Scholar]
- Kumar, S.; Dewan, L. A comparative analysis of LQR and SMC for Quanser Aero. In Control and Measurement Applications for Smart Grid: Selected Papers of SGESC; Springer: Singapore, 2022; pp. 453–463. [Google Scholar]
- Segerstrom, E.; Podlaski, M.; Khare, A.; Vanfretti, L. Parameter optimization and model validation of Quanser Aero using Modelica and RaPId. In Proceedings of the AIAA/IEEE Electric Aircraft Technology Symposium (EATS), Denver, CO, USA, 11–13 August 2021; pp. 1–9. [Google Scholar]
- Kumar, S.; Dewan, L. Set-point tracking of Quanser Aero using SMC in the presence of uncertainties. In Proceedings of the International Conference on Intelligent Computing and Control Systems (ICICCS), Madurai, India, 6–8 May 2021; pp. 1595–1601. [Google Scholar]
- Quanser, I. Quanser Real-Time Control (QUARC): User Documentation. Available online: https://docs.quanser.com/quanser-sdk/documentation/quarc.html (accessed on 26 August 2026).
- MathWorks. Control a Quanser QUBE Pendulum with a Raspberry Pi Using Reinforcement Learning. Available online: https://www.mathworks.com/help/reinforcement-learning/ug/use-reinforcement-learning-to-control-quanser-qube-pendulum-via-raspberry-pi.html (accessed on 26 August 2026).
- Polzounov, K.; Redden, L.; Sundar, R. Blue River Controls: A Toolkit for Reinforcement Learning Control Systems on Hardware. In Proceedings of the Workshop on Deep Reinforcement Learning, 33rd Conference on Neural Information Processing Systems, Vancouver, BC, Canada, 14 December 2019. [Google Scholar] [CrossRef] [Scilit]
- Kumawat, G.; Goswami, N.K.; Vajpai, J. Design of fuzzy controller for tracking of desired trajectory of 2-DOF Aero system. In Proceedings of the IEEE Power India International Conference (PIICON), Jaipur, India, 10–12 December 2024; pp. 1–6. [Google Scholar]
- Volpi, V. Study on Electric Propulsion Solutions for Vertical Takeoff and Landing Air Vehicles. Master’s Thesis, Politecnico di Milano, Milan, Italy, 2021. [Google Scholar]
- Yang, S.; Xi, L.; Hao, J.; Wang, W. Aerodynamic-parameter identification and attitude control of quad-rotor model with CIFER and adaptive LADRC. Chin. J. Mech. Eng. 2021, 34, 1. [Google Scholar] [CrossRef] [Scilit]
- Labdai, S.; Chrifi-Alaoui, L.; Drid, S.; Delahoche, L.; Bussy, P. Real-time implementation of an optimized fractional sliding mode controller on the Quanser-Aero helicopter. In Proceedings of the International Conference on Control, Automation and Diagnosis (ICCAD), Paris, France, 7–9 October 2020; pp. 1–6. [Google Scholar]
- Lopes, A.N.D.; Arcese, L.; Guelton, K.; Cherifi, A. Sampled-data controller design with application to the Quanser Aero 2-DOF helicopter. In Proceedings of the IEEE International Conference on Automation, Quality and Testing, Robotics (AQTR), Cluj-Napoca, Romania, 21–23 May 2020; pp. 1–6. [Google Scholar]
- AlHamouch, A.; Tuqan, M.; Bardawil, C.; Daher, N. Investigating performance of adaptive and robust control schemes for Quanser Aero. In Proceedings of the International Conference on Advanced Computational Tools for Engineering Applications (ACTEA), Zouk Mosbeh, Lebanon, 3–5 July 2019; pp. 1–6. [Google Scholar]
- Mehndiratta, M.; Kayacan, E. Receding horizon control of a 3 DOF helicopter using online estimation of aerodynamic parameters. Proc. Inst. Mech. Eng. Part G J. Aerosp. Eng. 2018, 232, 1442–1453. [Google Scholar] [CrossRef] [Scilit]
- Rojas-Cubides, H.; Cortes-Romero, J.; Coral-Enriquez, H.; Rojas-Cubides, H. Sliding mode control assisted by GPI observers for tracking tasks of a nonlinear multivariable twin-rotor aerodynamical system. Control Eng. Pract. 2019, 88, 1–15. [Google Scholar] [CrossRef] [Scilit]
- Arabi, E.; Yucelen, T. Experimental results with the set-theoretic model reference adaptive control architecture on an aerospace testbed. In Proceedings of the AIAA Scitech Forum, San Diego, CA, USA, 7–11 January 2019; p. 0930. [Google Scholar]
- Li, H.; Luo, P.; Li, Z.; Zhu, G.; Zhang, X. Finite time-adaptive full-state quantitative control of quadrotor aircraft and QDrone experimental platform verification. Drones 2024, 8, 351. [Google Scholar] [CrossRef] [Scilit]
- Chen, Z.; Zhong, W.; Xie, S.; Zhang, Y.; Yuen, C. Observer-based robust integral reinforcement learning for attitude regulation of quadrotors. Knowl. Based Syst. 2024, 303, 112360. [Google Scholar] [CrossRef] [Scilit]
- Abro, G.E.M.; Abdallah, A.M.; Elshaar, M.E. Swarm coordination and trajectory tracking in quadrotor UAVs using fractional-order PID control strategy. IEEE Trans. Autom. Sci. Eng. 2025, 23, 2995–3008. [Google Scholar] [CrossRef] [Scilit]
- Liu, X.; Yuan, Z.; Gao, Z.; Zhang, W. Reinforcement learning-based fault-tolerant control for quadrotor UAVs under actuator fault. IEEE Trans. Ind. Inform. 2024, 20, 13926–13935. [Google Scholar] [CrossRef] [Scilit]
- Borbolla-Burillo, P.; Sotelo, D.; Frye, M.; Garza-Castanon, L.E.; Juarez-Moreno, L.; Sotelo, C. Design and real-time implementation of a cascaded model predictive control architecture for unmanned aerial vehicles. Mathematics 2024, 12, 739. [Google Scholar] [CrossRef] [Scilit]
- Yang, P.; Xuan, Y.; Li, W. Adaptive nonsingular fast-reaching terminal sliding mode control based on observer for aerial robots. Actuators 2024, 13, 98. [Google Scholar] [CrossRef] [Scilit]
- Muthusamy, P.K.; Suthar, B.; Muthusamy, R.; Garratt, M.; Pota, H.; Seneviratne, L.; Zweiri, Y. Self-organizing BFBEL control system for a UAV under wind disturbance. IEEE Trans. Ind. Electron. 2023, 71, 5021–5033. [Google Scholar] [CrossRef] [Scilit]
- Li, S.; Chen, Z.; Zhang, Q.; Hu, T.; Zhu, B. Maneuver synchronization of networked rotating platforms using historical nominal command. Control Eng. Pract. 2024, 153, 106081. [Google Scholar] [CrossRef] [Scilit]
- Schafer, G.; Rehrl, J.; Huber, S.; Hirlaender, S. Comparison of model predictive control and proximal policy optimization for a 1-DOF helicopter system. In Proceedings of the IEEE International Conference on Industrial Informatics (INDIN), Beijing, China, 17–20 August 2024; pp. 1–7. [Google Scholar]
- Fellag, R.; Belhocine, M. 2-DOF helicopter control via state feedback and full/reduced-order observers. In Proceedings of the International Conference on Electrical Engineering and Automation Control (ICEEAC), Setif, Algeria, 12–14 May 2024; pp. 1–6. [Google Scholar]
- Schafer, G.; Schirl, M.; Rehrl, J.; Huber, S.; Hirlaender, S. Python-based reinforcement learning on Simulink models. In International Conference on Soft Methods in Probability and Statistics; Springer: Cham, Switzerland, 2024; pp. 449–456. [Google Scholar]
- Rezoug, A.; Messah, A.; Messaoud, W.A.; Baizid, K.; Iqbal, J. Adaptive-optimal MIMO nonsingular terminal sliding mode control of twin-rotor helicopter system. J. Braz. Soc. Mech. Sci. Eng. 2024, 46, 162. [Google Scholar] [CrossRef] [Scilit]
- Schlanbusch, S.M.; Zhou, J. Adaptive predictor-based control for a helicopter system with input delays: Design and experiments. J. Autom. Intell. 2024, 3, 50–56. [Google Scholar] [CrossRef] [Scilit]
- Ouerdane, F.; Mysorewala, M.F. Visual servoing of a 3 DOF hover quadcopter using 2D markers. In Proceedings of the IEEE Symposium on Industrial Electronics (ISIE), Ulsan, Republic of Korea, 18–21 June 2024; pp. 1–6. [Google Scholar]
- Heemels, W.P.M.H.; Johansson, K.H.; Tabuada, P. An introduction to event-triggered and self-triggered control. In Proceedings of the IEEE Conference on Decision and Control (CDC), Maui, HI, USA, 10–13 December 2012; pp. 3270–3285. [Google Scholar]
- Amin, R.U.; Li, A. Modelling and robust attenuation tracking control of 3-DOF four rotor hover vehicle. Aircr. Eng. Aerosp. Technol. 2017, 89, 87–98. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Jiang, B.; Lu, N.; Pan, J. Hybrid modeling based double-granularity fault detection and diagnosis for quadrotor helicopter. Nonlinear Anal. Hybrid. Syst. 2016, 21, 22–36. [Google Scholar] [CrossRef] [Scilit]
- Abro, G.E.M.; Abdallah, A.M.; Elshaar, M.E. Helical trajectory control of quadrotor UAVs using fractional-order PID controller. In Proceedings of the IEEE International Conference on Automation Science and Engineering (CASE), Bari, Italy, 28 August–1 September 2024; pp. 2085–2090. [Google Scholar] [CrossRef] [Scilit]
- Sanz, R.; Garcia, P.; Zhong, Q.-C.; Albertos, P. Predictor-based control of a class of time-delay systems and its application to quadrotors. IEEE Trans. Ind. Electron. 2016, 64, 459–469. [Google Scholar] [CrossRef] [Scilit]
- Haddad, A.G.; Boiko, I.; Zweiri, Y. Fuzzy ensembles of reinforcement learning policies for robotic systems with varied parameters. arXiv 2023, arXiv:2311.05655. [Google Scholar]
- Haddad, A.G.; Boiko, I.; Zweiri, Y. Reinforcement learning generalization for nonlinear systems through dual-scale homogeneity transformations. arXiv 2023, arXiv:2311.05013. [Google Scholar]
- Stauffer, L.; Manjunath, P.; Kim, D.; Korpela, C. Tactical autonomous maneuver testbed for multi-agent air-ground teams. In Proceedings of the International Conference on Electrical, Computer, Communications and Mechatronics Engineering (ICECCME), Tenerife, Spain, 19–21 July 2023; pp. 1–6. [Google Scholar]
- Ullah, N.; Mehmood, Y.; Aslam, J.; Ali, A.; Iqbal, J. UAVs-UGV leader follower formation using adaptive non-singular terminal super twisting sliding mode control. IEEE Access 2021, 9, 74385–74405. [Google Scholar] [CrossRef] [Scilit]
- Sun, H.; Li, J.; Wang, R.; Yang, K. Attitude control of the quadrotor UAV with mismatched disturbances based on fractional-order sliding mode and backstepping control subject to actuator faults. Fractal Fract. 2023, 7, 227. [Google Scholar] [CrossRef] [Scilit]
- Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A.A.; Veness, J.; Bellemare, M.G.; Graves, A.; Riedmiller, M.; Fidjeland, A.K.; Ostrovski, G.; et al. Human-level control through deep reinforcement learning. Nature 2015, 518, 529–533. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lillicrap, T.P.; Hunt, J.J.; Pritzel, A.; Heess, N.; Erez, T.; Tassa, Y.; Silver, D.; Wierstra, D. Continuous control with deep reinforcement learning. In Proceedings of the International Conference on Learning Representations, San Juan, Puerto Rico, 2–4 May 2016. [Google Scholar]
- Fujimoto, S.; van Hoof, H.; Meger, D. Addressing function approximation error in Actor–Critic methods. In Proceedings of the 35th International Conference on Machine Learning, Stockholm, Sweden, 10–15 July 2018; Volume 80, pp. 1587–1596. [Google Scholar]
- Haarnoja, T.; Zhou, A.; Abbeel, P.; Levine, S. Soft Actor–Critic: Off-policy maximum-entropy deep reinforcement learning with a stochastic actor. In Proceedings of the 35th International Conference on Machine Learning, Stockholm, Sweden, 10–15 July 2018; Volume 80, pp. 1861–1870. [Google Scholar]
- Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal policy optimization algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar] [CrossRef] [Scilit]
- Lowe, R.; Wu, Y.; Tamar, A.; Harb, J.; Abbeel, P.; Mordatch, I. Multi-agent Actor–Critic for mixed cooperative–competitive environments. In Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA, 4–9 December 2017; Volume 30. [Google Scholar]
- Kumar, A.; Zhou, A.; Tucker, G.; Levine, S. Conservative Q-learning for offline reinforcement learning. Adv. Neural Inf. Process. Syst. 2020, 33, 1179–1191. [Google Scholar]
- Zhang, Q.; Wang, H.; Cai, Y.; Xie, W.-F.; Sun, X.; Chen, L. Stability-constrained coordinated control strategy for vehicle chassis integrated AFS and DYC via reinforcement learning. Control Eng. Pract. 2026, 175, 107122. [Google Scholar] [CrossRef] [Scilit]
- Wang, H.; Zhang, W.; Xu, Y.; Li, H.; Ren, P. WaterCycleDiffusion: Visual-textual fusion empowered underwater image enhancement. Inf. Fusion 2026, 127, 103693. [Google Scholar] [CrossRef] [Scilit]
- Ma, G.; Wu, H.; Zhao, Z.; Zou, T.; Hong, K.-S. Adaptive neural network control for a nonlinear 2-DOF helicopter system with prescribed performance. IET Control Theory Appl. 2023, 17, 1789–1799. [Google Scholar] [CrossRef] [Scilit]
- Schlanbusch, S.M.; Zhou, J. Adaptive quantized control of uncertain nonlinear rigid body systems. Syst. Control Lett. 2023, 175, 105513. [Google Scholar] [CrossRef] [Scilit]
- Kim, S.-K.; Ahn, C.K. Performance-Boosting Attitude Control for 2-DOF Helicopter Applications via Surface Stabilization Approach. IEEE Trans. Ind. Electron. 2022, 69, 7234–7243. [Google Scholar] [CrossRef] [Scilit]
- Rinciog, A.; Meyer, A. Fabricatio-RL: A Reinforcement Learning Simulation Framework for Production Scheduling. In Proceedings of the 2021 Winter Simulation Conference (WSC), Phoenix, AZ, USA, 15–17 December 2021; pp. 1–12. [Google Scholar] [CrossRef] [Scilit]
- Zhao, Z.; He, W.; Mu, C.; Zou, T.; Hong, K.-S.; Li, H.-X. Reinforcement learning control for a 2-DOF helicopter with state constraints. IEEE Trans. Autom. Sci. Eng. 2022, 21, 157–167. [Google Scholar] [CrossRef] [Scilit]
- Jia, J.; Guo, K.; Yu, X.; Guo, L.; Xie, L. Reliability based LQR fault-tolerant control for a quadrotor UAV. In Advances in Guidance, Navigation and Control: Proceedings of ICGNC 2020; Springer: Singapore, 2021; pp. 4471–4481. [Google Scholar]
- Liao, T.; Haridevan, A.; Liu, Y.; Shan, J. Autonomous vision-based UAV landing with collision avoidance using deep learning. In Science and Information Conference; Springer: Cham, Switzerland, 2022; pp. 79–87. [Google Scholar]
- Matthews, M.T. Adaptive and Neural Network Based Control of Unmanned Aerial Vehicles. Ph.D. Dissertation, North Carolina A&T State University, Greensboro, NC, USA, 2021. [Google Scholar]
- Wahbah, M.; Chehadeh, M.; Zweiri, Y. Dynamic based estimator for UAVs with real-time identification using DNN and the modified relay feedback test. arXiv 2021, arXiv:2106.07299. [Google Scholar] [CrossRef] [Scilit]
- Alkayas, A.; Chehadeh, M.; Ayyad, A.; Zweiri, Y. Systematic online tuning of multirotor UAVs for accurate trajectory tracking under wind disturbances and in-flight dynamics changes. IEEE Access 2022, 10, 6798–6813. [Google Scholar] [CrossRef] [Scilit]
- Ayyad, A.; Chehadeh, M.; Silva, P.H.; Wahbah, M.; Hay, O.A.; Boiko, I.; Zweiri, Y. Multirotors from takeoff to real-time full identification using the modified relay feedback test and deep neural networks. IEEE Trans. Control Syst. Technol. 2021, 30, 1561–1577. [Google Scholar] [CrossRef] [Scilit]
- Abro, G.E.M.; Abdallah, A.M. Digital twins and control theory: A critical review on revolutionizing quadrotor UAVs. IEEE Access 2024, 12, 43291–43307. [Google Scholar] [CrossRef] [Scilit]
- Fandel, A.; Birge, A.; Miah, M.S. Development of reinforcement learning algorithm for 2-DOF helicopter model. In Proceedings of the IEEE Symposium on Industrial Electronics (ISIE), Cairns, QLD, Australia, 12–15 June 2018; pp. 553–558. [Google Scholar]
- Zhao, Z.; Weng, Y.; Liu, Z.; Liu, Y.; Hong, K.-S. Integral reinforcement learning control of an uncertain 2-DOF helicopter system with input quantization and state constraints. IEEE Trans. Ind. Electron. 2025, 72, 9250–9259. [Google Scholar] [CrossRef] [Scilit]
- Chen, T.; Shan, J. A novel cable-suspended quadrotor transportation system: From theory to experiment. Aerosp. Sci. Technol. 2020, 104, 105974. [Google Scholar] [CrossRef] [Scilit]
- Durdevic, P.; Ortiz-Arroyo, D. A deep neural network sensor for visual servoing in 3D spaces. Sensors 2020, 20, 1437. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Durdevic, P.; Ortiz-Arroyo, D.; Li, S.; Yang, Z. Vision aided navigation of a quad-rotor for autonomous wind-farm inspection. IFAC-PapersOnLine 2019, 52, 61–66. [Google Scholar] [CrossRef] [Scilit]
- Pitarch, J.L.; Sala, A. Multicriteria fuzzy-polynomial observer design for a 3DoF nonlinear electromechanical platform. Eng. Appl. Artif. Intell. 2014, 30, 96–106. [Google Scholar] [CrossRef] [Scilit]
- Chen, F.; Lu, F.; Jiang, B.; Tao, G. Adaptive compensation control of the quadrotor helicopter using quantum information technology and disturbance observer. J. Frankl. Inst. 2014, 351, 442–455. [Google Scholar] [CrossRef] [Scilit]
- Cavalca, M.S.M.; Kienitz, K.H. Application of TFL/LTR robust control techniques to failure accommodation. In Proceedings of the 20th International Congress of Mechanical Engineering, Gramado, Brazil, 15–20 November 2009; pp. 1–8. [Google Scholar]
- Bermudez-Ortega, J.; Besada-Portas, E.; Lopez-Orozco, J.A.; Chacon, J.; de la Cruz, J.M. Developing web and TwinCAT PLC-based remote control laboratories for modern web-browsers or mobile devices. In Proceedings of the IEEE Conference on Control Applications (CCA), Buenos Aires, Argentina, 19–22 September 2016; pp. 810–815. [Google Scholar]
- Chacon, J.; Besada-Portas, E.; Garcia-Perez, L.; Lopez-Orozco, J.A. An integrated framework for the agile development and deployment of low cost remote laboratories. Multimed. Tools Appl. 2025, 84, 29207–29227. [Google Scholar] [CrossRef] [Scilit]
- Maraoui, S.; Bouzrara, K. ARX model decomposed on Meixner-like orthonormal bases. ISA Trans. 2019, 95, 278–294. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Arabi, E.; Yucelen, T. A set-theoretic model reference adaptive control architecture with dead-zone effect. Control Eng. Pract. 2019, 89, 12–29. [Google Scholar] [CrossRef] [Scilit]
- Gruenwald, B.C.; Yucelen, T.; Muse, J.A. Direct uncertainty minimization in model reference adaptive control: Experimental results. In Proceedings of the AIAA Scitech Forum, San Diego, CA, USA, 7–11 January 2019; p. 2186. [Google Scholar]
- Lambert, P.; Reyhanoglu, M. Observer-based sliding mode control of a 2-DOF helicopter system. In Proceedings of the IECON—Annual Conference of the IEEE Industrial Electronics Society, Washington, DC, USA, 21–23 October 2018; pp. 2596–2600. [Google Scholar]
- Sadi, M.A.; Jamali, A.; Kamaruddin, A.M.N.A.; Jun, V.Y.S. Optimizing UAV performance in turbulent environments using cascaded model predictive control algorithm and Pixhawk hardware. J. Braz. Soc. Mech. Sci. Eng. 2025, 47, 396. [Google Scholar] [CrossRef] [Scilit]
- Mushitha, L.; Kumar, K.K.A.; Priyadharshini, S. Robust and optimal control of Quanser Aero. In Proceedings of the International Conference on Advancements in Electrical, Electronics, Communication, Computer and Automation (ICAECA), Coimbatore, India, 4–5 April 2025; pp. 1–6. [Google Scholar]
- Sadi, M.A.; Jamali, A.; Kamaruddin, A.M.N.A.; Jun, V.Y.S. Cascade model predictive control for enhancing UAV quadcopter stability and energy efficiency in wind turbulent mangrove forest environment. e-Prime-Adv. Electr. Eng. Electron. Energy 2024, 10, 100836. [Google Scholar] [CrossRef] [Scilit]
- Chiem, N.X. Synthesis of an orbit tracking controller for a 2DOF helicopter based on sequential manifolds with stabilization time in the presence of disturbances. Eng. Technol. Appl. Sci. Res. 2024, 14, 15083–15089. [Google Scholar] [CrossRef] [Scilit]
- Zhou, W.; Zhou, L.; Yuan, T.; Chen, R.; Liu, D. Robust performance optimization of UAV dynamic systems using MPC-PID hybrid control. Sci. Rep. 2026, 16, 2585. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Abdelkader, K.; Kais, B. Robust H∞ gain neuro-adaptive observer design for nonlinear uncertain systems. Trans. Inst. Meas. Control 2019, 41, 2293–2309. [Google Scholar] [CrossRef] [Scilit]
- Mohamed, S.I.A. Hybrid Active Force Control for Fixed Based Rotorcraft. Ph.D. Dissertation, Universiti Teknologi Malaysia, Skudai, Malaysia, 2022. [Google Scholar]
- Abdelmaksoud, S.I.; Mailah, M.; Hing, T.H. System enhancement on perturbations and wind gusts for twin-rotor helicopter using intelligent active force control. Int. J. Model. Identif. Control 2023, 43, 166–176. [Google Scholar] [CrossRef] [Scilit]
- Abdelmaksoud, S.I.; Mailah, M.; Abdallah, A.M. Enhancing disturbance rejection capability and body jerk performance of a twin-rotor helicopter model using intelligent active force control. J. Mek. 2021, 44, 1–20. [Google Scholar]
- Reyhanoglu, M.; Jafari, M.; Rehan, M. Simple learning-based robust trajectory tracking control of a 2-DOF helicopter system. Electronics 2022, 11, 2075. [Google Scholar] [CrossRef] [Scilit]
- Feng, Y.; Zhou, Y.; Ho, H.W. Reinforcement learning based robust tracking control for unmanned helicopter with state constraints and input saturation. Aerosp. Sci. Technol. 2024, 155, 109549. [Google Scholar] [CrossRef] [Scilit]
- Quanser. Academic Institutions Worldwide. Available online: https://www.quanser.com/community/our-customers/ (accessed on 28 August 2026).
- Schäfer, G.; Rehrl, J.; Huber, S.; Hirlaender, S. Safe reinforcement learning using ideas from model predictive control. arXiv 2026, arXiv:2607.07252. [Google Scholar] [CrossRef] [Scilit]
- Schäfer, G.; Rehrl, J.; Huber, S. Integrating physics-informed neural networks for safe reinforcement learning in a 1-DoF helicopter system. In International Conference on Database and Expert Systems Applications; Springer Nature: Cham, Switzerland, 2026; pp. 107–111. [Google Scholar] [CrossRef] [Scilit]
- Rajappa, S.; Chriette, A.; Chandra, R.; Khalil, W. Modelling and dynamic identification of 3 DOF Quanser helicopter. In Proceedings of the 16th International Conference on Advanced Robotics (ICAR), Montevideo, Uruguay, 25–29 November 2013. [Google Scholar] [CrossRef] [Scilit]
- Wu, T.; Acharya, S.; Khalil, A.; Aljanaideh, A.F.; Al Janaideh, M.; Kundur, D. Multi-head attention machine learning for fault classification in mixed autonomous and human-driven vehicle platoons. In Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), London, UK, 29 May–2 June 2023; pp. 10040–10046. [Google Scholar] [CrossRef] [Scilit]
- Shiyas, A.; Rao, S. Design of planar collision-free trochoidal paths for a multi-robot swarm. Eur. J. Control 2025, 81, 101143. [Google Scholar] [CrossRef] [Scilit]
- Fareh, R.; Baziyad, M.; Khadraoui, S.; Brahmi, B.; Bettayeb, M. Logarithmic potential field: A new leader–follower robotic control mechanism to enhance the execution speed and safety attributes. IEEE Access 2023, 11, 85451–85466. [Google Scholar] [CrossRef] [Scilit]
- Bal, L.; Mbakop, S.; Espindola-Winck, G.; Sueur, C.; Merzouki, R. Cooperative curve-based synchronized control of a fleet of autonomous robots. IEEE/ASME Trans. Mechatron. 2025, 30, 2900–2909. [Google Scholar] [CrossRef] [Scilit]
- Gąsieniec, L.; Kuszner, Ł.; Latif, E.; Parasuraman, R.; Spirakis, P.G.; Stachowiak, G. Brief Announcement: Anonymous Distributed Localisation via Spatial Population Protocols. In Proceedings of the 4th Symposium on Algorithmic Foundations of Dynamic Networks (SAND 2025), Liverpool, UK, 9–11 June 2025; Volume 330, pp. 19:1–19:5. [Google Scholar] [CrossRef]
- Rapalski, A.; Dudzik, S. Energy consumption analysis of navigation algorithms for wheeled mobile robots. Energies 2023, 16, 1532. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Z.; Bian, J.; Wu, K. Relay-switching-based fixed-time tracking controller for nonholonomic state-constrained systems. IEEE/CAA J. Autom. Sin. 2022, 10, 1778–1780. [Google Scholar] [CrossRef] [Scilit]
- Gao, S.; Zhang, H.; Wang, Z.; Huang, C.; Yan, H. Optimal injection attack strategy for cyber-physical systems under resource constraint. IEEE Trans. Control Netw. Syst. 2022, 10, 636–646. [Google Scholar] [CrossRef] [Scilit]
- Wu, Y.; Zuo, Z.; Han, Q.; Wang, Y.; Yang, H. Formation control of wheeled mobile robots with multiple virtual leaders under communication failures. IEEE Trans. Control Syst. Technol. 2022, 31, 295–305. [Google Scholar] [CrossRef] [Scilit]
- Tassanbi, A.; Iskakov, A.; Do, T.D.; Ali, M.H. Interactive real-time leader follower control system for UAV and UGV. In Proceedings of the International Conference on Electrical, Computer, Communications and Mechatronics Engineering (ICECCME), Male, Maldives, 16–18 November 2022; pp. 1–8. [Google Scholar]






| Parameter | Value |
|---|---|
| Device mass | 3.6 kg |
| Device height (ground to top of base) | 45 cm |
| Helicopter body mass | 1.39 kg |
| Helicopter body length | 48 cm |
| Base dimensions (W × L) | 17.5 cm × 17.5 cm |
| Encoder resolution (in quadrature) | 512 counts/rev |
| Pitch angle range | ° (±°) |
| Yaw angle range | ° |
| Motor/propeller force–thrust constant | 0.119 N/V |
| Motor/propeller torque–thrust constant | 0.0036 Nm/V |
| Propeller diameter | 12.7 cm |
| Propeller pitch | 15.2 cm |
| Motor armature resistance | 0.83 |
| Motor current–torque constant | 57.7 mN.m/A |
| Parameter | Value |
|---|---|
| Device Dimensions (D × W × H) | 18 cm × 52 cm × 40 cm |
| Operating Space (D × W × H) | 52 cm × 52 cm × 62 cm |
| Mass | 4.7 kg |
| Pitch Angle Range | ° (±° from horizontal) |
| Yaw Angle Range | ° continuous |
| Pitch Encoder Resolution | 2880 counts/revolution |
| Yaw Encoder Resolution | 4096 counts/revolution |
| Prop Thrust Constant | N·s/rad |
| Inertial Thrust Constant | 0.042 Nm/A |
| Inertial Measurement Unit (IMU) | IIM-42652 compact six-axis MEMS device |
| Tri-axis Gyroscope Range | ±500 dps |
| Tri-axis Accelerometer Range | ±2 g |
| Parameter | Value |
|---|---|
| Device Mass | 6.2 kg |
| Device Height (ground to top of base) | 45 cm |
| Device Length (counterweight to front of propellers) | 127 cm |
| Base Dimensions (W × L) | 17.5 cm × 17.5 cm |
| Pitch Encoder Resolution (quadrature mode) | 4096 counts/rev |
| Travel Encoder Resolution (quadrature mode) | 8192 counts/rev |
| Pitch Angle Range | ±° |
| Elevation Angle Range | ° |
| Travel Angle Range | ° |
| Platform | Hardware Features | Sensors | API/Software Support | ROS/Multi-Agent Support |
|---|---|---|---|---|
| Quanser Aero | Lightweight (3.46 kg), compact base (17.5 cm × 17.5 cm), 1–2 DOF (pitch, yaw), dual-rotor testbed | High-resolution encoder (8192 counts/rev) for pitch and yaw | Seamless with MATLAB/Simulink, QUARC | Not natively ROS; primarily single-agent, suitable for introductory RL |
| Quanser Aero 2 | Moderate size (4.7 kg), enhanced modular design, 2 DOF (pitch, yaw), improved mechanical structure | Encoders: pitch (2880 counts/rev) and yaw (4096 counts/rev); onboard six-axis MEMS IMU (IIM-42652), gyro (±500 dps), accel (±2 g) | MATLAB/Simulink, QUARC, Python APIs | Python-friendly; partial ROS integration possible; single-agent RL with extended workflows |
| Quanser 3-DOF Helicopter | Heavier (6.2 kg), larger footprint (127 cm length), 3 DOF (pitch, elevation, travel), gimbaled pivot for attitude control | Encoders: pitch (4096 counts/rev), travel (8192 counts/rev), and wide angular ranges; lacks modern onboard IMU | MATLAB/Simulink, QUARC | Not designed for ROS; focus on single-vehicle control validation; limited RL scalability |
| AVRS (QDrone/QBot) | Full-stack lab setup with aerial (QDrone/QDrone 2) and ground (QBot/QBot 2e) vehicles; requires larger workspace; onboard compute (Jetson Xavier NX/Intel Aero) | Rich sensor suite: vision cameras, IMUs, depth sensors, localisation cameras, and Wi-Fi-based swarm networking | MATLAB/Simulink, QUARC, Python, C++, ROS integration | ROS-native, multi-agent ready; supports swarm robotics, vision-based RL, and scalable experiments |
| Dimension | Category | Typical Methods | Objective | Key Evaluation Criteria |
|---|---|---|---|---|
| Control task | Stabilisation | PPO, SAC, DDPG, TD3 | Regulate attitude or position | Error, overshoot, settling time |
| Tracking | PPO, DDPG, TD3, SAC | Follow references or paths | Tracking error, effort, constraints | |
| Energy-aware | Multi-objective RL, PPO, SAC | Balance accuracy and energy | Energy–accuracy trade-off | |
| Fault-tolerant | Robust/constrained RL | Operate under faults or disturbances | Robustness, safety, saturation | |
| Multi-agent | MADDPG, MAPPO, QMIX | Coordinate vehicles | Scalability, communication, safety | |
| Algorithm | Value-based | Q-learning, DQN | Discrete action control | Discretisation, scalability, bias |
| Policy gradient | REINFORCE, PPO | Direct policy optimisation | Stability, interaction cost | |
| Deterministic Actor–Critic | DDPG, TD3 | Continuous control | Exploration, critic bias, reuse | |
| Stochastic Actor–Critic | SAC | Entropy-based continuous control | Exploration, efficiency, tuning | |
| Offline RL | CQL, IQL, BCQ | Learn from fixed datasets | Coverage, unseen actions | |
| Multi-agent RL | MADDPG, MAPPO, QMIX | Cooperative or competitive control | Non-stationarity, credit assignment | |
| Training | Simulation | Any RL method | Train without hardware | Model fidelity, hardware validation |
| Sim2Real | PPO, SAC, DDPG, TD3 | Transfer policy to hardware | Transfer loss, model mismatch | |
| Online hardware | Adaptive/integral RL | Learn directly on the testbed | Safety, wear, time, repeatability | |
| Hybrid | Pretraining + refinement | Simulate, transfer, and fine-tune | Robustness versus hardware cost |
| Ref. | Year | Platform | Task | RL Method | Training Mode | Reported Outcome |
|---|---|---|---|---|---|---|
| [65] | 2018 | Aero, 2-DOF | Pitch–yaw stabilisation | RL controller | Simulation | Feasible coupled control; quantitative hardware result NR. |
| [57] | 2022 | 2-DOF Helicopter | Constrained tracking | Robust Actor–Critic RL | Simulation + hardware | Improved tracking while satisfying state constraints. |
| [39] | 2023 | QDrone with load | Robust trajectory tracking | Fuzzy RL ensemble | Sim2Real | The 3D RMSE decreased to 0.0343 m and 0.0524 m under wind. |
| [40] | 2023 | QDrone with load | Load-position control | DDPG with homogeneity transformation | Sim2Real | Reported 96% success and 0.0253 m 3D RMSE. |
| [28] | 2024 | Aero 2, 1-DOF | Pitch tracking | PPO vs. MPC/LQR | Simulation + hardware fine-tuning | Mean error: ° simulation, ° hardware, and ° after fine-tuning. |
| [30] | 2024 | Aero 2, 1-DOF | Reference tracking | PPO | Sim2Real | Best simulation mean deviation was °, demonstrated by hardware transfer. |
| [66] | 2025 | 2-DOF Helicopter | Quantised constrained tracking | Integral Actor–Critic RL | Simulation + hardware | Tracking achieved under input quantisation and state constraints. |
| Gap Source | Platform | Transfer Effect | Mitigation | Evidence |
|---|---|---|---|---|
| Motor dead zone and nonlinear thrust | Aero, Aero 2, 3-DOF Helicopter | Ineffective small actions, tracking error, and oscillations | Identify dead zones, model nonlinear thrust, and randomise parameters | Addressed in control studies; isolated RL evidence is unavailable. |
| Encoder quantisation and missing states | All platforms; Aero 2 transfer study | Noisy observations, poor velocity estimates, and action switching | Use filtering, observers, measurement history, and realistic sensor models | State redesign is shown in [28]; related methods appear in [51,52,66]. |
| Aerodynamic coupling | Aero/Aero 2 (2-DOF) and 3-DOF Helicopter | Unintended cross-axis motion and altered response | Use coupled nonlinear models, domain randomisation, and residual adaptation | Modelling is reported in [1,66]; RL ablation evidence is unavailable. |
| Clearance, backlash, and friction | Aero and 3-DOF mechanical joints | Delayed motion, hysteresis, and limit cycles | Identify mechanical effects, randomise parameters, and adapt online | Friction is modelled in [1]; RL-specific evidence is limited. |
| Sensor and communication latency | AVRS and embedded systems | Delayed actions, instability, and poor coordination | Randomise latency and use timestamped observations or delay-aware policies | Matched AVRS simulation–hardware results remain scarce. |
| Different simulated and measured states | Aero 2 and sensor-limited systems | Policy requires states unavailable on hardware | Align observation spaces and simulate realistic sensors | Directly demonstrated in [28]. |
| Residual transfer mismatch | Aero 2 and other platforms | Remaining hardware tracking error | Apply limited hardware fine-tuning or residual learning | Fine-tuning improved performance in [28], but required extensive hardware use. |
| Method | Key Parameters | Simulation Starting Point | Main Hardware Risk |
|---|---|---|---|
| PPO | Learning rate, rollout, batch size, clipping, entropy, GAE | ; ; ; clip ; batch = 64–256 | Large updates or high entropy may cause abrupt actions. |
| SAC | Learning rates, replay buffer, batch size, target update, entropy | ; ; ; batch | High entropy may increase unsafe exploration. |
| DDPG | Actor/critic rates, exploration noise, replay buffer, target update | Actor ; critic ; ; batch –256 | Critic errors or noise may produce saturated commands. |
| TD3 | Learning rates, target noise, policy delay, replay buffer | ; delay ; ; batch –256 | Noise must remain within physical action limits. |
| DQN | Action discretisation, exploration decay, replay buffer, target update | ; buffer –; gradual exploration decay | Coarse actions reduce resolution; fine actions increase complexity. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Abro, G.E.M.; Memon, S.A.; Tanveer, J. Reinforcement Learning for Real-Time Control Using Quanser Platforms: A Structured Narrative Review. Electronics 2026, 15, 4111. https://doi.org/10.3390/electronics15184111
Abro GEM, Memon SA, Tanveer J. Reinforcement Learning for Real-Time Control Using Quanser Platforms: A Structured Narrative Review. Electronics. 2026; 15(18):4111. https://doi.org/10.3390/electronics15184111
Chicago/Turabian StyleAbro, Ghulam E Mustafa, Sufyan Ali Memon, and Jawad Tanveer. 2026. "Reinforcement Learning for Real-Time Control Using Quanser Platforms: A Structured Narrative Review" Electronics 15, no. 18: 4111. https://doi.org/10.3390/electronics15184111
APA StyleAbro, G. E. M., Memon, S. A., & Tanveer, J. (2026). Reinforcement Learning for Real-Time Control Using Quanser Platforms: A Structured Narrative Review. Electronics, 15(18), 4111. https://doi.org/10.3390/electronics15184111

