Next Article in Journal
ArchLock: Dynamic-Target Architectural Backdoor with Correlation-Based Statistical Triggers
Previous Article in Journal
Smoothly Weighted Hybrid NMPC–LQR Control for Slope-Dependent Uphill Motion of a Two-Wheeled Self-Balancing Wheelchair
Previous Article in Special Issue
AoI- and DS-Enhanced Cooperative Search for Multi-UAV Systems Under Spatially Structured Communication Constraints
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

Reinforcement Learning for Real-Time Control Using Quanser Platforms: A Structured Narrative Review

by
Ghulam E Mustafa Abro
1,
Sufyan Ali Memon
2 and
Jawad Tanveer
3,*
1
Research Unit for Robophilosophy and Integrative Social Robotics (RISR), Aarhus University, Aarhus C, 8000 Aarhus, Denmark
2
Department of Defense and AI Systems Engineering, Sejong University, Gwangjin-gu, Seoul 05006, Republic of Korea
3
Department of Computer Science and Engineering, Sejong University, Gwangjin-gu, Seoul 05006, Republic of Korea
*
Author to whom correspondence should be addressed.
Electronics 2026, 15(18), 4111; https://doi.org/10.3390/electronics15184111
Submission received: 10 August 2026 / Revised: 30 August 2026 / Accepted: 8 September 2026 / Published: 10 September 2026

Abstract

Reinforcement learning (RL) is increasingly used for real-time control of complex dynamical systems, but its practical performance must be evaluated under hardware constraints that are often simplified in simulation. This paper presents a comprehensive review of published RL-based control studies using the Quanser Aero, Aero 2, 3-DOF Helicopter, and Autonomous Vehicles Research Studio (AVRS) platforms. The reviewed studies are compared according to the RL algorithm, control objective, hardware configuration, implementation environment, and reported experimental performance. The synthesis shows that Aero and Aero 2 are used primarily for stabilisation, trajectory tracking, and energy-aware control, whereas the 3-DOF Helicopter provides a more demanding benchmark for adaptive and Actor–Critic methods under nonlinear and coupled dynamics. AVRS offers significant potential for vision-based and multi-agent RL; however, the available experimental literature remains limited. Across the reviewed comparisons, policy-gradient and Actor–Critic methods, including PPO and SAC, generally demonstrate greater adaptability and smoother continuous-control behaviour, while conventional controllers frequently retain advantages in steady-state accuracy, computational predictability, and safety verification. Nevertheless, no RL algorithm can be identified as universally superior because the published studies employ heterogeneous reward functions, reference trajectories, sampling rates, performance measures, and hardware configurations. Recurring limitations include sample inefficiency, simulation-to-hardware discrepancies, computational latency, safety constraints, and incomplete reporting of experimental protocols. The review therefore identifies standardised evaluation procedures, reproducible reporting, safety-aware RL, interoperable software interfaces, and higher-fidelity digital twins as priorities for future research. No new experimental data are generated; the contribution is a comparative synthesis of experimental evidence reported in the literature.

1. Introduction

In the swiftly advancing fields of control systems and reinforcement learning (RL), experimental testbeds are essential instruments for connecting theoretical design with practical applications. Hardware-in-the-loop (HIL) systems, which directly connect algorithms with emulations of physical plants, offer a dependable method for validating control strategies amidst practical uncertainties such as noise, delay, and nonlinear dynamics [1,2]. Such frameworks are essential in robotics, aerospace, and autonomous systems, where simulation-based methods frequently inadequately represent the intricacies of real-world situations [3]. Throughout the years, numerous benchmark configurations, such as inverted pendulums, rotary servos, and unmanned aerial vehicle test rigs, have been employed to evaluate reinforcement learning algorithms within realistic restrictions [4].
Among the commercially available laboratory control platforms, Quanser systems are widely used in engineering education and experimental control research because they combine modular hardware, sensing, real-time interfaces, and integrated software toolchains [5]. Quanser reports that its systems have been adopted by more than 2500 academic institutions worldwide (https://www.quanser.com/community/our-customers/, accesssed on 8 September 2026). This manufacturer-reported adoption figure demonstrates broad institutional reach, but it should not be interpreted as bibliometric evidence that Quanser constitutes a universal or de facto research standard. Accordingly, this review evaluates Quanser as a widely deployed family of academic testbeds rather than claiming that it dominates all hardware-based RL research. The Quanser Aero, Aero 2, Quanser 3-DOF Helicopter, and Autonomous Vehicle Research Studio (AVRS) are examples of systems that have found widespread use in the fields of education and advanced research [6,7], as shown in Figure 1. Quanser platforms are integrated with MATLAB/SimulinkR2025b through the QUARC real-time control environment, which supports real-time execution, code generation, parameter tuning, and deployment on supported targets [8]. Furthermore, MathWorks provides a documented reinforcement learning toolbox workflow for training and deploying an RL controller on the Quanser QUBE-Servo 2 pendulum through a Raspberry Pi interface [9]. Polzounov et al. introduced Blue River Controls, an OpenAI-Gym-compatible Python framework supporting common simulation and hardware interfaces for RL experiments on the QUBE-Servo 2 [10]. These sources directly substantiate the MATLAB/Simulink, Python 3.14.7, RL-toolbox, and Gym-style interface capabilities discussed here. Despite the growing body of research on reinforcement learning, the published experimental evidence concerning the use of Quanser platforms for RL-based control has not yet been comprehensively reviewed. Existing surveys generally address broad RL benchmarks, simulation environments, or instructional control systems, but provide limited synthesis of how Quanser testbeds have been used to evaluate RL controllers under physical hardware constraints. This article addresses this gap through a structured review of previously published studies. It does not report original experiments; instead, it critically synthesises the experimental configurations, control objectives, algorithms, findings, and implementation challenges documented in the reviewed literature. Accordingly, this review aims to:
  • Synthesise the experimental findings reported in the selected literature according to the RL algorithm, control objective, hardware configuration, and reported performance.
  • Critically assess recurring limitations related to hardware–software integration, computational latency, sample efficiency, safety, reproducibility, and simulation-to-hardware transfer.
  • Identify research opportunities for improving Quanser-based RL evaluation through enhanced sensing, modular multi-agent configurations, standardised machine learning interfaces, safety-aware control mechanisms, and higher-fidelity digital twins.
For clarity, the term “experimental validation” in this review refers exclusively to experiments conducted and reported in the cited primary studies. The present article does not introduce a new RL controller, perform new hardware experiments, or generate an original experimental dataset. Its contribution is the comparative analysis and critical synthesis of the experimental evidence available in the published literature. Through this synthesis of published experimental studies, the paper evaluates the suitability of Quanser systems as adaptable benchmark platforms for RL-based real-time control and identifies the conditions under which findings obtained in simulation may be transferred to physical implementations. The principal contributions of this review are as follows:
  • A comprehensive overview and structured classification of published research employing Quanser systems for reinforcement learning (RL)-based real-time control.
  • A comparative synthesis of the RL algorithms, control objectives, hardware platforms, and experimentally observed outcomes reported in the reviewed primary studies.
  • A critical assessment of recurring challenges related to hardware–software integration, computational latency, sample efficiency, safety constraints, reproducibility, and simulation-to-hardware transfer.
  • A platform-level comparison of the sensing, actuation, software-interface, and multi-agent capabilities of the Quanser Aero, Aero 2, 3-DOF Helicopter, and Autonomous Vehicles Research Studio (AVRS) systems.
  • A research roadmap for improving the suitability of Quanser platforms as reproducible RL benchmarks through advanced sensing, modular architectures, standardised software interfaces, safety-aware RL, and higher-fidelity digital twins.
Figure 1. Quanser testbeds: (a) Quanser Aero, (b) Aero 2, (c) Quanser 3-DOF Helicopter, (d) AVRS System.
Figure 1. Quanser testbeds: (a) Quanser Aero, (b) Aero 2, (c) Quanser 3-DOF Helicopter, (d) AVRS System.
Electronics 15 04111 g001
This work was conducted as a structured narrative review. Relevant publications were identified through systematic searches of Scopus, Web of Science, and IEEE Xplore using combinations of the keywords “reinforcement learning,” “real-time control,” “Quanser,” “Quanser Aero,” “Aero 2,” “3-DOF Helicopter,” “AVRS,” “QDrone,” and “QBot.” The search encompassed publications available up to 2026. Studies were considered eligible if they investigated an RL-based control method using a Quanser platform in simulation, hardware-in-the-loop operation, or physical experiments and provided sufficient information regarding the learning algorithm, control objective, experimental platform, or reported performance. Publications unrelated to RL-based control or lacking an explicit connection to a Quanser system were excluded. Through this screening and selection process, 129 relevant research articles were identified, shortlisted, and considered in the preparation of this manuscript. The selected studies were subsequently compared according to the Quanser platform employed, RL algorithm, control task, implementation environment, and principal experimental findings. Owing to the substantial heterogeneity among the reviewed studies in terms of control tasks, performance metrics, and experimental configurations, the evidence was synthesised qualitatively rather than through a formal meta-analysis.
The remainder of this article is organised as follows: Section 2 describes the principal hardware, sensing, and software characteristics of the Quanser Aero, Aero 2, 3-DOF Helicopter, and AVRS platforms. Section 3 provides background on RL for control and discusses the distinction between simulation-based and hardware-based evaluation. Section 4 reviews and compares published RL studies using Quanser platforms. Section 5 summarises the principal capabilities of these platforms, whereas Section 6 examines their limitations and practical challenges. Section 7 presents future research recommendations. Section 8 concludes the review and summarises its principal findings.
  • Note: All Figures containing product images were obtained from the official Quanser website and are reproduced solely for scholarly identification and discussion of the reviewed experimental platforms. In accordance with the Quanser Terms and Conditions (https://www.quanser.com/terms-conditions/, accesssed on 8 September 2026), which permit website content to be used as a resource when appropriate copyright notices, references, and credits are provided, the original source is acknowledged in each applicable figure caption. All product names, images, trademarks, and logos remain the property of Quanser Inc., and their inclusion in this review does not imply sponsorship or endorsement.

2. Overview of Quanser Platforms

This section evaluates the primary Quanser platforms (Quanser Inc., Markham, ON, Canada) utilised in reinforcement learning (RL) and control system validation [11,12,13], examining their advantages and drawbacks, and concludes with a comparative analysis [14,15,16].

2.1. Quanser Aero

Quanser Aero, a small rotor-based testbed depicted in Figure 2, interfaces seamlessly with MATLAB/Simulink [17,18,19] and simulates quadrotor dynamics with either one (pitch-only) or two (pitch plus yaw) degrees of freedom. The system is MIMO and nonlinear, enabling experimentation with sophisticated control strategies such as LQR, sliding-mode control, and versions of reinforcement learning [20,21,22].
Furthermore, Table 1 summarises the physical and electromechanical specifications of the Quanser helicopter-based platform. The relatively low device mass (3.46 kg) and compact base dimensions make it suitable for laboratory-scale experiments. The high-resolution encoder (8192 counts/rev) enables precise measurement of pitch and yaw, critical for reinforcement learning validation [23,24,25]. The wide yaw range ( 360 °) allows for full rotational studies, while the pitch range supports nonlinear maneuver testing. Motor and propeller constants provide a realistic force–torque relationship, ensuring experimental fidelity, though the modest torque–thrust constant limits scalability to larger UAV dynamics. Overall, these specifications balance safety, realism, and repeatability, making the platform well suited for both control education and reinforcement learning research [26].
The Quanser Aero platform is concise and cost-effective, simulates authentic aerial dynamics, and is supported by a robust ecosystem of software, curriculum, and community resources [25]. Nonetheless, its highly interdependent and nonlinear dynamics render modelling and control difficult, and the restricted onboard sensing limits studies necessitating high-speed manoeuvres or vision-based reinforcement learning.

2.2. Quanser Aero 2

The Quanser Aero 2 platform depicted in Figure 3 enhances the Aero system with improved modularity, superior abstraction layers, and multi-input functionality [24,25,26,27,28,29]. The integration of Simulink models with Python-based reinforcement learning libraries was demonstrated by Schäfer et al., who trained a PPO agent using Stable-Baselines3 and transferred the resulting policy to a physical Quanser Aero 2 system [30]. Moreover, multi-objective reinforcement learning can concurrently optimise energy efficiency and tracking performance [9,27]. The Quanser Aero 2 platform has numerous benefits, such as enhanced abstraction layers, adaptable interaction with Python and Simulink, and robust interoperability with contemporary reinforcement learning frameworks. Nonetheless, its effectiveness is still limited by hardware constraints regarding control precision and velocity, while the sample efficiency of physical reinforcement learning training frequently remains suboptimal.
Table 2 [28] delineates the characteristics of the Quanser platform, emphasising its mechanical dimensions, operational envelope, sensor capabilities, and motor parameters. The gadget, weighing 4.7 kg and measuring 18 cm × 52 cm × 40 cm , is appropriate for laboratory settings, providing adequate operational space of 52 cm × 52 cm × 62 cm . The pitch and yaw angle ranges ( 90 ° and 360 ° continuous, respectively) provide comprehensive attitude control studies. High-resolution encoders (2880 counts/rev for pitch and 4096 counts/rev for yaw) provide precise state feedback, essential for reinforcement learning and control validation. The thrust constants ensure realistic actuation dynamics, while the onboard IIM-42652 six-axis MEMS IMU (TDK InvenSense, San Jose, CA, USA) offers dependable motion sensing. The tri-axis gyroscope (±500 dps) and accelerometer (±2 g) enhance the platform’s sensing accuracy, facilitating both traditional control and data-driven reinforcement learning applications.

2.3. Quanser 3-DOF Helicopter Model

The Quanser 3-DOF Helicopter model shown in Figure 4 is a tri-axial rotor system that facilitates roll, pitch, and yaw, affixed to a gimbaled pivot for accurate attitude control [29]. It has been widely utilised in observer design, controller validation, and fault-tolerant control research [31,32,33,34,35,36,37,38]. The main advantages consist of realistic attitude dynamics, high-fidelity encoders, and appropriateness for verifying sophisticated control schemes. Nonetheless, the platform exhibits specific constraints: it does not incorporate thruster-based translational motion characteristic of complete UAVs, its modelling may overlook aerodynamic influences and friction [12,38] and its relevance to multi-agent reinforcement learning contexts is restricted.
According to Table 3, the Quanser 3-DOF Helicopter system, a twin-rotor airborne testbed, is intended for laboratory-scale validation of sophisticated control and reinforcement learning algorithms. The device has a weight of 6.2 kg, a height of 45 cm, and an overall length of 127 cm from counterweight to rotors, while featuring a compact base footprint of 17.5 cm × 17.5 cm , rendering it appropriate for restricted laboratory settings. It features high-precision encoders, providing 4096 counts/rev for pitch measurement and 8192 counts/rev for journey measurement, facilitating meticulous tracking of angular positions. The technology facilitates a pitch angle range of ± 32 °, an elevation range of 63.5 °, and a complete 360 ° journey range, thereby ensuring authentic aerial dynamics. The parameters together indicate that the platform is lightweight, small, and precise, offering extensive angular freedom, hence providing an exceptional experimental setting for secure and reproducible testing of control and reinforcement learning systems.

2.4. Quanser Autonomous Vehicle Research Studio (AVRS)

The Autonomous Vehicles Research Studio (AVRS), depicted in Figure 5, incorporates airborne and terrestrial vehicles, specifically the QDrone/QDrone 2 and QBot/QBot 2e, alongside a ground control station, localisation cameras, and networking equipment [13,25]. The QDrone features onboard computing capabilities such the Jetson Xavier NX or Intel Aero and vision sensors, and facilitates networked swarm operations [34,39,40]. AVRS, as a comprehensive platform, offers significant benefits: it enables multi-agent and vision-based reinforcement learning, is reasonably easy to deploy, and is ideally suited for educational and research applications. Nonetheless, the platform exhibits constraints, such as elevated costs, increased spatial demands, and the intricacies of sensor fusion and safety considerations, which may hinder scalability. Moreover, the reproducibility of investigations may be influenced by dependencies on wireless connectivity and localisation systems [16].

2.5. Comparison of Hardware, Sensors, and APIs

A comparative investigation of the Quanser Aero, Aero 2, 3-DOF Helicopter, and Autonomous Vehicles Research Studio (AVRS) identifies unique strengths and weaknesses in hardware, sensing capabilities, and software integration. The Aero and Aero 2 are compact and lightweight, ideal for desktop laboratory research, while the 3-DOF Helicopter is larger and heavier, providing greater motion ranges. Conversely, AVRS constitutes a comprehensive platform encompassing both airborne and terrestrial vehicles, necessitating greater specialised areas and resulting in elevated expenses. Aero offers fundamental high-resolution encoders for pitch and yaw measurements, but Aero 2 improves this configuration with an integrated six-axis MEMS IMU, facilitating more precise attitude estimation. The 3-DOF Helicopter features dual high-fidelity encoders for accurate pitch and travel angle measurements, but it does not possess contemporary onboard inertial sensor capabilities. AVRS provides the most comprehensive sensor suite, incorporating cameras, IMUs, depth sensors, and localisation systems, thus facilitating multi-agent vision-based reinforcement learning experiments.
Table 4 shows all platforms integrating effortlessly with MATLAB/Simulink and Quanser’s QUARC real-time control software regarding APIs and ecosystem support. Aero and the 3-DOF Helicopter predominantly depend on this integration, whereas Aero 2 enhances capability through Python APIs, facilitating reinforcement learning workflows beyond MATLAB. AVRS offers exceptional flexibility, ensuring interoperability with MATLAB/Simulink, QUARC, Python, ROS, and C++ libraries, rendering it optimal for multi-agent robotics and swarm research. Aero and Aero 2 are optimal for introductory reinforcement learning and control experiments due to their cost-effectiveness and accessibility. The 3-DOF Helicopter offers highly realistic single-vehicle dynamics for controller validation, whereas AVRS connects academic testbeds with real-world multi-robot systems, facilitating state-of-the-art reinforcement learning validation in swarm and vision-based applications.

3. Reinforcement Learning for Control: Background

Reinforcement learning (RL) has developed into a robust paradigm for control, wherein agents acquire decision-making skills through interaction with an environment to optimise cumulative reward, generally articulated within the Markov Decision Process formalism [29,41]. This has facilitated the development of learning policies that manage complicated, unpredictable, or nonlinear dynamics without the need for explicit modelling in control systems. Initial applications concentrated on conventional control methodologies; however, recent advancements in deep reinforcement learning have empowered agents to function directly from high-dimensional sensory inputs across various domains, including robotics and industrial process control [42,43].

3.1. Analytical Taxonomy of RL-Based Control Studies

RL-based control studies vary considerably in their objectives, algorithms, and implementation environments. To support consistent comparison, this review classifies the literature along three complementary dimensions: (i) control task, (ii) RL algorithm paradigm, and (iii) training and deployment mode. This framework is used to organise the Quanser studies reviewed in Section 4.
The first dimension defines the control task. Attitude stabilisation regulates variables such as pitch, yaw, elevation, or travel and is evaluated using steady-state error, overshoot, rise time, settling time, and disturbance rejection. Trajectory tracking considers time-varying references or spatial paths and is assessed through tracking error, transient response, control effort, and constraint satisfaction. Energy-aware control additionally considers actuator usage or power consumption. Fault-tolerant control addresses disturbances, modelling uncertainty, actuator faults, saturation, and state constraints. Multi-agent coordination covers formation control, cooperative navigation, collision avoidance, swarm behaviour, and coordinated aerial–ground operation. The second dimension concerns the RL paradigm. Value-based methods, such as Q-learning and DQN, estimate state–action values and are mainly suited to discrete actions. Policy-gradient methods directly optimise stochastic or continuous policies. Actor–Critic methods combine a policy-generating actor with a value-estimating critic and include deterministic off-policy methods (DDPG and TD3), stochastic off-policy methods (SAC), and on-policy methods (PPO). Offline RL learns from fixed datasets, while multi-agent RL extends value-based or Actor–Critic learning to cooperative or competitive systems. The third dimension describes the training and deployment mode. Pure-simulation studies offer high throughput and low risk but do not establish physical performance. Sim2Real studies train policies in simulation before deploying them on hardware, either directly or with limited adaptation. Pure hardware online learning captures real dynamics but raises concerns regarding safety, wear, reset time, and sample efficiency. Hybrid training combines simulation pretraining with controlled hardware refinement, balancing safety and adaptation. These dimensions are related but distinct. The same task may use different algorithms and training modes; for example, stabilisation may employ simulation-trained PPO with hardware transfer or an adaptive Actor–Critic method trained directly on the testbed. Likewise, multi-agent coordination may use centralised training with decentralised execution in simulation or on AVRS hardware. Performance comparisons must therefore consider the control task, algorithm, and training mode together rather than attributing outcomes to the RL algorithm alone.
Table 5 summarises the analytical framework. In Section 4, each reviewed study is classified according to these dimensions wherever the original publication provides sufficient information. Algorithms included as part of the broader RL control taxonomy are distinguished from algorithms with verified experimental evidence on the reviewed Quanser platforms. Therefore, inclusion of DQN, DDPG, TD3, SAC, offline RL, or multi-agent RL in this taxonomy does not imply that each method has already been physically validated on Aero, Aero 2, the 3-DOF Helicopter, or AVRS.

3.2. Evolution of RL Algorithms for Control

The technical evolution of reinforcement learning for control has progressed from tabular value-learning methods to deep value approximation, continuous-action actor–critic algorithms, and more recent safety-aware, offline, and multi-agent formulations. Early Q-learning methods were primarily applicable to problems with relatively small discrete state and action spaces and therefore scaled poorly to high-dimensional continuous-control systems. Deep Q-networks (DQN) addressed part of this limitation by combining Q-learning with deep neural function approximation, experience replay, and a target network [44]. Nevertheless, DQN still requires a discrete action set. For real-time control applications, discretising continuous quantities such as motor voltages, torques, or thrust commands may reduce control resolution, while using a fine discretisation can produce a rapidly expanding action space. Deep deterministic policy gradient (DDPG) extended deep RL to continuous-action problems by combining a deterministic actor with an off-policy critic and a replay buffer [45]. Its ability to reuse previously collected transitions makes it relevant to continuous control, where physical interaction may be expensive. However, DDPG can be sensitive to exploration noise, critic approximation error, hyperparameter selection, and overestimation of action values. Twin delayed deep deterministic policy gradient (TD3) was subsequently introduced to mitigate these limitations through clipped double-Q learning, delayed actor updates, and target-policy smoothing [46]. These mechanisms reduce critic overestimation and improve the stability of deterministic continuous-control learning.
Soft Actor–Critic (SAC) introduced an alternative off-policy approach based on a stochastic maximum-entropy objective [47]. By optimising both expected return and policy entropy, SAC encourages exploration while retaining the sample-reuse advantages of off-policy learning. Proximal policy optimisation (PPO) followed a different approach as an on-policy policy-gradient method that constrains policy changes using a clipped optimisation objective [48]. Its relative implementation simplicity and stable policy updates have contributed to its adoption in robotic and real-time control research. However, because PPO is on-policy and cannot reuse older interactions as extensively as DDPG, TD3, or SAC, it may require a larger number of newly collected transitions. The development of RL control has also expanded beyond single-agent unconstrained optimisation. Multi-agent Actor–Critic methods enable cooperative or competitive decision-making among interacting agents by incorporating centralised information during training and decentralised policies during execution [49]. Such approaches are relevant to cooperative navigation, vehicle formations, and the aerial–ground coordination tasks supported by multi-vehicle environments such as AVRS. Offline RL provides another important direction by learning policies from fixed datasets rather than requiring unrestricted online exploration. Conservative Q-learning, for example, regularises value estimation to reduce the selection of actions that are insufficiently represented in the training dataset [50].
This paradigm is particularly relevant to physical control systems, for which unsafe exploration, hardware wear, and experimental reset time can restrict online learning. More recent RL control research has placed greater emphasis on stability, safety, and integration with established control architectures. For example, Zhang et al. developed a stability-constrained RL-based coordinated-control strategy for the integrated active front-steering and direct-yaw-control problem of a vehicle chassis [51]. This study illustrates the broader evolution from unconstrained reward optimisation to RL formulations that explicitly incorporate closed-loop stability and coordinated actuator behaviour. Although this application does not employ a Quanser platform, it provides relevant context for the development of safety- and stability-aware RL in real-time control. Within the Quanser literature examined in this review, PPO and adaptive Actor–Critic formulations currently have clearer physical-platform evidence than DQN, DDPG, TD3, SAC, multi-agent Actor–Critic methods, and offline RL. The latter algorithms are nevertheless included because they represent mainstream developments in RL-based control and provide important reference points for identifying gaps in Quanser-based experimental validation. Their inclusion as algorithmic context should not be interpreted as evidence that each method has already been experimentally validated on the reviewed Quanser platforms. This taxonomy is used in Section 4 to distinguish algorithms that are broadly established in RL control from algorithms for which physical Quanser-based evidence is identified. Inclusion of an algorithm in the taxonomy does not imply that it has already been experimentally validated on every reviewed Quanser platform.
Recent advances in adjacent learning-based perception tasks also illustrate the increasing importance of multimodal and generative models for autonomous systems. For example, WaterCycleDiffusion integrates visual–textual information with a diffusion model to improve underwater images, demonstrating how enhanced visual perception can support downstream scene understanding and decision-making in visually degraded environments [52]. Although this method is not based on reinforcement learning, such perception-oriented developments are relevant to vision-enabled autonomous platforms, where the reliability of sensory observations can influence the state information supplied to an RL controller.

3.3. Simulation-Based, Sim2real and Hardware-Based Training

The validation modes considered in this review are distinguished as follows: Simulation-only evaluation executes both the learned controller and plant model in a virtual environment without a real-time requirement. Software-in-the-loop (SIL) evaluation connects the implemented policy software to a simulated plant and is primarily used to verify interfaces and algorithm execution. Hardware-in-the-loop (HIL) evaluation executes part of the control system on real-time computing or interface hardware while the controlled plant remains simulated. Hardware inference deploys a previously trained policy on the physical Quanser plant without updating its parameters. Hardware fine-tuning begins with a simulation-trained policy and subsequently updates it through limited physical interaction, whereas direct online hardware learning trains or adapts the policy on the physical plant from the outset. Hybrid validation combines two or more of these stages. These modes expose controllers to progressively different sources of uncertainty and should therefore not be treated as equivalent evidence of real-world performance. Research in reinforcement learning for control is mostly divided between simulation-based investigations and hardware-focused applications. Simulation-centric reinforcement learning facilitates swift experimentation and algorithmic refinement, leveraging low-risk, high-throughput settings. Benchmark suites, including the continuous control tasks in MuJoCo and OpenAI Gym, along with frameworks like Safe-Control-Gym that combine safety-certified control with reinforcement learning, have facilitated methodological advancements while guaranteeing reproducibility. Nonetheless, simulation frequently inadequately represents real-world intricacies, prompting an increasing focus on hardware-in-the-loop and real-robot reinforcement learning research. Benchmark studies conducted on physical robots reveal issues related to hyperparameter sensitivity and reproducibility inherent to actual configurations [53]; analogous initiatives encompass real-robot offline reinforcement learning datasets that assist in identifying discrepancies between simulation and reality [54].
Benchmarking and reproducibility are crucial for achieving significant advancements in control-oriented reinforcement learning. The community has reacted with open-source platforms and standardised protocols, including RL Unplugged for offline reinforcement learning, SURREAL for manipulation benchmarking, and Open RL Benchmark for comparative experiment tracking. These initiatives facilitate the transparent assessment of RL methodologies and expedite technology validation. Recent research underscores the evaluation of reinforcement learning (RL) algorithms within real-time restrictions, acknowledging that computational latencies might impair control efficacy in time-sensitive applications, and advocates for the comparison of RL with model-based control methodologies utilising standardised robotic platforms. In process control, reinforcement learning (RL) has progressed, with benchmark studies assessing various RL algorithms (e.g., fourteen versions across six benchmark settings) to determine their appropriateness for particular control scenarios [55]. Moreover, frameworks in operations research, such as Fabricatio-RL, seek to connect scheduling management issues with clearly specified reinforcement learning interfaces and reproducibility standards [56]. Ultimately, although simulation provides scalability and safety, hardware implementations deliver realism and reveal essential restrictions like delays, noise, and reproducibility. Dependable benchmarks and standardised frameworks, encompassing simulation, hybrid, and physical domains, are essential for the objective progression of RL control research.
Figure 6 presents five reinforcement learning (RL) training and deployment pathways for simulation and Quanser hardware. Pathway A represents a simulation-only approach, in which the RL algorithm is trained, optimised, and evaluated entirely within a simulated environment [44,45,46,47,48,57]. Pathway B illustrates the Sim2Real workflow, where a policy trained in simulation is transferred to the physical platform for real-time inference. Pathway C extends this process through hardware fine-tuning, allowing the transferred policy to adapt using data obtained from the physical system. Pathway D depicts direct online learning, beginning with a safety-oriented initialisation and proceeding through physical interaction and continuous policy updates. Finally, Pathway E presents a hybrid approach that combines simulation, software-in-the-loop (SIL), and hardware-in-the-loop (HIL) pretraining with hardware inference and controlled adaptation. These pathways enable RL algorithms to be evaluated and refined across simulated and physical environments using Quanser testbeds, including the Aero, Aero 2, 3-DOF Helicopter, and Autonomous Vehicles Research Studio (AVRS), thereby connecting theoretical development with safe and practical implementation.

4. Review of RL-Based Research Contributions Using Quanser Products

The hardware platforms offered by Quanser have greatly aided in the use of reinforcement learning (RL) in control research. This section examines RL-based contributions across four notable Quanser testbeds—Aero, Aero 2, 3-DOF Helicopter, and AVRS—alongside comparative analyses of algorithmic performance and a synthesis of experimental findings. The Quanser Aero platform has been a foundational bed for RL research in aerial control. DRL algorithms such as PPO and SAC have been tested for pitch–yaw stabilisation and trajectory tracking in both simulation and hardware implementations. Comparative evaluations show that PPO offers faster rise time and adaptability, while classical controllers like LQR achieve superior steady-state accuracy on Aero 2 in 1-DOF settings [58]. Although Aero lacks rich onboard sensing, it remains valuable for rapid RL prototyping and education. Examining the reinforcement learning-based research on Aero 2 focusses on multi-objective reinforcement learning, especially regarding energy-efficient control architectures. A multi-objective reinforcement learning framework was established to optimise pitch tracking and energy consumption, utilising Aero 2 in 1-DOF mode, demonstrating considerable trade-offs affected by reward-shaping factors [59].

4.1. Study-Level Evidence Base

Moreover, research on problem formulation has shown that meticulous design of state-action spaces and incentive structures can significantly improve real-world reinforcement learning performance and sample efficiency when implementing policies on Aero 2 hardware [60,61,62]. Moreover, the 3-DOF Helicopter presents a more intricate and tightly interrelated dynamic environment. Researchers have utilised Actor–Critic reinforcement learning approaches, derived from heuristic dynamic programming, to attain real-time adaptive control, showcasing effective attitude stabilisation and disturbance rejection on a physical platform [62,63,64]. These model-free methodologies provide online optimisation among uncertainty, demonstrating the helicopter’s value as a demanding reinforcement learning benchmark. Lastly, AVRS provides hardware and software capabilities for investigating vision-based control, cooperative navigation, decentralised decision-making, and multi-agent RL using aerial and ground vehicles. However, the existence of these platform capabilities should not be interpreted as evidence that complete multi-agent RL pipelines have already been extensively validated on AVRS hardware. The reviewed literature predominantly addresses enabling components or conventional coordination methods, while reproducible experimental evidence concerning multi-agent RL performance remains limited. However, comparatively few peer-reviewed studies currently report complete RL training and hardware evaluation using the integrated AVRS environment. Accordingly, the available evidence is not yet sufficient for a robust comparison of RL algorithms on this platform. The following subsection examines the technical and practical factors that may have contributed to this limited publication record.
Table 6 lists the individual studies that satisfied the review’s eligibility criteria and for which a Quanser-based RL control application could be identified. Unlike the analytical taxonomy in Table 5, which classifies the wider RL control field, this table distinguishes the physical platform, control task, algorithm, training mode, hardware-validation status, and principal reported outcome of each eligible study. The evidence is concentrated on Aero/Aero 2 and 2-DOF Helicopter systems, while peer-reviewed hardware results for the 3-DOF Helicopter and AVRS remain scarce. PPO, DDPG, and adaptive or integral Actor–Critic formulations have received some physical validation, whereas comparable Quanser experiments involving DQN, TD3, SAC, offline RL, and multi-agent RL were not identified in the reviewed corpus. Because the studies employ different tasks, trajectories, reward functions, training budgets, and performance measures, their numerical results should be interpreted as study-specific evidence rather than as a controlled ranking of algorithms.
Real-time performance was also examined as a comparison dimension; however, the reviewed Quanser RL studies report timing information inconsistently. Although several studies demonstrate successful closed-loop hardware execution, most do not provide sufficient information on control-loop frequency, mean or worst-case policy-inference time, communication delay, timing jitter, processor specifications, or missed deadlines. Hardware implementation is therefore interpreted here as evidence of practical closed-loop feasibility rather than formal confirmation of deadline-compliant real-time operation. This distinction is important because a controller operating at 10 Hz and one consistently satisfying a 1 kHz deadline represent fundamentally different computational capabilities. Future studies should report the nominal sampling period, processor and software configuration, end-to-end sensing-to-actuation latency, mean and maximum inference time, jitter distribution, and missed-deadline rate over repeated experiments. The limited availability of these quantities prevents a reliable quantitative timing comparison and represents a significant reporting and benchmarking gap in current Quanser-based RL research.
The reviewed studies differ considerably in their experimental specifications. In [65], the 2-DOF Aero stabilisation problem used pitch–yaw states and motor-voltage actions, with LQR as the baseline, although hardware trials, sampling frequency, safety provisions, and code availability were not reported. The constrained Actor–Critic method in [57] used helicopter states and tracking errors to generate two motor inputs, with barrier-Lyapunov constraints providing an explicit safety mechanism during physical validation. For the QDrone suspended-load experiments, ref. [39] employed an ensemble of RL policies and reported three-dimensional RMSE values of 0.0343 m and 0.0524 m under nominal and wind-disturbance conditions, respectively, whereas the DDPG formulation in [40] reported a 96% success rate and 0.0253 m RMSE. In the Aero 2 study [28], PPO generated continuous motor commands from measurable pitch information; comparison with MPC and LQR produced mean deviations of 3.80 ° in simulation, 3.75 ° after initial hardware transfer, and 3.2 ° following hardware fine-tuning. Study [30] reported five PPO training runs of 500,000 steps and a best simulated mean deviation of 4.6 °, followed by physical deployment. Finally, ref. [66] applied integral Actor–Critic RL to constrained tracking with quantised motor inputs and reported successful hardware validation. Several publications did not disclose reward coefficients, trial counts, control frequencies, complete safety procedures, or public code and datasets; these omissions are identified as “not reported” rather than inferred and represent an important reproducibility limitation of the existing evidence.

4.2. Factors Limiting Published RL Experiments on AVRS

Although AVRS provides a technically capable environment for vision-based, multi-agent, and heterogeneous aerial–ground RL, relatively few peer-reviewed studies report complete RL experiments using the platform. Several interrelated factors may explain this scarcity. First, the entry barrier is higher than for compact, single-device testbeds such as Aero or Aero 2. A representative AVRS configuration requires multiple aerial or ground vehicles, an indoor localisation system, wireless communication infrastructure, a ground-control workstation, safety equipment, software integration, and a dedicated experimental area. Consequently, fewer laboratories can reproduce an identical physical configuration. Second, an AVRS experiment requires the simultaneous integration of several technical layers, including vehicle dynamics, onboard sensing, localisation, wireless communication, perception, inter-agent coordination, and real-time control. Failure or uncertainty in any of these layers can affect the RL policy, making system debugging and the attribution of performance considerably more difficult than in a single-agent, low-degree-of-freedom testbed. Multi-agent RL further increases the dimensionality of the state and action spaces, introduces non-stationarity as multiple agents learn or operate simultaneously, and requires the management of communication delays and synchronisation errors. Third, hardware-based RL is sample-intensive, whereas aerial experiments involve limited battery duration, repeated charging, manual resets, safety supervision, and the risk of vehicle damage during exploratory actions. These factors reduce experimental throughput and encourage researchers to conduct most policy training in simulation, followed by only limited hardware deployment. Fourth, reproducibility is complicated by differences in the number and type of vehicles, sensor calibration, localisation layout, communication network, software version, safety constraints, and physical workspace used by individual laboratories. Results may therefore be difficult to reproduce without detailed reporting and access to a comparable studio configuration. Finally, the AVRS ecosystem has continued to evolve in terms of its vehicles, onboard computing capabilities, digital twins, application programming interfaces (APIs), and product nomenclature. The time required to establish a laboratory, develop stable RL environments, conduct safe multi-agent experiments, and complete the peer-review process may create a publication lag relative to the platform’s technological development.
The available AVRS-related literature does not yet provide a mature body of directly comparable multi-agent RL experiments. Published studies more commonly demonstrate individual enabling components—such as aerial–ground coordination, leader–follower formation, vision-based navigation, localisation, or communication—than complete multi-agent RL pipelines involving centralised training, decentralised execution, hardware policy deployment, and repeated comparative evaluation. Consequently, the principal finding of this review is not that a particular multi-agent RL algorithm has been validated successfully on AVRS, but that the platform possesses relevant experimental capabilities while substantive and reproducible hardware evidence for multi-agent RL remains limited. Product functionality should therefore not be presented as equivalent to demonstrated algorithmic performance. Three technical bottlenecks are particularly important. First, wireless latency, jitter, packet loss, and unequal communication delays can cause agents to act on stale observations or neighbouring-agent states, degrading formation accuracy and violating the timing assumptions used during centralised training. Second, AVRS localisation depends on coordinated sensing and reference-frame calibration; occlusion, measurement noise, update delay, and temporary pose loss can propagate through the policies of several agents and increase collision or coordination risk. Third, collaborative scheduling becomes more difficult as the number and heterogeneity of vehicles increase because task allocation, shared-resource access, action synchronisation, episode reset, battery availability, and conflict resolution must be coordinated simultaneously. These factors interact: delayed communication can corrupt scheduling decisions, while localisation uncertainty can invalidate assigned trajectories. Future AVRS studies should therefore report end-to-end latency and jitter, packet-loss rate, localisation error and dropout frequency, synchronisation error, task-completion rate, collision or safety-intervention frequency, and performance as the number of agents increases. These factors should be interpreted as plausible and mutually reinforcing explanations derived from the platform requirements and the broader hardware-RL literature, rather than as independently verified causes. The current scarcity of publications therefore appears to reflect barriers related to accessibility, system integration, experimental throughput, and reproducibility rather than a lack of research potential.

4.3. Cross-Comparison of RL Algorithms

The reviewed studies investigate different RL families under substantially different experimental conditions. Value-based, policy-gradient, and Actor–Critic methods differ in their treatment of discrete or continuous actions, sample reuse, exploration, and policy updating; however, these methodological differences do not establish that one algorithm is inherently better suited to a particular Quanser platform. The available studies employ different control objectives, state and action definitions, reward functions, reference trajectories, training budgets, sampling rates, hardware configurations, baselines, and performance measures [67,68,69,70]. Moreover, direct controlled comparisons among DQN, PPO, SAC, DDPG, and TD3 on an identical Quanser system were not identified in the reviewed evidence. Consequently, the reported results should be interpreted as study-specific demonstrations rather than evidence of a general algorithm hierarchy. A reliable ranking would require evaluating the algorithms on the same hardware, task, training budget, safety constraints, random seeds, and evaluation protocol [71,72]. The current literature therefore supports identifying algorithm-specific opportunities and experimental gaps, but does not permit a platform-independent conclusion regarding the superiority of DQN, PPO, SAC, or other RL families [73].

5. Capabilities of Quanser Platforms for RL

Quanser platforms provide significant benefits for reinforcement learning (RL) research through the integration of high-fidelity sensors and real-time hardware-in-the-loop (HIL) capabilities. Their devices are meticulously designed for repeatability and deterministic behaviour, facilitating realistic physical testing of algorithms inside a hardware-in-the-loop research environment, as detailed in a lengthy whitepaper [74]. Furthermore, the open-architecture design enables modular reconfiguration of hardware configurations for various experiments, while high-resolution encoders, IMUs, vision sensors, and localisation systems on platforms such as QBot, AVRS, and Aero 2 significantly improve sensing accuracy. Quanser’s educational platforms, such as QBot and the Controls and Dynamics Lab, are designed for accessibility and adaptability, catering to both classroom instruction and advanced research. These technologies are augmented by comprehensive course-ware, scaffolding learning modules, and virtual twins, facilitating swift implementation and hybrid learning experiences [75,76,77,78]. Entire laboratories can be established or replicated in a virtual environment within hours, providing scalable infrastructure appropriate for both educational and research purposes.
Quanser systems exhibit exceptional adaptability with many software ecosystems. They provide support for MATLAB/Simulink and QUARC for real-time control, in addition to supporting Python APIs and ROS interoperability for control tasks and algorithm implementation [79,80,81,82]. The QBot platform offers explicit multi-language support (MATLAB, Simulink, Python, C++, ROS), facilitating easy connection with RL toolkits and expeditious control prototyping [83,84,85]. Complementary digital twins enhance accessibility by enabling experimentation with virtual representations that accurately replicate actual platforms [86,87,88,89]. Quanser reports that its products have been used by more than 2500 academic institutions worldwide [90], indicating substantial adoption in engineering education and laboratory research. However, institutional adoption does not establish the relative publication share of Quanser compared with other experimental platforms. Within the narrower corpus reviewed in this article, Quanser systems provide reusable hardware configurations and common software interfaces that can support cross-institutional replication. Their actual contribution to reproducibility nevertheless depends on complete reporting of hardware versions, model parameters, sampling rates, software configurations, reward functions, trained policies, and evaluation protocols. Figure 7 provides a comparative overview of the reported reinforcement learning (RL) capabilities across selected Quanser platforms, considering simulation, hardware inference, hardware fine-tuning or online learning, and multi-agent RL. The Aero and Aero 2 platforms exhibit evidence of both simulation-based training and hardware inference, while hardware-level adaptation has been reported only to a limited extent and no clear multi-agent RL implementation has been identified. The 2-DOF Helicopter is comparatively well established for simulation, hardware inference, and adaptive or integral RL [91,92], whereas multi-agent learning is not applicable to this single-system platform. Although simulation studies are reported for the 3-DOF Helicopter, evidence of RL deployment on physical hardware remains limited, and hardware fine-tuning has not been established. For QDrone and AVRS, simulation-based RL is reported, but evidence concerning QDrone hardware inference, online adaptation, and AVRS-based multi-agent RL remains scarce. Overall, the matrix distinguishes mature areas of RL application from capabilities that require further experimental validation, thereby identifying opportunities for future research on real-time adaptation, hardware-based learning, and cooperative multi-agent control.
These capabilities make Quanser platforms useful for repeatable proof-of-concept evaluation; however, their laboratory scale, constrained dynamics, standardised environments, and safety-protected operation limit the extent to which successful results can be generalised directly to full-scale or industrial systems. Their principal role should therefore be understood as intermediate validation between simulation and application-specific field testing.

6. Limitations and Challenges

Although Quanser systems have shown to be valuable in reinforcement learning (RL) validation, they encounter some intrinsic limits and obstacles that need to be resolved for wider applicability.

6.1. Simulator-to-Reality Gap: Evidence from Quanser Studies

The simulation-to-reality gap arises when an RL policy trained using a mathematical model or simulated plant encounters physical dynamics, measurements, and constraints that were omitted or inaccurately represented during training [30]. On Quanser platforms, this discrepancy is not limited to generic noise and delays; it can be attributed to specific electromechanical and aerodynamic characteristics, including motor and propeller dead zones, encoder quantisation, cross-axis aerodynamic coupling, friction, actuator saturation, mechanical clearance or backlash, unmeasured states, and communication latency [1,93]. These effects modify the state transitions and responses to control actions experienced by the physical system and may therefore invalidate the relationships learned by the policy in simulation. Motor dead zones and nonlinear thrust characteristics are particularly important near an equilibrium point. Simulated motor models commonly assume that thrust varies continuously or linearly with the commanded voltage. In the physical system, however, small commands may produce little or no rotor response because they must overcome motor friction and the effective thrust threshold [1,93]. Consequently, an RL policy trained using an idealised command-to-thrust relationship may generate frequent low-amplitude actions that influence the simulated plant but fail to move the physical system. Following transfer, this discrepancy may result in steady-state error, oscillatory corrective actions, or delayed responses.
Encoder quantisation introduces a different form of discrepancy. Simulated environments generally provide continuous angle and angular-velocity states, whereas a physical controller receives discrete encoder counts and may derive angular velocity through numerical differentiation or state estimation. The resulting staircase-like observations and derivative noise may cause rapid switching between neighbouring actions, particularly when a policy is sensitive to small changes in the observed state. For example, the encoder specifications of the reviewed platforms correspond to finite angular increments rather than continuous measurements [26]. Nevertheless, the practical effect of quantisation also depends on the sampling frequency, filtering method, observation normalisation, and policy architecture; therefore, it cannot be determined from encoder resolution alone. Aerodynamic coupling is especially significant in the 2-DOF Aero/Aero 2 and 3-DOF Helicopter configurations [1,93]. Variations in the speed of one rotor can influence multiple rotational axes through reaction torque, asymmetric airflow, and cross-axis forces. A policy trained using decoupled or linearised pitch and yaw dynamics may therefore produce unintended motion along an unmodelled axis when deployed on the physical hardware. This challenge is further amplified in the 3-DOF Helicopter because its pitch, elevation, and travel dynamics are strongly coupled. Mechanical clearance, joint flexibility, and backlash may similarly create a direction-dependent region in which a change in the control command does not immediately produce a corresponding measurable motion. Because these mechanical characteristics can vary with assembly condition, component wear, payload, and operating direction, they may not be represented consistently by a nominal simulation model. These platform-specific factors operate through different mechanisms but may occur simultaneously. Motor dead zones and mechanical clearance primarily distort the relationship between control actions and physical motion; encoder quantisation alters the observations supplied to the policy; and aerodynamic coupling introduces unmodelled interactions among state variables. Their combined influence explains why a controller may remain stable following transfer while still exhibiting changes in tracking error, response time, oscillatory behaviour, or control effort [28].
For Quanser systems, this gap is associated with imperfect representations of friction, rotor thrust, aerodynamic coupling, sensor resolution, state-estimation errors, actuator nonlinearities, communication delays, and safety limits. Because the significance of these effects varies across platforms, the simulation-to-reality gap should not be treated solely as a general limitation of RL. Studies involving the Quanser Aero demonstrate how plant-model fidelity can influence subsequent controller transfer. Dyvik et al. [1] showed that accurate modelling of the Aero requires explicit consideration of friction and centripetal forces, whereas Segerstrom et al. [6] employed parameter optimisation and model validation to improve the agreement between simulated and physical Aero responses. Although these studies did not directly evaluate RL policies, they identified modelling discrepancies that were particularly relevant to RL because a policy may exploit inaccuracies in the simulated dynamics during training. Consequently, a policy trained using an idealised thrust relationship or an incomplete friction model may generate actions that are effective in simulation but produce substantially different transient or steady-state behaviour on the physical testbed.
A direct RL example was reported by Schäfer et al. [28], who trained a proximal policy optimisation (PPO) controller in simulation and subsequently evaluated it on a physical Quanser Aero 2 operating in a 1-DOF configuration. Their implementation revealed an important difference between the two environments: the simulation provided direct access to the pitch angle and angular velocity, whereas the physical platform measured the pitch angle but did not directly measure the angular velocity. To improve transferability, the RL state was reformulated using measurable pitch information and pitch variation instead of relying on an unavailable velocity measurement. The simulation-trained agent achieved an average pitch deviation of approximately 3.80 ° in simulation and 3.75 ° during its initial evaluation on the physical system. Subsequent training on the physical platform reduced the reported average deviation to approximately 3.2 °; however, this improvement required an additional 500,000 interaction steps and more than 15 h of hardware operation. This example demonstrates that successful initial transfer does not eliminate the simulation-to-reality problem, as observation design, sensor availability, hardware interaction time, and post-transfer adaptation remain important practical considerations. The results reported by Schäfer et al. [28] permit an aggregate, although not factor-by-factor, assessment of policy transfer. The average pitch deviation changed from approximately 3.80 ° in simulation to 3.75 ° during the initial hardware evaluation. For this specific metric, the absolute difference between the simulation and hardware results was 0.05 °, corresponding to approximately 1.3 % of the simulation value. Thus, the initial transfer did not result in a measurable degradation in average tracking performance. Subsequent training on the physical system reduced the average deviation from 3.75 ° to approximately 3.2 °, representing an improvement of about 0.55 °, or 14.7 % , relative to the initial hardware result. However, this improvement required an additional 500,000 hardware-interaction steps and more than 15 h of physical operation. These findings quantify the aggregate transfer outcome for a specific Aero 2 configuration; they do not isolate the individual contributions of motor dead zones, encoder quantisation, aerodynamic coupling, or mechanical clearance. The reviewed literature does not report controlled RL ablation experiments in which these physical factors are introduced individually while the same trained policy is evaluated under otherwise identical conditions. Consequently, separate quantitative effects or percentages cannot currently be assigned to their respective influence on transfer performance. Encoder resolution provides a physical measure of observation quantisation, but it does not independently determine the resulting RL tracking error. Similarly, the cited modelling studies identify friction, nonlinear thrust characteristics, and coupled dynamics without quantifying their individual effects on RL-policy performance. The absence of factor-isolated experimental evidence represents an important research gap and motivates the development of standardised simulation-to-hardware evaluation protocols in which each source of discrepancy is varied independently. The manifestation of the simulation-to-reality gap differs across the reviewed platforms, as shown in Table 7.
First, experimental system identification and model validation can be used to estimate friction, thrust coefficients, inertial parameters, and cross-axis coupling before policy training [1,6]. Second, the simulated observation space should reproduce the signals that are genuinely available on hardware. Schäfer et al. [28], for example, reformulated the PPO state to avoid the dependence on an angular-velocity measurement that was unavailable on the physical Aero 2. Third, dead-zone- and quantisation-aware control studies [51,52,66,84] indicate that actuator thresholds and discrete inputs or measurements should be represented explicitly or addressed through adaptive compensation. Although not all of these studies employ RL, they identify physical effects and compensation mechanisms that can be incorporated into RL environments. Fourth, domain randomisation can vary friction, thrust gain, dead-zone width, sensor offset, quantisation, time delay, payload, and coupling coefficients during training so that the learned policy is not optimised for a single nominal model. Fifth, limited physical-system fine-tuning can compensate for residual mismatch, as illustrated in [28], although its interaction time and hardware-wear costs should be reported. Finally, physics-informed or experimentally calibrated digital twins can combine nominal plant equations with parameters obtained from measured Quanser responses. At present, however, the reviewed literature provides stronger evidence for system identification, observation alignment, and post-transfer adaptation than for controlled comparisons of domain randomisation or residual learning on Quanser hardware. Future studies should therefore compare these strategies using identical policies, test trajectories, random seeds, and pre-defined transfer metrics.
The manifestation of the simulation-to-reality gap differs across the reviewed platforms. For Aero and Aero 2, the principal sources include thrust approximation, friction, aerodynamic coupling, sensor resolution, and unmeasured states. For the 3-DOF Helicopter, the gap is intensified by strongly coupled nonlinear motion, input constraints, and interactions among the pitch, elevation, and travel dynamics. In AVRS, the transfer problem additionally encompasses vision and localisation errors, wireless latency, synchronisation uncertainty, variations among individual vehicles, and changes in the physical environment. These platform-specific differences explain why a policy-transfer procedure demonstrated using a single-axis Aero configuration cannot be assumed to generalise directly to a 3-DOF system or a multi-agent AVRS experiment.
Taken together, the available evidence indicates that Quanser Sim2Real evaluation should report at least: (i) simulation and hardware tracking error before adaptation; (ii) relative transfer degradation; (iii) success and safety-violation rates; (iv) changes in rise time, settling time, overshoot, and control effort; and (v) the number and duration of physical interactions required for adaptation. Studies should also document the dead-zone model, encoder resolution, observation filtering, actuator saturation, sampling rate, delay distribution, coupling parameters, and mechanical configuration. Reporting these quantities would permit future reviews to distinguish zero-shot transfer effectiveness from performance obtained after hardware-specific fine-tuning and would support quantitative attribution of transfer error to individual platform characteristics.

6.2. Limited Scalability for Multi-Agent Learning

The Autonomous Vehicles Research Studio (AVRS) endorses swarm robotics; nonetheless, its scalability is limited by communication bottlenecks, spatial requirements, and the intricacy of localisation systems. Experiments in multi-agent reinforcement learning using extensive teams of robots are challenging to replicate reliably due to networking delays and synchronisation problems that diminish robustness [94]. Consequently, AVRS can authenticate small-scale cooperative techniques but it encounters difficulties in facilitating large-scale swarm trials. These scalability constraints are particularly relevant to AVRS. In this platform, the difficulty of RL experimentation arises not only from the number of agents but also from the need to coordinate localisation, wireless communication, onboard perception, safety supervision, and real-time control across heterogeneous vehicles. The resulting setup and replication burden provides one explanation for why the number of complete AVRS-based RL experiments remains smaller than the literature associated with compact single-agent Quanser testbeds [65,86,87,88,89,91].

6.3. Hardware and Safety Constraints

Reinforcement learning requires exploration, which can damage actuators, sensors, and structures. Long-term training on real hardware increases wear and tear, limiting experimental cycles and maintenance costs. Safety limitations must also be applied to prevent accidents and overheating, which might limit learning space and algorithm performance [95,96,97].

6.4. Academic Scope and Engineering Transferability

Quanser platforms provide an intermediate validation stage between simulation and full-scale deployment. They enable safe, repeatable, and observable experiments, but their results should not be interpreted as direct evidence of industrial readiness. Aero and Aero 2 are fixed-base systems with one or two dominant rotational degrees of freedom. They do not fully represent free-flight translation, position–attitude coupling, payload variation, battery depletion, structural flexibility, environmental disturbances, wide operating ranges, or realistic failure consequences. The 3-DOF Helicopter captures nonlinear and cross-axis dynamics but remains mechanically constrained within bounded pitch, elevation, and travel ranges. These platforms therefore demonstrate RL control feasibility under controlled physical conditions rather than robustness in unrestricted flight. AVRS adds mobile aerial and ground vehicles, onboard perception, wireless communication, and multi-agent operation. However, experiments are generally performed indoors with structured environments, controlled lighting, assisted localisation, few vehicles, and short missions. This differs from real deployments involving weather, uncertain terrain, large communication networks, component degradation, regulatory constraints, and uncontrolled agents.
Standardised testbeds improve reproducibility but may cause policies to overfit to specific hardware, sensor calibrations, workspaces, trajectories, or disturbance ranges. Moreover, laboratory metrics such as tracking error, rise time, and cumulative reward do not fully capture reliability, certification, maintainability, energy endurance, fault coverage, or long-term operation. Quanser experiments should therefore be regarded as proof of concept or intermediate validation. Stronger transfer claims require varied parameters and disturbances, unseen payloads and trajectories, trials across multiple hardware units, long-duration operation, realistic outdoor or industrial testing, and comparison with established safety-certified controllers.

6.5. Computational Challenges in Real-Time Deployment

Real-time RL deployment faces computational bottlenecks due to the high processing requirements of deep neural policies [98,99], particularly when running on embedded processors within Quanser hardware. Latency in policy inference, especially under complex tasks or multi-agent settings, can degrade control performance and hinder real-world applicability [100,101,102,103].

6.6. Constructive Solutions for the Limitations

Therefore, to mitigate these challenges, several constructive solutions can be proposed:
  • Bridging the Simulator-to-Reality Gap: Methods including domain randomisation, system identification, and physics-informed digital twins can mitigate the discrepancies between simulated and physical settings.
  • Improving Multi-Agent Scalability: Implementing lightweight communication protocols, modular testbed expansions, and edge-computing technologies can improve the scalability of AVRS for extensive multi-agent reinforcement learning studies.
  • Reducing Hardware Stress: Hybrid training methodologies, wherein policies are initially established in simulation and subsequently refined on hardware, might mitigate wear and tear. Safety-conscious reinforcement learning and confined exploration frameworks can enhance hardware protection during the learning process.
  • Addressing Computational Demands: Utilising optimised neural architectures, hardware accelerators (such as GPUs and TPUs), and real-time inference frameworks helps mitigate latency challenges in embedded reinforcement learning deployment.
In general, despite the fact that these problems provide obstacles to scaling RL validation on Quanser platforms, ongoing advancements in simulation quality, scalable system design, and algorithmic efficiency have the potential to considerably improve their long-term sustainability as RL benchmarking and validation platforms.

6.7. Recommended Experimental Protocol for Reproducible and Safe RL Validation

A five-stage protocol is recommended to improve the safety and reproducibility of Quanser-based RL experiments. Rather than prescribing a universal algorithm, it defines the minimum procedures and reporting requirements for training, deployment, and evaluation.
  • Hardware characterisation and baseline control: Calibrate the platform and verify encoder offsets, actuator directions, command limits, sampling rate, communication delay, and emergency-stop operation. Use safe, low-amplitude excitation to estimate friction, dead zones, thrust gains, coupling, and delays. Implement PID, LQR, or MPC as both a performance baseline and, where appropriate, a backup controller.
  • Environment and reward specification: Report the observation vector, action limits, reward terms, termination conditions, reference signals, control frequency, filtering, normalisation, and safety constraints. Simulated observations should reproduce the noise, resolution, delay, filtering, and state availability of the physical system. Individual reward components should be provided to allow reconstruction.
  • Simulation training and tuning: Tune hyperparameters in simulation using a documented search procedure and multiple random seeds. Report interactions, episodes, wall-clock time, convergence criteria, and variability. For Sim2Real transfer, vary model parameters, sensor noise, latency, friction, actuator gains, and payload during training.
  • Safety-gated hardware deployment: Before deployment, test the policy under uncertainty, disturbances, noise, delay, saturation, and unseen references. Begin hardware trials with reduced state and action limits. Action clipping, rate limits, software end stops, watchdog timers, automatic termination, and a backup controller should operate independently of the RL policy.
  • Evaluation and reproducibility reporting: Evaluate using predefined, unseen trajectories and disturbances. Compare against at least one conventional controller under identical conditions. Report repeated-trial statistics, tracking error, overshoot, settling time, control effort, constraint violations, safety interventions, inference time, hardware interaction time, and failures. Where possible, release code, configurations, model parameters, random seeds, trained policies, and hardware/software versions.
Table 8 provides practical simulation-based starting points. These values are not universal settings and should be tuned for the selected platform, task, control frequency, and action scale.
The values in Table 8 are initial simulation settings commonly used in the RL literature [44,45,46,47,48] and should not be interpreted as platform-independent optimal values. Final settings must be selected according to the sampling rate, normalised observation and action ranges, control objective, and safety envelope of the specific Quanser system. Hyperparameter selection should be completed primarily in simulation, and the search space, selection criterion, random seeds, and final configuration should be reported. In addition to algorithm-specific tuning, safe and reproducible hardware deployment requires systematic management of common experimental failures. Encoder miscalibration, noisy state estimates, actuator dead zones and saturation, communication loss, control-loop overruns, localisation dropout, battery variation, and unintended policy oscillations should be detected through calibration checks, signal-validity monitoring, watchdog timers, and automatic termination criteria. Safety constraints should operate independently of the learned policy through conservative action and slew-rate limits, software end stops, state-safety envelopes, and a verified backup controller that assumes control when timing, communication, tracking, or physical limits are violated. All safety interventions, failed trials, hardware resets, and recovery actions should be recorded and reported, because excluding them can overstate policy reliability and weaken experimental reproducibility.

7. Future Recommendations and Directions

The future advancement of Quanser testbeds may concentrate on augmenting their function as dependable validation platforms for reinforcement learning (RL).

7.1. Platform-Specific Digital Twins and Benchmarking

A key research direction is the development of sophisticated digital-twin models that accurately reproduce real-world system dynamics. The integration of domain randomisation into these models could further reduce the simulation-to-reality gap by exposing RL agents to variations in system parameters, measurement noise, disturbances, and sensor latency during training. The evidence reviewed in Section 6.1 further indicates that simulation-to-reality mitigation strategies should be tailored to the specific characteristics of each Quanser platform. For Aero and Aero 2, digital twins should incorporate experimentally identified friction, thrust nonlinearities, sensor resolution, and actuator dynamics. For the 3-DOF Helicopter, simulations should accurately represent cross-axis coupling, state constraints, and input saturation. For AVRS, domain randomisation should additionally account for variations in camera and lighting conditions, localisation uncertainty, wireless communication latency, inter-vehicle variability, and environmental changes. Evaluation protocols should report performance both before and after hardware adaptation, thereby distinguishing zero-shot transfer performance from improvements achieved through subsequent fine-tuning on the physical system. Collectively, these platform-specific measures could enhance policy robustness and support safer, more reliable, and more reproducible transfer from simulation to physical hardware.
Improvement priorities should reflect the limitations of each Quanser platform. For Aero and Aero 2, standardised calibration should identify motor dead zones, thrust nonlinearities, friction, coupling, encoder offsets, and delay, enabling unit-specific digital twins and common stabilisation and tracking benchmarks. For the 3-DOF Helicopter, benchmark models should represent cross-axis coupling, saturation, quantisation, clearance, and state constraints, supported by an independent safety controller. For AVRS, timestamped communication and configurable injection of latency, packet loss, localisation noise, and pose dropout should support reproducible multi-agent tasks involving formation, waypoint allocation, collision avoidance, and aerial–ground coordination. Across all platforms, calibrated models, versioned RL environments, baseline policies, safety configurations, and common evaluation metrics should be released to enable meaningful multi-laboratory replication. These platform-specific pathways should produce measurable research outputs. Aero and Aero 2 studies should report calibrated simulation–hardware prediction error and policy-transfer degradation; 3-DOF Helicopter studies should report simultaneous three-axis tracking accuracy, constraint violations, and safety interventions; and AVRS studies should quantify latency, packet loss, localisation error, task completion, collision frequency, and scalability with agent count. Releasing the corresponding identification data, environment configurations, baseline controllers, trained policies, random seeds, and raw experimental logs would enable transparent comparison and multi-laboratory replication.
Building on the system-identification and model-validation evidence reported for Quanser platforms [1,6], the digital-twin roadmap could be extended through a predict–explain–corroborate workflow. First, an RL policy would generate control actions from synchronised hardware observations while recording the associated states, rewards, timing information, and constraint activity. Second, the policy behaviour would be interpreted by identifying the observations and learned relationships that most strongly influence its actions. Third, these interpretations would be independently examined using experimentally calibrated plant models, digital-twin predictions, known physical relationships, operating limits, and targeted hardware tests. This corroboration stage is particularly important because successful simulation-to-hardware transfer [28,30] does not establish that the policy has learned physically meaningful or causal relationships. Agreement between policy interpretations and physical-model evidence would strengthen confidence in the controller, whereas disagreement could indicate spurious correlations, inadequate model fidelity, or unsafe policy dependence. Consistent with the broader role of digital twins in aerial-control validation [64], future evaluations should report control performance together with model–hardware residuals, constraint violations, interpretation consistency, and the computational overhead of the corroboration process. This workflow is proposed as a future validation architecture rather than an experimentally established capability of the reviewed Quanser literature.

7.2. Sensing, Communication, and AVRS Multi-Agent Reproducibility

Another crucial focus is enhancing sensor accuracy and transmission bandwidth. Enhanced-resolution inertial measurement units, low-latency motion capture systems, and sophisticated vision sensors could substantially improve the precision of state estimation, facilitating more complex reinforcement learning challenges. Enhancing wireless communication infrastructure within platforms such as the Autonomous Vehicles Research Studio (AVRS) would mitigate latency and synchronisation challenges, hence facilitating more dependable multi-agent reinforcement learning studies. Modular configurations tailored for swarm and cooperative robotics would facilitate scalability, allowing researchers to expand single-agent platforms into investigations of collective intelligence with consistent setups. To reduce the barriers currently limiting AVRS-based RL publications, future work should provide standardised RL environments, reference reward functions, calibrated digital-twin models, common evaluation scenarios, and detailed reporting templates for vehicle, sensor, localisation, communication, and software configurations. Public release of training code, trained policies, random seeds, safety constraints, and hardware-deployment protocols would further improve reproducibility and enable meaningful comparison across laboratories.

7.3. Open Software, Safety, and Computational Infrastructure

In addition to hardware issues, it is essential to reconcile open-source APIs with software toolchains orientated towards machine learning. Although Quanser presently accommodates MATLAB/Simulink, Python, and ROS, the incorporation of standardised open-source APIs in conjunction with machine learning libraries like TensorFlow, PyTorch, and RLlib will facilitate wider use within research communities. Open interfaces would guarantee reproducibility and reduce obstacles for researchers familiar with data-driven machine learning workflows. The integration of safety-aware reinforcement learning should be prioritised. Integrating runtime assurance frameworks and hardware fail-safes will enable reinforcement learning agents to explore efficiently while mitigating hazardous activities that could compromise hardware integrity. Methods such as shielded exploration, constrained policy optimisation, and automated emergency shutdown systems will not only protect testbeds but also facilitate the investigation of safety-critical reinforcement learning, a progressively significant research domain.
The future development of Quanser platforms can be enhanced through further integration with edge AI and Industry 4.0 testbeds. Integrating lightweight AI accelerators like NVIDIA Jetson or Google Coral into Quanser hardware will facilitate low-latency decision-making, consistent with Industry 4.0 frameworks of cyber–physical systems and intelligent manufacturing. Furthermore, cloud-based reinforcement learning facilitates the delegation of computationally demanding training tasks, enabling extensive policy optimisation and transfer learning that can subsequently be effectively implemented on local Quanser systems. Ultimately, the expansion of the Autonomous Vehicles Research Studio (AVRS) to encompass multi-agent reinforcement learning would facilitate swarm coordination, distributed autonomy, and collaborative robotics, thereby offering a realistic setting to validate novel multi-robot RL algorithms in educational environments and provide preliminary evidence for subsequent application-specific industrial testing. Collectively, these enhancements would establish Quanser systems as innovative and robust platforms for the progression of RL research in educational and academic settings.

8. Conclusions

This article reviewed published research on RL-based real-time control using the Quanser Aero, Aero 2, 3-DOF Helicopter, and Autonomous Vehicles Research Studio platforms. The reviewed literature indicates that these systems provide useful environments for investigating RL under physical constraints that are commonly absent or simplified in simulation, including sensor noise, actuator limits, computational delay, nonlinear dynamics, communication constraints, and safety requirements. Aero and Aero 2 have primarily supported stabilisation, trajectory-tracking, and energy-aware control studies; the 3-DOF Helicopter has enabled evaluation under strongly coupled nonlinear dynamics; and AVRS provides an extensible environment for aerial–ground, vision-based, and multi-agent research. The synthesis also identifies limitations that prevent straightforward comparison and reproduction of published results. These include heterogeneous reward functions and evaluation measures, incomplete reporting of hardware and software configurations, limited availability of implementation code and trained policies, sample inefficiency, safety constraints, and discrepancies between simulation and physical deployment. Consequently, the available evidence should be interpreted as a collection of platform- and task-specific findings rather than proof that a particular RL algorithm is universally superior. Future research should prioritise standardised experimental protocols, complete reporting of training and hardware configurations, open and interoperable software interfaces, safety-aware RL mechanisms, higher-fidelity digital twins, and common benchmarks that permit comparisons across algorithms and platforms. The reviewed evidence should therefore not be interpreted as demonstrating industrial readiness. Quanser experiments primarily establish algorithmic feasibility under controlled laboratory conditions, and further validation on full-scale systems, under broader disturbances and operational constraints, remains necessary before engineering deployment. This review does not report original experiments; its contribution is a structured synthesis and critical assessment of experimental evidence reported in the existing literature.

Author Contributions

Conceptualisation, G.E.M.A., S.A.M. and J.T.; methodology, G.E.M.A., S.A.M. and J.T.; literature investigation, G.E.M.A., S.A.M. and J.T.; formal analysis, G.E.M.A., S.A.M. and J.T.; writing—original draft preparation, G.E.M.A., S.A.M. and J.T.; writing—review and editing, G.E.M.A., S.A.M. and J.T.; visualisation, G.E.M.A., S.A.M. and J.T.; supervision, J.T. All authors have read and agreed to the published version of the manuscript.

Funding

This research has no funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article.

Acknowledgments

We thank Aarhus University, Denmark, for providing the research facility and support to conduct this work.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Dyvik, M.; Fjereide, D.E.; Rotondo, D. Modeling and identification of the Quanser Aero using a detailed description of friction and centripetal forces. In Proceedings of the Scandinavian Simulation Society Conference, Vasteras, Sweden, 25–28 September 2023; pp. 246–253. [Google Scholar]
  2. Rotondo, D.; Sanchez, H.S. Experiences with using Kahoot! in control theoretical courses. In Proceedings of the European Control Conference (ECC), Stockholm, Sweden, 25–28 June 2024; pp. 2678–2684. [Google Scholar]
  3. Iza, J.; Paredes, E.; Herrera, M.; Benítez, D.; Pérez-Pérez, N.; Camacho, O. Real-time experimental benchmarking of control strategies for a coupled 2-DOF helicopter. Eng 2026, 7, 170. [Google Scholar] [CrossRef] [Scilit]
  4. Pereda Perez, G. Modeling and Control Using Feedback Linearization of a Quanser Aero 2 Device. Master’s Thesis, Universitat Politecnica de Catalunya, Barcelona, Spain, 2024. [Google Scholar]
  5. Kumar, S.; Dewan, L. A comparative analysis of LQR and SMC for Quanser Aero. In Control and Measurement Applications for Smart Grid: Selected Papers of SGESC; Springer: Singapore, 2022; pp. 453–463. [Google Scholar]
  6. Segerstrom, E.; Podlaski, M.; Khare, A.; Vanfretti, L. Parameter optimization and model validation of Quanser Aero using Modelica and RaPId. In Proceedings of the AIAA/IEEE Electric Aircraft Technology Symposium (EATS), Denver, CO, USA, 11–13 August 2021; pp. 1–9. [Google Scholar]
  7. Kumar, S.; Dewan, L. Set-point tracking of Quanser Aero using SMC in the presence of uncertainties. In Proceedings of the International Conference on Intelligent Computing and Control Systems (ICICCS), Madurai, India, 6–8 May 2021; pp. 1595–1601. [Google Scholar]
  8. Quanser, I. Quanser Real-Time Control (QUARC): User Documentation. Available online: https://docs.quanser.com/quanser-sdk/documentation/quarc.html (accessed on 26 August 2026).
  9. MathWorks. Control a Quanser QUBE Pendulum with a Raspberry Pi Using Reinforcement Learning. Available online: https://www.mathworks.com/help/reinforcement-learning/ug/use-reinforcement-learning-to-control-quanser-qube-pendulum-via-raspberry-pi.html (accessed on 26 August 2026).
  10. Polzounov, K.; Redden, L.; Sundar, R. Blue River Controls: A Toolkit for Reinforcement Learning Control Systems on Hardware. In Proceedings of the Workshop on Deep Reinforcement Learning, 33rd Conference on Neural Information Processing Systems, Vancouver, BC, Canada, 14 December 2019. [Google Scholar] [CrossRef] [Scilit]
  11. Kumawat, G.; Goswami, N.K.; Vajpai, J. Design of fuzzy controller for tracking of desired trajectory of 2-DOF Aero system. In Proceedings of the IEEE Power India International Conference (PIICON), Jaipur, India, 10–12 December 2024; pp. 1–6. [Google Scholar]
  12. Volpi, V. Study on Electric Propulsion Solutions for Vertical Takeoff and Landing Air Vehicles. Master’s Thesis, Politecnico di Milano, Milan, Italy, 2021. [Google Scholar]
  13. Yang, S.; Xi, L.; Hao, J.; Wang, W. Aerodynamic-parameter identification and attitude control of quad-rotor model with CIFER and adaptive LADRC. Chin. J. Mech. Eng. 2021, 34, 1. [Google Scholar] [CrossRef] [Scilit]
  14. Labdai, S.; Chrifi-Alaoui, L.; Drid, S.; Delahoche, L.; Bussy, P. Real-time implementation of an optimized fractional sliding mode controller on the Quanser-Aero helicopter. In Proceedings of the International Conference on Control, Automation and Diagnosis (ICCAD), Paris, France, 7–9 October 2020; pp. 1–6. [Google Scholar]
  15. Lopes, A.N.D.; Arcese, L.; Guelton, K.; Cherifi, A. Sampled-data controller design with application to the Quanser Aero 2-DOF helicopter. In Proceedings of the IEEE International Conference on Automation, Quality and Testing, Robotics (AQTR), Cluj-Napoca, Romania, 21–23 May 2020; pp. 1–6. [Google Scholar]
  16. AlHamouch, A.; Tuqan, M.; Bardawil, C.; Daher, N. Investigating performance of adaptive and robust control schemes for Quanser Aero. In Proceedings of the International Conference on Advanced Computational Tools for Engineering Applications (ACTEA), Zouk Mosbeh, Lebanon, 3–5 July 2019; pp. 1–6. [Google Scholar]
  17. Mehndiratta, M.; Kayacan, E. Receding horizon control of a 3 DOF helicopter using online estimation of aerodynamic parameters. Proc. Inst. Mech. Eng. Part G J. Aerosp. Eng. 2018, 232, 1442–1453. [Google Scholar] [CrossRef] [Scilit]
  18. Rojas-Cubides, H.; Cortes-Romero, J.; Coral-Enriquez, H.; Rojas-Cubides, H. Sliding mode control assisted by GPI observers for tracking tasks of a nonlinear multivariable twin-rotor aerodynamical system. Control Eng. Pract. 2019, 88, 1–15. [Google Scholar] [CrossRef] [Scilit]
  19. Arabi, E.; Yucelen, T. Experimental results with the set-theoretic model reference adaptive control architecture on an aerospace testbed. In Proceedings of the AIAA Scitech Forum, San Diego, CA, USA, 7–11 January 2019; p. 0930. [Google Scholar]
  20. Li, H.; Luo, P.; Li, Z.; Zhu, G.; Zhang, X. Finite time-adaptive full-state quantitative control of quadrotor aircraft and QDrone experimental platform verification. Drones 2024, 8, 351. [Google Scholar] [CrossRef] [Scilit]
  21. Chen, Z.; Zhong, W.; Xie, S.; Zhang, Y.; Yuen, C. Observer-based robust integral reinforcement learning for attitude regulation of quadrotors. Knowl. Based Syst. 2024, 303, 112360. [Google Scholar] [CrossRef] [Scilit]
  22. Abro, G.E.M.; Abdallah, A.M.; Elshaar, M.E. Swarm coordination and trajectory tracking in quadrotor UAVs using fractional-order PID control strategy. IEEE Trans. Autom. Sci. Eng. 2025, 23, 2995–3008. [Google Scholar] [CrossRef] [Scilit]
  23. Liu, X.; Yuan, Z.; Gao, Z.; Zhang, W. Reinforcement learning-based fault-tolerant control for quadrotor UAVs under actuator fault. IEEE Trans. Ind. Inform. 2024, 20, 13926–13935. [Google Scholar] [CrossRef] [Scilit]
  24. Borbolla-Burillo, P.; Sotelo, D.; Frye, M.; Garza-Castanon, L.E.; Juarez-Moreno, L.; Sotelo, C. Design and real-time implementation of a cascaded model predictive control architecture for unmanned aerial vehicles. Mathematics 2024, 12, 739. [Google Scholar] [CrossRef] [Scilit]
  25. Yang, P.; Xuan, Y.; Li, W. Adaptive nonsingular fast-reaching terminal sliding mode control based on observer for aerial robots. Actuators 2024, 13, 98. [Google Scholar] [CrossRef] [Scilit]
  26. Muthusamy, P.K.; Suthar, B.; Muthusamy, R.; Garratt, M.; Pota, H.; Seneviratne, L.; Zweiri, Y. Self-organizing BFBEL control system for a UAV under wind disturbance. IEEE Trans. Ind. Electron. 2023, 71, 5021–5033. [Google Scholar] [CrossRef] [Scilit]
  27. Li, S.; Chen, Z.; Zhang, Q.; Hu, T.; Zhu, B. Maneuver synchronization of networked rotating platforms using historical nominal command. Control Eng. Pract. 2024, 153, 106081. [Google Scholar] [CrossRef] [Scilit]
  28. Schafer, G.; Rehrl, J.; Huber, S.; Hirlaender, S. Comparison of model predictive control and proximal policy optimization for a 1-DOF helicopter system. In Proceedings of the IEEE International Conference on Industrial Informatics (INDIN), Beijing, China, 17–20 August 2024; pp. 1–7. [Google Scholar]
  29. Fellag, R.; Belhocine, M. 2-DOF helicopter control via state feedback and full/reduced-order observers. In Proceedings of the International Conference on Electrical Engineering and Automation Control (ICEEAC), Setif, Algeria, 12–14 May 2024; pp. 1–6. [Google Scholar]
  30. Schafer, G.; Schirl, M.; Rehrl, J.; Huber, S.; Hirlaender, S. Python-based reinforcement learning on Simulink models. In International Conference on Soft Methods in Probability and Statistics; Springer: Cham, Switzerland, 2024; pp. 449–456. [Google Scholar]
  31. Rezoug, A.; Messah, A.; Messaoud, W.A.; Baizid, K.; Iqbal, J. Adaptive-optimal MIMO nonsingular terminal sliding mode control of twin-rotor helicopter system. J. Braz. Soc. Mech. Sci. Eng. 2024, 46, 162. [Google Scholar] [CrossRef] [Scilit]
  32. Schlanbusch, S.M.; Zhou, J. Adaptive predictor-based control for a helicopter system with input delays: Design and experiments. J. Autom. Intell. 2024, 3, 50–56. [Google Scholar] [CrossRef] [Scilit]
  33. Ouerdane, F.; Mysorewala, M.F. Visual servoing of a 3 DOF hover quadcopter using 2D markers. In Proceedings of the IEEE Symposium on Industrial Electronics (ISIE), Ulsan, Republic of Korea, 18–21 June 2024; pp. 1–6. [Google Scholar]
  34. Heemels, W.P.M.H.; Johansson, K.H.; Tabuada, P. An introduction to event-triggered and self-triggered control. In Proceedings of the IEEE Conference on Decision and Control (CDC), Maui, HI, USA, 10–13 December 2012; pp. 3270–3285. [Google Scholar]
  35. Amin, R.U.; Li, A. Modelling and robust attenuation tracking control of 3-DOF four rotor hover vehicle. Aircr. Eng. Aerosp. Technol. 2017, 89, 87–98. [Google Scholar] [CrossRef] [Scilit]
  36. Wang, Y.; Jiang, B.; Lu, N.; Pan, J. Hybrid modeling based double-granularity fault detection and diagnosis for quadrotor helicopter. Nonlinear Anal. Hybrid. Syst. 2016, 21, 22–36. [Google Scholar] [CrossRef] [Scilit]
  37. Abro, G.E.M.; Abdallah, A.M.; Elshaar, M.E. Helical trajectory control of quadrotor UAVs using fractional-order PID controller. In Proceedings of the IEEE International Conference on Automation Science and Engineering (CASE), Bari, Italy, 28 August–1 September 2024; pp. 2085–2090. [Google Scholar] [CrossRef] [Scilit]
  38. Sanz, R.; Garcia, P.; Zhong, Q.-C.; Albertos, P. Predictor-based control of a class of time-delay systems and its application to quadrotors. IEEE Trans. Ind. Electron. 2016, 64, 459–469. [Google Scholar] [CrossRef] [Scilit]
  39. Haddad, A.G.; Boiko, I.; Zweiri, Y. Fuzzy ensembles of reinforcement learning policies for robotic systems with varied parameters. arXiv 2023, arXiv:2311.05655. [Google Scholar]
  40. Haddad, A.G.; Boiko, I.; Zweiri, Y. Reinforcement learning generalization for nonlinear systems through dual-scale homogeneity transformations. arXiv 2023, arXiv:2311.05013. [Google Scholar]
  41. Stauffer, L.; Manjunath, P.; Kim, D.; Korpela, C. Tactical autonomous maneuver testbed for multi-agent air-ground teams. In Proceedings of the International Conference on Electrical, Computer, Communications and Mechatronics Engineering (ICECCME), Tenerife, Spain, 19–21 July 2023; pp. 1–6. [Google Scholar]
  42. Ullah, N.; Mehmood, Y.; Aslam, J.; Ali, A.; Iqbal, J. UAVs-UGV leader follower formation using adaptive non-singular terminal super twisting sliding mode control. IEEE Access 2021, 9, 74385–74405. [Google Scholar] [CrossRef] [Scilit]
  43. Sun, H.; Li, J.; Wang, R.; Yang, K. Attitude control of the quadrotor UAV with mismatched disturbances based on fractional-order sliding mode and backstepping control subject to actuator faults. Fractal Fract. 2023, 7, 227. [Google Scholar] [CrossRef] [Scilit]
  44. Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A.A.; Veness, J.; Bellemare, M.G.; Graves, A.; Riedmiller, M.; Fidjeland, A.K.; Ostrovski, G.; et al. Human-level control through deep reinforcement learning. Nature 2015, 518, 529–533. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  45. Lillicrap, T.P.; Hunt, J.J.; Pritzel, A.; Heess, N.; Erez, T.; Tassa, Y.; Silver, D.; Wierstra, D. Continuous control with deep reinforcement learning. In Proceedings of the International Conference on Learning Representations, San Juan, Puerto Rico, 2–4 May 2016. [Google Scholar]
  46. Fujimoto, S.; van Hoof, H.; Meger, D. Addressing function approximation error in Actor–Critic methods. In Proceedings of the 35th International Conference on Machine Learning, Stockholm, Sweden, 10–15 July 2018; Volume 80, pp. 1587–1596. [Google Scholar]
  47. Haarnoja, T.; Zhou, A.; Abbeel, P.; Levine, S. Soft Actor–Critic: Off-policy maximum-entropy deep reinforcement learning with a stochastic actor. In Proceedings of the 35th International Conference on Machine Learning, Stockholm, Sweden, 10–15 July 2018; Volume 80, pp. 1861–1870. [Google Scholar]
  48. Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal policy optimization algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar] [CrossRef] [Scilit]
  49. Lowe, R.; Wu, Y.; Tamar, A.; Harb, J.; Abbeel, P.; Mordatch, I. Multi-agent Actor–Critic for mixed cooperative–competitive environments. In Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA, 4–9 December 2017; Volume 30. [Google Scholar]
  50. Kumar, A.; Zhou, A.; Tucker, G.; Levine, S. Conservative Q-learning for offline reinforcement learning. Adv. Neural Inf. Process. Syst. 2020, 33, 1179–1191. [Google Scholar]
  51. Zhang, Q.; Wang, H.; Cai, Y.; Xie, W.-F.; Sun, X.; Chen, L. Stability-constrained coordinated control strategy for vehicle chassis integrated AFS and DYC via reinforcement learning. Control Eng. Pract. 2026, 175, 107122. [Google Scholar] [CrossRef] [Scilit]
  52. Wang, H.; Zhang, W.; Xu, Y.; Li, H.; Ren, P. WaterCycleDiffusion: Visual-textual fusion empowered underwater image enhancement. Inf. Fusion 2026, 127, 103693. [Google Scholar] [CrossRef] [Scilit]
  53. Ma, G.; Wu, H.; Zhao, Z.; Zou, T.; Hong, K.-S. Adaptive neural network control for a nonlinear 2-DOF helicopter system with prescribed performance. IET Control Theory Appl. 2023, 17, 1789–1799. [Google Scholar] [CrossRef] [Scilit]
  54. Schlanbusch, S.M.; Zhou, J. Adaptive quantized control of uncertain nonlinear rigid body systems. Syst. Control Lett. 2023, 175, 105513. [Google Scholar] [CrossRef] [Scilit]
  55. Kim, S.-K.; Ahn, C.K. Performance-Boosting Attitude Control for 2-DOF Helicopter Applications via Surface Stabilization Approach. IEEE Trans. Ind. Electron. 2022, 69, 7234–7243. [Google Scholar] [CrossRef] [Scilit]
  56. Rinciog, A.; Meyer, A. Fabricatio-RL: A Reinforcement Learning Simulation Framework for Production Scheduling. In Proceedings of the 2021 Winter Simulation Conference (WSC), Phoenix, AZ, USA, 15–17 December 2021; pp. 1–12. [Google Scholar] [CrossRef] [Scilit]
  57. Zhao, Z.; He, W.; Mu, C.; Zou, T.; Hong, K.-S.; Li, H.-X. Reinforcement learning control for a 2-DOF helicopter with state constraints. IEEE Trans. Autom. Sci. Eng. 2022, 21, 157–167. [Google Scholar] [CrossRef] [Scilit]
  58. Jia, J.; Guo, K.; Yu, X.; Guo, L.; Xie, L. Reliability based LQR fault-tolerant control for a quadrotor UAV. In Advances in Guidance, Navigation and Control: Proceedings of ICGNC 2020; Springer: Singapore, 2021; pp. 4471–4481. [Google Scholar]
  59. Liao, T.; Haridevan, A.; Liu, Y.; Shan, J. Autonomous vision-based UAV landing with collision avoidance using deep learning. In Science and Information Conference; Springer: Cham, Switzerland, 2022; pp. 79–87. [Google Scholar]
  60. Matthews, M.T. Adaptive and Neural Network Based Control of Unmanned Aerial Vehicles. Ph.D. Dissertation, North Carolina A&T State University, Greensboro, NC, USA, 2021. [Google Scholar]
  61. Wahbah, M.; Chehadeh, M.; Zweiri, Y. Dynamic based estimator for UAVs with real-time identification using DNN and the modified relay feedback test. arXiv 2021, arXiv:2106.07299. [Google Scholar] [CrossRef] [Scilit]
  62. Alkayas, A.; Chehadeh, M.; Ayyad, A.; Zweiri, Y. Systematic online tuning of multirotor UAVs for accurate trajectory tracking under wind disturbances and in-flight dynamics changes. IEEE Access 2022, 10, 6798–6813. [Google Scholar] [CrossRef] [Scilit]
  63. Ayyad, A.; Chehadeh, M.; Silva, P.H.; Wahbah, M.; Hay, O.A.; Boiko, I.; Zweiri, Y. Multirotors from takeoff to real-time full identification using the modified relay feedback test and deep neural networks. IEEE Trans. Control Syst. Technol. 2021, 30, 1561–1577. [Google Scholar] [CrossRef] [Scilit]
  64. Abro, G.E.M.; Abdallah, A.M. Digital twins and control theory: A critical review on revolutionizing quadrotor UAVs. IEEE Access 2024, 12, 43291–43307. [Google Scholar] [CrossRef] [Scilit]
  65. Fandel, A.; Birge, A.; Miah, M.S. Development of reinforcement learning algorithm for 2-DOF helicopter model. In Proceedings of the IEEE Symposium on Industrial Electronics (ISIE), Cairns, QLD, Australia, 12–15 June 2018; pp. 553–558. [Google Scholar]
  66. Zhao, Z.; Weng, Y.; Liu, Z.; Liu, Y.; Hong, K.-S. Integral reinforcement learning control of an uncertain 2-DOF helicopter system with input quantization and state constraints. IEEE Trans. Ind. Electron. 2025, 72, 9250–9259. [Google Scholar] [CrossRef] [Scilit]
  67. Chen, T.; Shan, J. A novel cable-suspended quadrotor transportation system: From theory to experiment. Aerosp. Sci. Technol. 2020, 104, 105974. [Google Scholar] [CrossRef] [Scilit]
  68. Durdevic, P.; Ortiz-Arroyo, D. A deep neural network sensor for visual servoing in 3D spaces. Sensors 2020, 20, 1437. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  69. Durdevic, P.; Ortiz-Arroyo, D.; Li, S.; Yang, Z. Vision aided navigation of a quad-rotor for autonomous wind-farm inspection. IFAC-PapersOnLine 2019, 52, 61–66. [Google Scholar] [CrossRef] [Scilit]
  70. Pitarch, J.L.; Sala, A. Multicriteria fuzzy-polynomial observer design for a 3DoF nonlinear electromechanical platform. Eng. Appl. Artif. Intell. 2014, 30, 96–106. [Google Scholar] [CrossRef] [Scilit]
  71. Chen, F.; Lu, F.; Jiang, B.; Tao, G. Adaptive compensation control of the quadrotor helicopter using quantum information technology and disturbance observer. J. Frankl. Inst. 2014, 351, 442–455. [Google Scholar] [CrossRef] [Scilit]
  72. Cavalca, M.S.M.; Kienitz, K.H. Application of TFL/LTR robust control techniques to failure accommodation. In Proceedings of the 20th International Congress of Mechanical Engineering, Gramado, Brazil, 15–20 November 2009; pp. 1–8. [Google Scholar]
  73. Bermudez-Ortega, J.; Besada-Portas, E.; Lopez-Orozco, J.A.; Chacon, J.; de la Cruz, J.M. Developing web and TwinCAT PLC-based remote control laboratories for modern web-browsers or mobile devices. In Proceedings of the IEEE Conference on Control Applications (CCA), Buenos Aires, Argentina, 19–22 September 2016; pp. 810–815. [Google Scholar]
  74. Chacon, J.; Besada-Portas, E.; Garcia-Perez, L.; Lopez-Orozco, J.A. An integrated framework for the agile development and deployment of low cost remote laboratories. Multimed. Tools Appl. 2025, 84, 29207–29227. [Google Scholar] [CrossRef] [Scilit]
  75. Maraoui, S.; Bouzrara, K. ARX model decomposed on Meixner-like orthonormal bases. ISA Trans. 2019, 95, 278–294. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  76. Arabi, E.; Yucelen, T. A set-theoretic model reference adaptive control architecture with dead-zone effect. Control Eng. Pract. 2019, 89, 12–29. [Google Scholar] [CrossRef] [Scilit]
  77. Gruenwald, B.C.; Yucelen, T.; Muse, J.A. Direct uncertainty minimization in model reference adaptive control: Experimental results. In Proceedings of the AIAA Scitech Forum, San Diego, CA, USA, 7–11 January 2019; p. 2186. [Google Scholar]
  78. Lambert, P.; Reyhanoglu, M. Observer-based sliding mode control of a 2-DOF helicopter system. In Proceedings of the IECON—Annual Conference of the IEEE Industrial Electronics Society, Washington, DC, USA, 21–23 October 2018; pp. 2596–2600. [Google Scholar]
  79. Sadi, M.A.; Jamali, A.; Kamaruddin, A.M.N.A.; Jun, V.Y.S. Optimizing UAV performance in turbulent environments using cascaded model predictive control algorithm and Pixhawk hardware. J. Braz. Soc. Mech. Sci. Eng. 2025, 47, 396. [Google Scholar] [CrossRef] [Scilit]
  80. Mushitha, L.; Kumar, K.K.A.; Priyadharshini, S. Robust and optimal control of Quanser Aero. In Proceedings of the International Conference on Advancements in Electrical, Electronics, Communication, Computer and Automation (ICAECA), Coimbatore, India, 4–5 April 2025; pp. 1–6. [Google Scholar]
  81. Sadi, M.A.; Jamali, A.; Kamaruddin, A.M.N.A.; Jun, V.Y.S. Cascade model predictive control for enhancing UAV quadcopter stability and energy efficiency in wind turbulent mangrove forest environment. e-Prime-Adv. Electr. Eng. Electron. Energy 2024, 10, 100836. [Google Scholar] [CrossRef] [Scilit]
  82. Chiem, N.X. Synthesis of an orbit tracking controller for a 2DOF helicopter based on sequential manifolds with stabilization time in the presence of disturbances. Eng. Technol. Appl. Sci. Res. 2024, 14, 15083–15089. [Google Scholar] [CrossRef] [Scilit]
  83. Zhou, W.; Zhou, L.; Yuan, T.; Chen, R.; Liu, D. Robust performance optimization of UAV dynamic systems using MPC-PID hybrid control. Sci. Rep. 2026, 16, 2585. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  84. Abdelkader, K.; Kais, B. Robust H gain neuro-adaptive observer design for nonlinear uncertain systems. Trans. Inst. Meas. Control 2019, 41, 2293–2309. [Google Scholar] [CrossRef] [Scilit]
  85. Mohamed, S.I.A. Hybrid Active Force Control for Fixed Based Rotorcraft. Ph.D. Dissertation, Universiti Teknologi Malaysia, Skudai, Malaysia, 2022. [Google Scholar]
  86. Abdelmaksoud, S.I.; Mailah, M.; Hing, T.H. System enhancement on perturbations and wind gusts for twin-rotor helicopter using intelligent active force control. Int. J. Model. Identif. Control 2023, 43, 166–176. [Google Scholar] [CrossRef] [Scilit]
  87. Abdelmaksoud, S.I.; Mailah, M.; Abdallah, A.M. Enhancing disturbance rejection capability and body jerk performance of a twin-rotor helicopter model using intelligent active force control. J. Mek. 2021, 44, 1–20. [Google Scholar]
  88. Reyhanoglu, M.; Jafari, M.; Rehan, M. Simple learning-based robust trajectory tracking control of a 2-DOF helicopter system. Electronics 2022, 11, 2075. [Google Scholar] [CrossRef] [Scilit]
  89. Feng, Y.; Zhou, Y.; Ho, H.W. Reinforcement learning based robust tracking control for unmanned helicopter with state constraints and input saturation. Aerosp. Sci. Technol. 2024, 155, 109549. [Google Scholar] [CrossRef] [Scilit]
  90. Quanser. Academic Institutions Worldwide. Available online: https://www.quanser.com/community/our-customers/ (accessed on 28 August 2026).
  91. Schäfer, G.; Rehrl, J.; Huber, S.; Hirlaender, S. Safe reinforcement learning using ideas from model predictive control. arXiv 2026, arXiv:2607.07252. [Google Scholar] [CrossRef] [Scilit]
  92. Schäfer, G.; Rehrl, J.; Huber, S. Integrating physics-informed neural networks for safe reinforcement learning in a 1-DoF helicopter system. In International Conference on Database and Expert Systems Applications; Springer Nature: Cham, Switzerland, 2026; pp. 107–111. [Google Scholar] [CrossRef] [Scilit]
  93. Rajappa, S.; Chriette, A.; Chandra, R.; Khalil, W. Modelling and dynamic identification of 3 DOF Quanser helicopter. In Proceedings of the 16th International Conference on Advanced Robotics (ICAR), Montevideo, Uruguay, 25–29 November 2013. [Google Scholar] [CrossRef] [Scilit]
  94. Wu, T.; Acharya, S.; Khalil, A.; Aljanaideh, A.F.; Al Janaideh, M.; Kundur, D. Multi-head attention machine learning for fault classification in mixed autonomous and human-driven vehicle platoons. In Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), London, UK, 29 May–2 June 2023; pp. 10040–10046. [Google Scholar] [CrossRef] [Scilit]
  95. Shiyas, A.; Rao, S. Design of planar collision-free trochoidal paths for a multi-robot swarm. Eur. J. Control 2025, 81, 101143. [Google Scholar] [CrossRef] [Scilit]
  96. Fareh, R.; Baziyad, M.; Khadraoui, S.; Brahmi, B.; Bettayeb, M. Logarithmic potential field: A new leader–follower robotic control mechanism to enhance the execution speed and safety attributes. IEEE Access 2023, 11, 85451–85466. [Google Scholar] [CrossRef] [Scilit]
  97. Bal, L.; Mbakop, S.; Espindola-Winck, G.; Sueur, C.; Merzouki, R. Cooperative curve-based synchronized control of a fleet of autonomous robots. IEEE/ASME Trans. Mechatron. 2025, 30, 2900–2909. [Google Scholar] [CrossRef] [Scilit]
  98. Gąsieniec, L.; Kuszner, Ł.; Latif, E.; Parasuraman, R.; Spirakis, P.G.; Stachowiak, G. Brief Announcement: Anonymous Distributed Localisation via Spatial Population Protocols. In Proceedings of the 4th Symposium on Algorithmic Foundations of Dynamic Networks (SAND 2025), Liverpool, UK, 9–11 June 2025; Volume 330, pp. 19:1–19:5. [Google Scholar] [CrossRef]
  99. Rapalski, A.; Dudzik, S. Energy consumption analysis of navigation algorithms for wheeled mobile robots. Energies 2023, 16, 1532. [Google Scholar] [CrossRef] [Scilit]
  100. Zhang, Z.; Bian, J.; Wu, K. Relay-switching-based fixed-time tracking controller for nonholonomic state-constrained systems. IEEE/CAA J. Autom. Sin. 2022, 10, 1778–1780. [Google Scholar] [CrossRef] [Scilit]
  101. Gao, S.; Zhang, H.; Wang, Z.; Huang, C.; Yan, H. Optimal injection attack strategy for cyber-physical systems under resource constraint. IEEE Trans. Control Netw. Syst. 2022, 10, 636–646. [Google Scholar] [CrossRef] [Scilit]
  102. Wu, Y.; Zuo, Z.; Han, Q.; Wang, Y.; Yang, H. Formation control of wheeled mobile robots with multiple virtual leaders under communication failures. IEEE Trans. Control Syst. Technol. 2022, 31, 295–305. [Google Scholar] [CrossRef] [Scilit]
  103. Tassanbi, A.; Iskakov, A.; Do, T.D.; Ali, M.H. Interactive real-time leader follower control system for UAV and UGV. In Proceedings of the International Conference on Electrical, Computer, Communications and Mechatronics Engineering (ICECCME), Male, Maldives, 16–18 November 2022; pp. 1–8. [Google Scholar]
Figure 2. Quanser Aero with all labelled components.
Figure 2. Quanser Aero with all labelled components.
Electronics 15 04111 g002
Figure 3. Quanser Aero 2 with all labelled components.
Figure 3. Quanser Aero 2 with all labelled components.
Electronics 15 04111 g003
Figure 4. Quanser 3-DOF Helicopter Model with all labelled components.
Figure 4. Quanser 3-DOF Helicopter Model with all labelled components.
Electronics 15 04111 g004
Figure 5. Quanser AVRS setup with all devices.
Figure 5. Quanser AVRS setup with all devices.
Electronics 15 04111 g005
Figure 6. General workflow for the training and deployment over Quanser’s testbeds.
Figure 6. General workflow for the training and deployment over Quanser’s testbeds.
Electronics 15 04111 g006
Figure 7. Evidence maturity of RL validation across the reviewed Quanser platforms. “Reported” denotes identifiable published evidence, whereas “limited” indicates isolated or incompletely documented studies rather than a mature experimental body of literature.
Figure 7. Evidence maturity of RL validation across the reviewed Quanser platforms. “Reported” denotes identifiable published evidence, whereas “limited” indicates isolated or incompletely documented studies rather than a mature experimental body of literature.
Electronics 15 04111 g007
Table 1. Technical specifications of the Quanser AERO (https://www.quanser.com/wp-content/uploads/2017/01/Quanser_AERO_Product_Info_Sheet_v1.1.pdf, accesssed on 8 September 2026).
Table 1. Technical specifications of the Quanser AERO (https://www.quanser.com/wp-content/uploads/2017/01/Quanser_AERO_Product_Info_Sheet_v1.1.pdf, accesssed on 8 September 2026).
ParameterValue
Device mass3.6 kg
Device height (ground to top of base)45 cm
Helicopter body mass1.39 kg
Helicopter body length48 cm
Base dimensions (W × L)17.5 cm × 17.5 cm
Encoder resolution (in quadrature)512 counts/rev
Pitch angle range 75 ° (± 37.5 °)
Yaw angle range 360 °
Motor/propeller force–thrust constant0.119 N/V
Motor/propeller torque–thrust constant0.0036 Nm/V
Propeller diameter12.7 cm
Propeller pitch15.2 cm
Motor armature resistance0.83 Ω
Motor current–torque constant57.7 mN.m/A
Table 2. Technical specifications of the Quanser Aero 2 (https://www.quanser.com/products/aero-2/, accesssed on 8 September 2026).
Table 2. Technical specifications of the Quanser Aero 2 (https://www.quanser.com/products/aero-2/, accesssed on 8 September 2026).
ParameterValue
Device Dimensions (D × W × H)18 cm × 52 cm × 40 cm
Operating Space (D × W × H)52 cm × 52 cm × 62 cm
Mass4.7 kg
Pitch Angle Range 90 ° (± 45 ° from horizontal)
Yaw Angle Range 360 ° continuous
Pitch Encoder Resolution2880 counts/revolution
Yaw Encoder Resolution4096 counts/revolution
Prop Thrust Constant 5 × 10 4 N·s/rad
Inertial Thrust Constant0.042 Nm/A
Inertial Measurement Unit (IMU)IIM-42652 compact six-axis MEMS device
Tri-axis Gyroscope Range±500 dps
Tri-axis Accelerometer Range±2 g
Table 3. Technical specifications of the Quanser 3-DOF Helicopter (https://www.quanser.com/products/3-dof-helicopter/, accesssed on 8 September 2026).
Table 3. Technical specifications of the Quanser 3-DOF Helicopter (https://www.quanser.com/products/3-dof-helicopter/, accesssed on 8 September 2026).
ParameterValue
Device Mass6.2 kg
Device Height (ground to top of base)45 cm
Device Length (counterweight to front of propellers)127 cm
Base Dimensions (W × L)17.5 cm × 17.5 cm
Pitch Encoder Resolution (quadrature mode)4096 counts/rev
Travel Encoder Resolution (quadrature mode)8192 counts/rev
Pitch Angle Range± 32.0 °
Elevation Angle Range 63.5 °
Travel Angle Range 360 °
Table 4. Comparison of Quanser testbeds (https://www.quanser.com/resource-type/technical-resources/, accesssed on 8 September 2026): hardware, sensors, and software interfaces.
Table 4. Comparison of Quanser testbeds (https://www.quanser.com/resource-type/technical-resources/, accesssed on 8 September 2026): hardware, sensors, and software interfaces.
PlatformHardware FeaturesSensorsAPI/Software SupportROS/Multi-Agent Support
Quanser AeroLightweight (3.46 kg), compact base (17.5 cm × 17.5 cm), 1–2 DOF (pitch, yaw), dual-rotor testbedHigh-resolution encoder (8192 counts/rev) for pitch and yawSeamless with MATLAB/Simulink, QUARCNot natively ROS; primarily single-agent, suitable for introductory RL
Quanser Aero 2Moderate size (4.7 kg), enhanced modular design, 2 DOF (pitch, yaw), improved mechanical structureEncoders: pitch (2880 counts/rev) and yaw (4096 counts/rev); onboard six-axis MEMS IMU (IIM-42652), gyro (±500 dps), accel (±2 g)MATLAB/Simulink, QUARC, Python APIsPython-friendly; partial ROS integration possible; single-agent RL with extended workflows
Quanser 3-DOF HelicopterHeavier (6.2 kg), larger footprint (127 cm length), 3 DOF (pitch, elevation, travel), gimbaled pivot for attitude controlEncoders: pitch (4096 counts/rev), travel (8192 counts/rev), and wide angular ranges; lacks modern onboard IMUMATLAB/Simulink, QUARCNot designed for ROS; focus on single-vehicle control validation; limited RL scalability
AVRS (QDrone/QBot)Full-stack lab setup with aerial (QDrone/QDrone 2) and ground (QBot/QBot 2e) vehicles; requires larger workspace; onboard compute (Jetson Xavier NX/Intel Aero)Rich sensor suite: vision cameras, IMUs, depth sensors, localisation cameras, and Wi-Fi-based swarm networkingMATLAB/Simulink, QUARC, Python, C++, ROS integrationROS-native, multi-agent ready; supports swarm robotics, vision-based RL, and scalable experiments
Table 5. Compact taxonomy of RL-based control studies.
Table 5. Compact taxonomy of RL-based control studies.
DimensionCategoryTypical MethodsObjectiveKey Evaluation Criteria
Control taskStabilisationPPO, SAC, DDPG, TD3Regulate attitude or positionError, overshoot, settling time
TrackingPPO, DDPG, TD3, SACFollow references or pathsTracking error, effort, constraints
Energy-awareMulti-objective RL, PPO, SACBalance accuracy and energyEnergy–accuracy trade-off
Fault-tolerantRobust/constrained RLOperate under faults or disturbancesRobustness, safety, saturation
Multi-agentMADDPG, MAPPO, QMIXCoordinate vehiclesScalability, communication, safety
AlgorithmValue-basedQ-learning, DQNDiscrete action controlDiscretisation, scalability, bias
Policy gradientREINFORCE, PPODirect policy optimisationStability, interaction cost
Deterministic Actor–CriticDDPG, TD3Continuous controlExploration, critic bias, reuse
Stochastic Actor–CriticSACEntropy-based continuous controlExploration, efficiency, tuning
Offline RLCQL, IQL, BCQLearn from fixed datasetsCoverage, unseen actions
Multi-agent RLMADDPG, MAPPO, QMIXCooperative or competitive controlNon-stationarity, credit assignment
TrainingSimulationAny RL methodTrain without hardwareModel fidelity, hardware validation
Sim2RealPPO, SAC, DDPG, TD3Transfer policy to hardwareTransfer loss, model mismatch
Online hardwareAdaptive/integral RLLearn directly on the testbedSafety, wear, time, repeatability
HybridPretraining + refinementSimulate, transfer, and fine-tuneRobustness versus hardware cost
Table 6. Study-level summary of RL applications on Quanser platforms.
Table 6. Study-level summary of RL applications on Quanser platforms.
Ref.YearPlatformTaskRL MethodTraining ModeReported Outcome
[65]2018Aero, 2-DOFPitch–yaw stabilisationRL controllerSimulationFeasible coupled control; quantitative hardware result NR.
[57]20222-DOF HelicopterConstrained trackingRobust Actor–Critic RLSimulation + hardwareImproved tracking while satisfying state constraints.
[39]2023QDrone with loadRobust trajectory trackingFuzzy RL ensembleSim2RealThe 3D RMSE decreased to 0.0343 m and 0.0524 m under wind.
[40]2023QDrone with loadLoad-position controlDDPG with homogeneity transformationSim2RealReported 96% success and 0.0253 m 3D RMSE.
[28]2024Aero 2, 1-DOFPitch trackingPPO vs. MPC/LQRSimulation + hardware fine-tuningMean error: 3.80 ° simulation, 3.75 ° hardware, and 3.2 ° after fine-tuning.
[30]2024Aero 2, 1-DOFReference trackingPPOSim2RealBest simulation mean deviation was 4.6 °, demonstrated by hardware transfer.
[66]20252-DOF HelicopterQuantised constrained trackingIntegral Actor–Critic RLSimulation + hardwareTracking achieved under input quantisation and state constraints.
Note: NR indicates that a directly comparable numerical result was not reported. Differences in tasks, metrics, and experimental settings prevent direct algorithm ranking.
Table 7. Quanser-specific simulation-to-reality gaps and mitigation strategies.
Table 7. Quanser-specific simulation-to-reality gaps and mitigation strategies.
Gap SourcePlatformTransfer EffectMitigationEvidence
Motor dead zone and nonlinear thrustAero, Aero 2, 3-DOF HelicopterIneffective small actions, tracking error, and oscillationsIdentify dead zones, model nonlinear thrust, and randomise parametersAddressed in control studies; isolated RL evidence is unavailable.
Encoder quantisation and missing statesAll platforms; Aero 2 transfer studyNoisy observations, poor velocity estimates, and action switchingUse filtering, observers, measurement history, and realistic sensor modelsState redesign is shown in [28]; related methods appear in [51,52,66].
Aerodynamic couplingAero/Aero 2 (2-DOF) and 3-DOF HelicopterUnintended cross-axis motion and altered responseUse coupled nonlinear models, domain randomisation, and residual adaptationModelling is reported in [1,66]; RL ablation evidence is unavailable.
Clearance, backlash, and frictionAero and 3-DOF mechanical jointsDelayed motion, hysteresis, and limit cyclesIdentify mechanical effects, randomise parameters, and adapt onlineFriction is modelled in [1]; RL-specific evidence is limited.
Sensor and communication latencyAVRS and embedded systemsDelayed actions, instability, and poor coordinationRandomise latency and use timestamped observations or delay-aware policiesMatched AVRS simulation–hardware results remain scarce.
Different simulated and measured statesAero 2 and sensor-limited systemsPolicy requires states unavailable on hardwareAlign observation spaces and simulate realistic sensorsDirectly demonstrated in [28].
Residual transfer mismatchAero 2 and other platformsRemaining hardware tracking errorApply limited hardware fine-tuning or residual learningFine-tuning improved performance in [28], but required extensive hardware use.
Table 8. Practical starting points for simulation-based tuning before Quanser deployment.
Table 8. Practical starting points for simulation-based tuning before Quanser deployment.
MethodKey ParametersSimulation Starting PointMain Hardware Risk
PPOLearning rate, rollout, batch size, clipping, entropy, GAE α = 3 × 10 4 ; γ = 0.99 ; λ = 0.95 ; clip = 0.2 ; batch = 64–256Large updates or high entropy may cause abrupt actions.
SACLearning rates, replay buffer, batch size, target update, entropy α = 3 × 10 4 ; γ = 0.99 ; τ = 0.005 ; batch = 256 High entropy may increase unsafe exploration.
DDPGActor/critic rates, exploration noise, replay buffer, target updateActor = 10 4 ; critic = 10 3 ; τ = 0.005 ; batch = 128 –256Critic errors or noise may produce saturated commands.
TD3Learning rates, target noise, policy delay, replay buffer α = 3 × 10 4 ; delay = 2 ; τ = 0.005 ; batch = 128 –256Noise must remain within physical action limits.
DQNAction discretisation, exploration decay, replay buffer, target update α = 10 4 ; buffer = 10 5 10 6 ; gradual exploration decayCoarse actions reduce resolution; fine actions increase complexity.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Abro, G.E.M.; Memon, S.A.; Tanveer, J. Reinforcement Learning for Real-Time Control Using Quanser Platforms: A Structured Narrative Review. Electronics 2026, 15, 4111. https://doi.org/10.3390/electronics15184111

AMA Style

Abro GEM, Memon SA, Tanveer J. Reinforcement Learning for Real-Time Control Using Quanser Platforms: A Structured Narrative Review. Electronics. 2026; 15(18):4111. https://doi.org/10.3390/electronics15184111

Chicago/Turabian Style

Abro, Ghulam E Mustafa, Sufyan Ali Memon, and Jawad Tanveer. 2026. "Reinforcement Learning for Real-Time Control Using Quanser Platforms: A Structured Narrative Review" Electronics 15, no. 18: 4111. https://doi.org/10.3390/electronics15184111

APA Style

Abro, G. E. M., Memon, S. A., & Tanveer, J. (2026). Reinforcement Learning for Real-Time Control Using Quanser Platforms: A Structured Narrative Review. Electronics, 15(18), 4111. https://doi.org/10.3390/electronics15184111

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop