1. Introduction
Wheeled bipedal robots (WBRs) have emerged as compact and agile mobile platforms that combine the high-speed efficiency of wheeled locomotion with the posture-regulation capability of articulated legs [
1]. By actively modulating the center of mass (CoM), WBRs can maintain dynamic balance while operating within confined industrial environments, making them promising candidates for logistics, inspection, and surveillance applications. However, the inherently unstable inverted-pendulum dynamics and underactuated structure of WBRs render their balance and trajectory tracking highly sensitive to actuator saturation, parameter uncertainty, and persistent disturbances such as ground inclinations and payload variations [
2].
Ensuring both stable balance and accurate longitudinal tracking under sustained disturbances remains a central control challenge. While optimal state-feedback methods such as Linear Quadratic Regulator (LQR) have been widely adopted for WBR stabilization, their practical deployment often requires the explicit treatment of steady-state tracking errors and implementation constraints beyond idealized simulation environments [
3,
4,
5,
6]. Specifically, constant torque biases induced by slope or modeling mismatch may lead to residual offsets unless integral action is systematically incorporated into the control framework.
An example of the WBR platform considered in this study is shown in
Figure 1.
Although integral augmentation is theoretically well established in optimal servo control formulations, its comparative evaluation against conventional full-state tracking strategies under identical design conditions has not been sufficiently investigated for embedded WBR platforms [
1,
3,
5,
7,
8,
9,
10,
11]. Moreover, many prior studies emphasize simulation-level validation, leaving a gap in implementation-level analysis that accounts for sampling effects, actuator nonlinearities, and measurement noise.
To address these issues, this paper investigates the practical realization of an integral-augmented optimal tracking strategy for a three-degrees-of-freedom sagittal-plane WBR model. A Linear Quadratic Servo (LQ-Servo) controller is systematically compared with a conventional Full-State Feedback (FSF) tracking controller under identical weighting structures. Both controllers are discretized and executed on embedded hardware within a real-time Hardware-in-the-Loop Simulation (HILS) framework, enabling implementation-level performance evaluation.
The main contributions of this work are summarized as follows:
Development of a unified control-oriented 3-DOF sagittal-plane WBR model enabling structured comparison between FSF and LQ-Servo tracking controllers.
Practical embedded realization of an integral-augmented optimal controller targeting steady-state error elimination under constant disturbance torque.
Real-time HILS validation quantifying tracking accuracy and robustness through time-domain metrics and disk margin analysis.
This study provides practical guidance for selecting tracking architectures for cost-effective embedded WBR systems operating under sustained disturbances and parametric uncertainties.
6. Discussion
The real-time HILS results in
Section 5 enable a practical comparison of two state-feedback tracking architectures for an underactuated WBR.
Table 4,
Table 5,
Table 6 and
Table 7 provide complementary views of performance by combining tracking-error, step-response, and actuator-input metrics. Accordingly, the discussion below interprets the observed behaviors in terms of tracking accuracy, input-effort trade-offs, disturbance rejection, robustness notions, and implementation considerations.
6.1. Tracking-Error and Transient Performance Interpretation
The ISE results and the reported step-response metrics capture different aspects of closed-loop behavior. While the step-response indices of x emphasize transient characteristics around the commanded position transition, the ISE values reflect the accumulated deviation over the full experimental horizon and therefore include both transient and steady-state effects.
In the present discussion, the step-response indices are interpreted only for the cases in which the response can be consistently viewed as a single transition over the full evaluation interval, namely, the nominal, uncertainty-only, and sine-disturbance cases. For the constant-bias disturbance cases (D and DC), the disturbance is injected at s, whereas the original step-response metrics are determined mainly by the earlier reference transition around s. Therefore, those indices do not provide a meaningful basis for comparing the no-disturbance and constant-disturbance conditions and are not used here for disturbance-related interpretation.
First, the LQ-Servo consistently yields smaller in all tested cases, indicating tighter regulation of the body pitch around the equilibrium. This trend is consistent with the smoother attitude behavior observed in the time histories and suggests that the integral augmentation does not compromise the practical stabilization of .
Second, the x-tracking results reveal a more nuanced trade-off. In several cases, including the nominal and time-varying disturbance conditions, FSF achieves smaller , indicating reduced accumulated position error over the finite evaluation horizon. At the same time, the step-response indices reported for the nominal, uncertainty-only, and sine-disturbance cases show that FSF generally produces a more aggressive longitudinal response, whereas LQ-Servo tends to regulate the transition more conservatively and with a smaller residual offset near the final reference. However, the updated settling-time results indicate that convergence behavior does not uniformly favor one controller. In the nominal case, LQ-Servo settles earlier than FSF under the adopted criterion, whereas in the uncertainty-only case FSF settles earlier. At Hz, LQ-Servo satisfies the settling criterion whereas FSF does not, and at 1 Hz neither controller satisfies the criterion. Accordingly, the step-response indices should be interpreted as revealing a trade-off among response aggressiveness, residual offset, and convergence behavior, rather than a uniform transient advantage for either controller.
However, these ISE results do not imply uniformly better regulation quality in all respects. The step-response indices reported for the nominal, uncertainty-only, and sine-disturbance cases show that FSF generally produces a more aggressive longitudinal response, with larger overshoot and larger residual steady-state error, whereas LQ-Servo tends to regulate the transition more conservatively and with smaller offset near the final reference. The updated settling-time values further show that the convergence-time comparison is scenario-dependent: LQ-Servo is more favorable in the nominal case and in the Hz sine-disturbance case, FSF settles earlier in the uncertainty-only case, and neither controller exhibits finite settling under the adopted criterion at 1 Hz. Thus, the two sets of metrics should be interpreted together: finite-horizon reflects accumulated tracking performance, whereas the step-response indices reveal differences in transient shape, convergence behavior, and final-position regulation.
For the constant-bias disturbance cases (D and DC), the interpretation requires additional care. Because the disturbance is injected only after s, the full-horizon values reflect the disturbance effect only over the later part of the experiment. As a result, the similar values of LQ-Servo and FSF in these cases do not indicate equivalent steady-state disturbance rejection. Rather, they suggest that the finite evaluation horizon partially masks the structural advantage of LQ-Servo, which removes the post-disturbance steady-state offset, whereas FSF retains a persistent position bias. Accordingly, if the constant-disturbance interval were extended, the accumulated position error of FSF would be expected to increase more noticeably due to its nonzero residual offset.
The comparison between the two sine-disturbance frequencies also suggests a frequency-dependent trade-off. At Hz, LQ-Servo exhibits substantially smaller overshoot and much smaller steady-state error, and it is also the only controller of the two that satisfies the adopted settling criterion, although FSF still attains a smaller over the full horizon. At 1 Hz, the advantage of FSF becomes more pronounced, while neither controller satisfies the settling criterion and both exhibit degraded convergence behavior under the faster oscillatory disturbance. This suggests that the faster disturbance variation is less favorable to the integral compensation mechanism of LQ-Servo in terms of accumulated tracking error. Therefore, FSF is generally advantageous from the viewpoint of finite-horizon position ISE, whereas LQ-Servo is more favorable from the viewpoint of final-position regulation and offset suppression. The settling-time results further indicate that convergence characteristics become strongly scenario-dependent once uncertainty or oscillatory disturbances are introduced.
6.2. Input-Effort and Actuator-Usage Trade-Off
The input metrics provide an actuator-usage perspective that is not visible from tracking measures alone. For the body-torque channel, FSF consistently shows smaller values of , smaller RMS torque, and smaller peak absolute torque than LQ-Servo in all tested cases. This indicates that FSF places a systematically lower demand on the body-torque actuator.
For the wheel-torque channel, however, the comparison is more case-dependent. FSF produces a larger peak wheel torque in every case, and in the nominal, disturbance-only, and sine-disturbance cases it also requires larger cumulative wheel-torque effort and RMS torque. By contrast, in the uncertainty-only, disturbance-plus-uncertainty, and sine-disturbance cases, FSF achieves lower cumulative wheel-torque effort and smaller RMS torque than LQ-Servo, although its peak wheel torque still remains higher. These results indicate that no single controller is uniformly superior in actuator usage: FSF is consistently more economical on the body-torque channel and can be more efficient in cumulative wheel-torque usage under some challenging conditions, but LQ-Servo more effectively limits peak wheel-torque demand.
Taken together, the input metrics suggest a clear control-allocation difference between the two architectures. FSF tends to rely on sharper wheel-torque actions while keeping body-torque demand low, whereas LQ-Servo tends to distribute compensation more gradually, which increases body-torque usage and, in some cases, cumulative wheel-torque effort, but reduces peak wheel-torque demand. This trade-off is important from a practical actuator-design viewpoint because cumulative effort, sustained RMS demand, and instantaneous peak demand correspond to different physical concerns, such as power consumption, thermal loading, and saturation margin.
6.3. Disturbance Rejection and the Internal Model Principle
The clearest structural difference between the two controllers appears in the post-disturbance response to a constant wheel-torque bias. Because the bias is applied only after s, the comparison in the D and DC cases should be based on the time histories, accumulated-error measures, and post-disturbance steady-state behavior rather than on the initial step-response indices.
The FSF controller remains stable under the constant-bias disturbance, but it exhibits a persistent position offset in the disturbance cases. This behavior is consistent with the absence of an internal model for constant disturbances in the FSF structure, which does not include integral action on the tracking error.
In contrast, the LQ-Servo controller drives the post-disturbance position error close to zero and restores x to the reference neighborhood after the bias is applied. By augmenting the plant with integral states of the tracking errors, the LQ-Servo design embeds an internal model of step-like disturbances and reference variations, thereby enabling the steady-state rejection of constant bias torques. For applications such as transport on mild slopes or repetitive positioning tasks under persistent loads, this property is practically significant.
At the same time, the input tables show that this steady-state regulation capability is not obtained for free. The integral channel can increase cumulative actuation demand in some cases, particularly in the wheel-torque channel. Therefore, the main practical advantage of LQ-Servo is not a universally lower input usage, but rather the reliable removal of constant-bias tracking offsets.
6.4. Classical Stability Margins, Input Demand, and Task-Level Robustness
From the viewpoint of classical loop-robustness measures, the FSF controller appears more favorable because it achieves larger multiloop disk margins and phase margins than the LQ-Servo controller. However, such margins primarily characterize tolerance to gain/phase uncertainty around the nominal feedback loop and do not directly determine the rejection of constant disturbances or actuator-input distribution.
FSF is more favorable in the conventional loop-margin sense, whereas LQ-Servo is better suited to removing steady-state offsets under persistent wheel-torque bias. This distinction highlights that robustness, disturbance rejection, and actuator usage are related but non-identical performance dimensions.
6.5. Embedded Implementation and Anti-Windup Considerations
Both controllers were executed on an Arduino Due at , confirming that the proposed control laws are computationally feasible for embedded hardware. The recorded HILS inputs also remained within the actuator saturation limit of . Specifically, the body-torque peaks stayed well below the limit, and the maximum observed wheel-torque peaks also remained comfortably inside the available actuation range. The added input metrics therefore quantify not only controller effort but also the remaining margin to saturation.
For the LQ-Servo implementation, the anti-windup clamping strategy introduced in
Section 4 remains practically important. Although saturation did not occur in the reported experiments, stronger disturbances or larger reference changes could push the wheel-torque channel closer to its limit. In such cases, anti-windup logic prevents excessive integrator accumulation and reduces the risk of prolonged recovery or undesirable transients. Sensor noise and environmental variations, such as changes in friction, may affect the integral channel in practical operation by influencing error accumulation and transient recovery. Although no such instability was observed in the present HILS tests, the anti-windup clamping helps mitigate excessive integral buildup.
6.6. Modeling Contribution and Controller-Selection Implications
A modeling contribution of this study is the application of servo-oriented state-space control to a reduced-order sagittal-plane model incorporating a virtual-leg representation. Relative to rigid-leg abstractions frequently adopted in simplified WBR analyses, the present model is intended to reflect the role of leg kinematics more explicitly within the considered 3-DOF formulation.
The HILS results indicate that linear controllers designed from this model can maintain balance and track longitudinal commands under nominal, disturbed, and parameter-perturbed conditions. However, the controller choice remains application-dependent. When steady-state accuracy under persistent loads is the dominant requirement, LQ-Servo is preferred because it removes constant-bias offsets and limits peak wheel-torque demand more effectively. When lower body-actuator usage is prioritized, or when reduced cumulative wheel-torque effort under some uncertain or dynamic conditions is desirable, FSF can be attractive, provided that its larger wheel-torque peaks and possible steady-state offsets remain acceptable for the intended task.
Overall, the added actuator-input metrics clarify that the comparison between FSF and LQ-Servo is not a simple matter of which controller is “better” in an absolute sense. Rather, the choice depends on the desired balance among tracking accuracy, disturbance rejection, cumulative effort, RMS actuator loading, and peak input demand. Advanced nonlinear or predictive approaches such as MPC may further improve these trade-offs in constrained or more complex WBR operating conditions, and direct comparison with such methods remains an important topic for future work.
7. Conclusions
This paper presented a systematic design and real-time HILS-based comparative evaluation of two computationally efficient state-space tracking controllers for a WBR: a conventional FSF tracking controller and an LQ-Servo controller with integral augmentation. Unlike many prior studies that primarily rely on numerical simulation, the proposed controllers were discretized and implemented on embedded hardware, and their performance was validated through a real-time HILS framework.
A reduced 3-DOF sagittal-plane linear model, derived from a 6-DoF platform and incorporating a variable virtual-leg representation, was employed to obtain a control-oriented yet practically relevant description of balance and longitudinal motion. Both controllers were tuned under matched design conditions and executed in real time, with the controller running on an Arduino Due at and the linearized plant simulated on a Speedgoat real-time target via analog ADC/DAC interfaces. This configuration enabled implementation-level evaluation, including sampling effects and embedded computational constraints.
The real-time HILS results lead to the following main conclusions:
Nominal Operation: Both controllers achieved stable balancing and reference tracking under nominal conditions. The nominal HILS results revealed a trade-off rather than a uniformly superior controller: FSF yielded smaller finite-horizon and a shorter rise time, whereas LQ-Servo provided tighter body-pitch regulation, smaller overshoot, smaller residual position error, and shorter settling time under the adopted convergence criterion.
Robust Tracking Under Disturbance andUncertainty: Under constant wheel-torque bias, LQ-Servo removed the post-disturbance steady-state position offset, while FSF remained stable but exhibited a persistent bias because it lacked integral action. Under parametric uncertainty and time-varying sine disturbances, neither controller was uniformly superior: FSF often achieved smaller finite-horizon position , whereas LQ-Servo remained more favorable for final-position regulation and offset suppression and, in some cases, also for convergence under the adopted settling criterion. These results highlight the practical importance of integral augmentation when sustained disturbances are expected.
Implementation-Level Validation: The successful execution of both controllers on a low-cost embedded platform confirms the feasibility of real-time optimal state-space control for WBR applications. The added actuator-input analysis further showed controller-dependent implementation trade-offs: FSF consistently required less body-torque effort but produced larger peak wheel torque, whereas LQ-Servo generally reduced peak wheel-torque demand at the cost of greater cumulative actuation in some cases.
An important insight from this study is that frequency-domain robustness indicators (e.g., disk margins) did not directly translate to task-level tracking performance, convergence behavior, disturbance rejection, or actuator usage. Although FSF achieved larger multiloop disk margins and phase margins, its lack of integral action resulted in degraded steady-state tracking under constant bias loads. Therefore, for WBR applications in which persistent disturbances and steady-state accuracy are dominant concerns, LQ-Servo provides a more suitable baseline architecture. In contrast, FSF remains attractive when smaller finite-horizon position ISE, lower body-actuator usage, and structurally simpler feedback are prioritized, provided that larger wheel-torque peaks and residual offsets are acceptable.
Overall, this work contributes practical evidence that integral-augmented optimal state-space control, when carefully implemented in discrete time, offers clear benefits for steady-state disturbance rejection and offset suppression in embedded WBR platforms, but not uniform superiority across all tracking, convergence, and input-effort metrics. The present HILS comparison therefore suggests that controller selection should be made according to the desired balance among pitch regulation, finite-horizon position tracking, convergence behavior, disturbance rejection, cumulative actuator effort, RMS loading, and peak input demand.
Future work will extend the present framework to physical WBR hardware and broader operating conditions, including nonlinear and 3D dynamics, terrain variability, friction and slip effects, actuator saturation with anti-windup strategies, and gain-scheduled or robust optimal designs to enhance performance beyond the local linear regime while preserving embedded implementability. Direct comparison with advanced nonlinear or predictive controllers, such as MPC, also remains an important direction for clarifying the performance trade-offs observed in the present study.