Next Article in Journal
Autonomous Vehicles in the Traffic Ecosystem: A Comprehensive Review of Integration, Impacts, and Policy Implications
Previous Article in Journal
Towards AI-Assisted Motorcycle Safety: Multi-Modal Video Analysis for Hazard Detection and Contextual Risk Assessment
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Collaborative Control of Rear-Wheel Independent Drive Electric Vehicles During Tire Blowouts Using Broad-Extreme Reinforcement Learning: Simulation and Scaled Prototype Verification

1
Department of Electromechanical Engineering, University of Macau, Taipa, Macau, China
2
College of Science and Engineering, James Cook University, Douglas, QLD 4814, Australia
3
School of Mechanical Engineering and Automation, Fuzhou University, Fuzhou 350025, China
*
Author to whom correspondence should be addressed.
Vehicles 2026, 8(2), 40; https://doi.org/10.3390/vehicles8020040
Submission received: 22 December 2025 / Revised: 9 February 2026 / Accepted: 10 February 2026 / Published: 18 February 2026
(This article belongs to the Topic Vehicle Dynamics and Control, 2nd Edition)

Abstract

Tire blowouts represent one of the most hazardous fault scenarios for electric vehicles (EVs). While collaborative active steering control (ASC) and direct yaw moment control (DYC) can theoretically maintain stability during these events, the strong coupling effects between them make controller design challenging. To address this, an adaptive control algorithm based on broad-extreme reinforcement learning (RL), named broad critic extreme actor (BCEA), is proposed. Compared to traditional controllers, the proposed BCEA architecture is simpler to design and demonstrates enhanced robustness. Crucially, it achieves significantly faster training speed than traditional RL methods such as deep deterministic policy gradient (DDPG). Both simulation and scaled prototype tests verify the ability of the BCEA-based controller to maintain vehicle stability during different types of tire blowout scenarios. Furthermore, compared to traditional RL methods, the training efficiency is improved by more than 80%. These results indicate that the proposed BCEA controller is a promising advancement for vehicle stability control under critical failure conditions.

1. Introduction

Electric vehicles (EVs) represent a promising solution to current emission challenges [1]. However, like conventional internal combustion engine vehicles (ICEVs), EVs remain susceptible to tire blowouts. These blowouts are critical safety hazards that can cause severe path deviation, increased energy consumption and serious accidents [2]. The mechanical limitations of ICEVs prevent the real-time adjustment of driving forces required to maintain stability during such failures. In contrast, EVs can be controlled precisely, with the ability to rapidly adjust output torque to stabilize the vehicle immediately with tire blowouts. Furthermore, independent drive EVs (IDEVs) offer superior force distribution compared to conventional configurations. Among these, rear-wheel independent drive EVs (RWID EVs) feature a simpler mechanical design, optimized driving dynamics, reduced unsprung mass and enhanced directional stability compared to four-wheel independent drive EVs [3,4]. Moreover, while four-wheel systems provide comprehensive stability control, the RWID configuration offers an advantageous balance between performance and mechanical simplicity. Therefore, developing effective controllers for RWID EVs under tire blowout conditions represent a highly promising research direction. These controllers could significantly enhance the safety and reliability of EV operations, further strengthening their viability as a sustainable transportation solution.
Currently, active steering control (ASC) and direct yaw moment control (DYC) are two extensively employed methods for addressing EV stability issues induced by tire blowouts. ASC is widely recognized as an effective approach for maintaining stability, because it functions by adjusting front wheel steering angles to counteract the lateral force generated by a tire blowout. Several classic control strategies are applied to ASC. Proportional–integral–derivative (PID) control [5] and linear quadratic regulator (LQR) [6,7,8] are frequently used, with PID enabling rapid steering response and LQR ensuring high control precision. Additionally, strategies such as fuzzy control [9,10], adaptive control [11,12], and sliding mode control (SMC) [13,14,15] are also commonly applied to ASC. These methods excel at adapting to the nonlinear dynamic characteristics of EVs following a tire blowout, allowing the vehicle to precisely track its ideal trajectory. Apart from ASC, DYC is another method for stabilizing EVs during tire blowouts. The core operation of DYC involves two sequential parts. First, it calculates the desired yaw moment required for stability. Second it distributes torque to each driving wheel based on this target moment. DYC strategies are generally divided into two categories. The first includes traditional algorithms, including PID control [16], LQR [17,18,19], model predictive control (MPC) [20,21], and SMC [22,23]. The second category consists of intelligent algorithms, such as fuzzy control [24,25] and neural network control [26,27,28], which are particularly effective in handling the complex uncertainties of tire blowout scenarios. In summary, both ASC and DYC represent primary and reliable strategies for EV stability during tire blowouts.
However, practical applications of vehicle stability still face numerous challenges during emergency situations. First, conventional algorithms rely heavily on vehicle models. Tire blowout scenarios exhibit complex nonlinear dynamics that can cause a sudden decline in the performance of model-dependent controllers. Developing accurate models for these scenarios remain difficult, as simulations often oversimplify critical factors like dynamic tire behavior, resulting in significant deviations from real-world conditions. Second, traditional control algorithms struggle with parameter tuning. Parameters that are effective under specific conditions often prove inadequate when vehicle speeds or road conditions change, potentially degenerating vehicle stability. Third, standalone control systems, using either ASC or DYC in isolation, have fundamental structural limitations. Neither system alone can effectively control the vehicle during a tire blowout. ASC is mechanically constrained and often cannot provide adequate steering correction force. Similarly, the effectiveness of DYC is reduced when tire blowouts lead to sudden drops in cornering stiffness, disrupting the force and moment equilibrium of the vehicle. These combined challenges of model dependency, parameter sensitivity, and structural limitations imply the necessity of an integrated approach. Consequently, developing a control system that integrates both ASC and DYC represents a critical step towards enhancing vehicle stability during tire blowout emergencies.
Currently, collaborative ASC and DYC for electric vehicles are widely used. These methods are often developed through several approaches, including cost-based coordination [29], multi-level cooperation [30], optimized sliding mode control [31], and model predictive control [32]. Collaborative ASC and DYC are also applied to fault-tolerant control in vehicles [33,34,35,36]. These studies show that collaborative ASC and DYC are effective for both stability control and fault-tolerant control. However, existing methods still face significant challenges. They often rely on vehicle models that are difficult to establish accurately, especially when the tire blowout is considered. Additionally, these methods require complex real-time weight adjustments to adapt to nonlinear systems. Moreover, these methods suffer from high computational complexity, and this complexity hinders rapid response. Therefore, existing collaborative ASC and DYC cannot be directly used for vehicle stability control under tire blowing out conditions.
To overcome the above problems of the collaborative ASC and DYC, new methods should be proposed and introduced. The application of rapidly involving artificial intelligence (AI) technologies to vehicle control has emerged as a highly promising solution. Among various AI techniques, reinforcement learning (RL) offers unique advantages for handling complex scenarios including tire blowouts. Unlike traditional controllers, RL does not rely on complex physical vehicle models. Instead, it enables agents to construct adaptive controllers directly through interactions, thereby eliminating the need for manual parameter tuning. Building on these advantages, this research demonstrates how RL can effectively integrate ASC and DYC. The optimal weight distribution between ASC and DYC is automatically determined by an RL agent without manual intervention. Moreover, the RL-based controller adapts dynamically to varied driving conditions and specific tire blowout locations. Despite a blowout occurring at the front or rear axle, the RL-based controllers are capable of instantly recalibrating its control strategy to maintain vehicle stability. Consequently, RL represents a robust approach for achieving precise and efficient vehicle control during tire blowout scenarios. This significantly enhances vehicle safety by providing context-sensitive responses that potentially reduce accident rates.
Traditional RL algorithms often exhibit significant limitations when used in tire blowout control problems. For instance, Q-learning [37,38] requires discretization of state spaces, which proves impractical for the complex dynamics inherent in tire blowout scenarios. Deep Q network (DQN) [39,40] improves state representation using neural networks, yet it still suffers from the limitation of discrete action outputs, which restrict its applicability for precise vehicle control. Furthermore, deep deterministic policy gradient (DDPG) [41,42] addresses the discretization problem but requires the training of multiple neural networks, resulting in substantially reduced training efficiency. Many RL-based studies have been conducted for collaborative ASC and DYC. Twin delayed deep deterministic policy gradient (TD3) [43,44] and proximal policy optimization (PPO) are RL-based methods which are often used as controllers [45,46]. These methods are also applied to ASC and DYC [47,48,49,50]. However, these methods have fundamental limitations. PPO relies on a complex value model design, which increases training complexity. Additionally, TD3 requires high computational resources, which leads to slow training speeds. TD3 is also highly sensitive to hyperparameters that require extensive fine-tuning. Therefore, using reinforcement learning alone for collaborative ASC and DYC is insufficient to overcome the limitations of current methods for collaborative ASC and DYC.
To overcome the above limitations, a broad critic deep actor (BCDA) algorithm is developed [51], which enhances the DDPG architecture by replacing the critic network component with a broad learning system (BLS) [52]. As demonstrated in [51], this innovation significantly improves training efficiency while maintaining precision control. Moreover, our previous work successfully implemented this algorithm in several truck control scenarios [53]. These implementations highlight the potential of BCDA as a viable solution for addressing the limitations of traditional RL algorithms in vehicle stability control during tire blowouts.
Despite its high potential, the BCDA algorithm requires further optimization to ensure vehicle safety during critical failure scenarios. While the BLS-based critic in the BCDA algorithm significantly accelerates training efficiency, our preliminary analysis indicates that it induces instability during catastrophic events, such as tire blowouts. Specifically, inherent limitations in BLS node configuration [54] cause the critic to generate extreme or abnormal value estimates when encountering these outlier states. These anomalies propagate to the deep actor, which is based on a deep neural network, resulting in ill-conditioned policy gradients. This scenario often manifests as exploding or vanishing updates in the deep actor that degenerate vehicle control performance. Thus, re-evaluating BLS node selection parameters is essential to prevent these cascading failures.
To overcome the limitations of gradient descent, specifically the risk of exploding gradients and slow convergence in complex environments, a meta-hierarchy algorithm is proposed. Instead of relying on gradient-based backpropagation, this framework adopts cuckoo search (CS) for network optimization. Compared with other derivative-free meta-heuristic methods like genetic algorithms [55], CS is selected for its superior flexibility and lower computational cost in high-dimensional spaces [56,57]. To further address the computational efficiency, the deep actor network in the BCDA algorithm is replaced by an extreme learning machine (ELM) because ELM allows the system to increase the training speed [58]. The proposed training protocol operates in two stages. First, the BLS critic calculates Q-values and updates weights via pseudoinverse. Second, for a given state, the CS algorithm identifies the optimal action that maximizes the Q-value. The ELM-based actor is then analytically updated to map the state to this optimal action. This approach can eliminate gradient dependencies entirely and achieve rapid convergence while maximizing Q-value stability.
This study aims to enhance stability control of EVs during tire blowout scenario through a novel broad critic extreme actor (BCEA)-based RL approach. Existing collaborative ASC and DYC mainly rely on conventional control algorithms. However, active steering and yaw moment control show significant coupling effect in vehicle dynamics. Conventional strategies require manual decoupling and sensitivity-based control weight allocation. This process is computationally complex, likely to have design errors, and further limited by the highly nonlinear and strongly perturbed characteristics of the vehicle model under extreme conditions such as tire blowouts. Consequently, the adaptability and the robustness of traditional methods are limited. In contrast, the proposed BCEA method has good adaptability. Using neural network-based training mechanisms, control decoupling and weight allocation are automatically achieved without manual intervention. This approach greatly improves the environmental adaptability and control robustness of the system under complex and emergency conditions (such as tire blowouts). The overall performance is believed to be superior to that of traditional control strategies.
Compared with other RL methods, the proposed BCEA algorithm replaces gradient descent-based optimization strategies with CS method. This change effectively improves global optimization capability. At the same time, the ELM is also utilized. As a result, computational efficiency, convergence speed and stability of the algorithm can be improved.
By using physical experiments with a custom-developed prototype and computer simulations using a data-driven numerical framework, the effectiveness of the proposed controller is verified. The core innovations and contributions of this study are as follows:
(1)
This study proposes a new RL-based collaborative ASC and DYC controller based on BCEA that addresses the challenges related to structural complexity and low computational efficiency in traditional RL and BCDA approaches for vehicle control. Compared to existing collaborative ASC and DYC controllers, the proposed BCEA reduces system complexity and computational cost while enhancing the stability of algorithm convergence.
(2)
This study proposes a method that directly integrates tire blowout detection signals into the input layer, effectively addressing stability control challenges for electric vehicles during tire blowout events. Compared to traditional control strategies, the proposed approach enhances the system robustness and trajectory tracking accuracy in such emergency scenarios.
(3)
This study proposes road tests utilizing prototype vehicles, mitigating safety risks associated with real vehicle tests under tire blowout conditions. Compared to traditional testing methods, the proposed approach enhances scenario diversity and test repeatability while eliminating potential safety hazards.
The remainder of this paper is structured as follows. Section 2 reviews vehicle tire fault diagnosis methods. Section 3 details the architecture, working principle and training protocol of the BCEA controller. Section 4 describes the experimental setups and discusses results. Finally, Section 5 summarizes the research findings and outlines future directions.

2. Fault Diagnosis Methods for Tire Blowout

Previous studies have established that tire blowouts are primarily characterized by low internal air pressure in tires [59]. In this study, tire blowout means rapid loss of tire pressure. Tire explosion conditions are not considered in this study. Consequently, a threshold value for tire internal pressure is employed to identify blowout events. The fault detector in this study uses only the absolute value of tire pressure instead of the decay rate of the tire pressure. This choice is based on two reasons. First, a decay in tire pressure is a critical factor for the dynamic performance of the vehicle. In contrast, the pressure decay rate is usually hard to measure directly. It requires algorithms and filters to get accurate value. Therefore, it is not suitable to be treated as a primary indicator. Second, the proposed fault detector activates once the air pressure in a tire falls below the threshold pressure P s . It then helps the controller control the vehicle under post-blowout conditions effectively by collaborative ASC and DYC. Such control is always working no matter whether the tire pressure decays slowly or quickly. As a result, the rate of pressure loss is not used in this study.
Let P i denote the real-time pressure for each tire, where i L F , R F , L R , R R . A tire is classified as blown-out when its internal air pressure falls below a threshold pressure P s . Accordingly, the blowout coefficient λ i is defined as
λ i = { 0 ,   i f   P i > P s 1 , i f   P i P s .
When λ i = 0 , the i t h tire is considered healthy. Otherwise, when λ i = 1 , it is considered blown-out.

3. Broad Critic Extreme Actor (BCEA) Controller

3.1. Structure of Controller

The proposed controller employs a collaborative architecture integrating ASC and DYC, as shown in Figure 1.
As shown in Figure 1, the controller comprises two distinct layers. Their functions and operational sequences are detailed below:
(a)
Upper-layer controller
The upper-layer controller integrates a 2-degree-of-freedom (2-DOF) reference model to calculate the desired extra front wheel steering angle (denoted as Δ d F ) and yaw moment (denoted as M z ). The operation is as follows:
Step 1: Driver inputs (i.e., longitudinal speed and steering angle) are processed by the 2-DOF reference model to determine the desired yaw rate (denoted as r z r e f ) and side-slip angle (denoted as β r e f ).
Step 2: The actual yaw rate (denoted as r z ) and side-slip angle (denoted as β ) are continuously monitored. The system computes the error values, along with their integrals and derivatives.
Step 3: Tire blowout coefficients ( λ L F , λ R F , λ L R , λ R R ) are incorporated into the neural network-based (NN-based) module. By processing error terms and blowout coefficients, the controller generates two key outputs: (1) the reference yaw moment M z ; and (2) the desired additional front wheel steering angle Δ d F . Finally, these are transmitted to the lower-layer controller.
(b)
Lower-layer controller
The lower-layer controller converts the upper-layer signals M z and Δ d F into motor-specific control signals. The operation sequence is as follows:
Step 1: M z and Δ d F values received from the upper layer are transformed into electrical signals (voltage and current).
Step 2: The converted signals are delivered to the rear wheel motor, enabling precise vehicle actuation and real-time stability control.
This dual-layer system ensures robust vehicle stability management during tire blowout scenarios. The dimension of the NN-based controller in the upper-layer controller is determined through training. The resulting architecture, which represents a significant evolution of our previous BCDA model, is defined as the broad critic extreme actor (BCEA).

3.2. Two-Degree-of-Freedom Reference Model

The primary objective of the control problem addressed in this study is to maintain stability during tire blowouts. The 2-DOF vehicle model is widely adopted in analysis of dynamics and provides a reference for controller design in this study. Consequently, the controller aims to meet the actual state of the vehicle with the reference state defined by the 2-DOF model, thereby ensuring optimal stability under tire blowout conditions.
The 2-DOF reference model is derived from the Newton’s second law, as shown in Figure 2.
In Figure 2, m v represents the mass of the vehicle; I z z denotes the moment of inertia of the vehicle; L f and L r represent the distance from the center of gravity (C.G.) to the front and rear axles, respectively; and d F indicates the front wheel steering angle. The lateral tire forces for the front and rear wheels are denoted by F w y F and F w y R , respectively. The reference longitudinal speed, lateral speed and yaw rate are represented by u x r e f , v y r e f and r z r e f , respectively.
The dynamic equation based on Newton’s second law is defined as
{ m v v ˙ y r e f + u x r e f r z r e f = F y I z z r ˙ z r e f = M z ,
where F y and M z represent the integral lateral force and yaw moment, respectively. These are calculated by
{ F y = F y F + F y R M z = L f F y F L r F y R ,
where F y F and F y R denote the lateral tire forces on the front and rear axles, respectively. These forces are calculated by
{ F y F = F w y F cos d F F y R = F w y R .
Assuming small slip angles, the lateral tire forces are proportional to the tire slip angles. This linear relationship is expressed by
{ F w y F = C f α F F w y R = C r α R ,
where C f and C r represent the cornering stiffness of the front and rear tires, respectively. The side-slip angles for front wheel α F and rear wheel α R are defined as
{ α F = tan 1 v y r e f + L f r z r e f u x r e f d F α R = tan 1 v y r e f L r r z r e f u x r e f .
The reference side-slip angle of the vehicle β r e f is calculated by
β r e f = tan 1 U y U x θ r e f ,
where θ r e f is the reference yaw angle of the vehicle and it is calculated by
θ r e f = 0 t r z r e f d t ;
U x and U y represent the vehicle velocities in the global X and Y directions, respectively, and these velocities are calculated by
{ U x = u x r e f cos θ r e f v y r e f sin θ r e f U y = u x r e f sin θ r e f + v y r e f cos θ r e f

3.3. Design of BCEA Controller

The fundamental process of RL is illustrated in Figure 3.
As shown in Figure 3, an RL system comprises four essential components. The first component is the agent, which is the central entity in the RL system. Next, the environment represents the external scenario in which the agent operates. The third component is the policy, which defines the decision-making rule. Finally, the reward signal serves as a feedback mechanism from the environment that evaluates the quality of the agent’s actions. This learning process is executed through a cyclic agent–environment interaction loop, as detailed below.
Step 1: Upon operating within the environment, the agent detects its current state, which includes speed and positional data.
Step 2: The detected state is processed by the policy to determine the required action using predefined computational methods.
Step 3: After the action is executed, the environment provides a reward signal based on the state transition and resulting outcome.
Step 4: With the objective of maximizing cumulative rewards, the agent updates its policy parameters.
This iterative framework enables adaptive learning and optimal decision-making in dynamic environments. The flowchart of the proposed BCEA algorithm is illustrated in Figure 4 that integrates learning mechanisms to ensure robust performance across dynamic conditions.
The BCEA architecture comprises four distinct networks. The first is the extreme actor network (EAN), which receives the state s and outputs the action a . The EAN is structured as a single-hidden-layer neural network, with optimization performed using the extreme learning machine (ELM) algorithm. The second is the broad critic network (denoted as “BCN”), which receives both the state s and action a to output a Q-value, which evaluates the quality of the action. The BCN is designed based on the BLS, with its parameters optimized by the cuckoo search in this study. The third is the target EAN (denoted as “t-EAN”), which shares the same structure as the EAN. It calculates the next action a based on the next state s . The fourth network is the target BCN (denoted as “t-BCN”), which is structurally similar to the BCN. The t-BCN calculates the Q-value of the predicted action a (denoted as Q ) by processing the next state s and the predicted action a .
The training process proceeds sequentially. First, the reward (denoted as R s ) is calculated based on the current state s . The target Q-value (denoted as Q t a r g e t ) is then computed using the Bellman equation by combining this reward with Q , which is generated based on the state on the next time step. The set consisting of state s , action a and reward is stored in the experience replay buffer D . After a fixed number of iteration, random samples are retrieved from replay buffer D to train BCN and EAN, optimizing their parameters to minimize error. The design of these networks is further discussed in the following sections.

3.3.1. Design of EAN

The architecture of the EAN is shown in Figure 5.
The EAN architecture comprises the input layer, the hidden layer, and the output layer. The implementation of a single hidden layer significantly improves computational efficiency. The input layer processes the state s while the output layer generates the action a based on the features extracted from the hidden layer. These outputs depend on the network parameters including the weight and bias matrices, denoted respectively as w 1 and b 1 , which connect the input and hidden layers as illustrated in Figure 5. The state s is composed of n input variables as formulated by
s = s 1 , s 2 , , s n T ,
where s 1 , s 2 , , s n represent the variables in s .
In this study, the state s is specifically defined to include tracking errors and tire fault signals, as expressed by
s = Δ r z ,   Δ ˙ r z , Δ r z d t ,   Δ β , Δ ˙ β , Δ β d t , λ L F , λ R F , λ L R , λ R R T ,
where Δ r z and Δ β represent the tracking errors of the yaw rate and side-slip angle, respectively; λ L F , λ R F , λ L R and λ R R denote the blowout coefficients for the left-front, right-front, left-rear and right-rear tires, respectively. The tracking errors Δ r z and Δ β are defined as
{ Δ r z = r z r e f r z Δ β = β r e f β .
The inactivated hidden layer output matrix H 1 o is calculated by
H 1 o = w 1 s + b 1 .
The activated neurons are then obtained by applying a nonlinear activation function f . , as expressed by
H 1 = f H 1 o .
In this study, a Leaky ReLU function is selected as the activated function, as expressed by
f x = { x , i f   x 0 λ k x , i f   x < 0 .
The reason for utilizing the Leaky ReLU is proposed in by (Wong et al., 2025) [53].
Finally, the EAN produces action a which is computed by using the following weight matrix w s :
a = w s H 1 .
The action vector a consists of the desired additional front wheel steering angle ( Δ d F ) and the desired yaw moment M z :
a = Δ d F ,   M z

3.3.2. Design of BCN

The BCN is proposed based on an optimized version of the original BCDA algorithm by [51]. To improve the stability and efficiency, a novel initialization strategy is proposed in which a substantial number of enhancement nodes are pre-configured, ensuring their count significantly exceeds that of the input layer nodes. Similarly, the number of featured nodes is maintained which is significantly larger than the input dimension. This strategy enables precise fitting during initialization, avoiding iterative calculations and enhancing the stability of the algorithm. The improved BCN architecture is illustrated in Figure 6.
The input for the BCN combines all the action and state vectors:
Z n B C N = s 1 , s 2 , , s n s B C N , a 1 , a 2 , , a n o E A N
Let w n B C N and b n B C N denote the weight and bias matrices that map inputs to feature nodes. Thus, the inactivated feature nodes are calculated by
H o f B C N = w n B C N Z n B C N + b n B C N
Applying the activation function from Equation (15), the featured nodes become
H f B C N = f H o f B C N
Subsequently, the enhancement nodes are generated by the using the weights w e n B C N and b e n B C N :
H o e n B C N = w e n B C N H f B C N + b e n B C N
Applying Equation (14), the enhancement nodes become
H e n B C N = f H o e n B C N
Combining the featured and enhancement nodes yields the comprehensive hidden layer matrix Z h :
Z h = H f B C N H e n B C N
The final output Q-value is computed by
Q = w q Z h ,
where w q is the output weight matrix.

3.3.3. Design of Target Networks (t-EAN and t-BCN)

The target network architecture comprises the t-EAN and t-BCN. The t-EAN retains an identical structure to the EAN. When the vehicle executes action a , the t-EAN computes the predicted next action a based on the subsequent state s . Similarly, the t-BCN mirrors the structure of the BCN. It processes the anticipated state s and predicted action a to generate the estimated target Q-value Q . The coordination of these target networks enables stable prediction of future system behavior, which is essential to effective vehicle control.

3.4. Training of BCEA

3.4.1. Computation of Target Q-Value

An appropriately designed reward function is critical for evaluating system performance in safety-critical RL applications. Since trajectory stability is very important in this study, the reward focuses on minimizing tracking errors rather than maximizing the computational speed. The reward function R s is defined as
R s = 1 ε + X X r e f 2 + Y Y r e f 2
where ( X ,   Y ) and ( X r e f , Y r e f ) represent the trajectory coordinates of the actual vehicle and the 2-DOF reference model, respectively. ε is a small positive constant to prevent division by zero.
The target Q-value is computed using the Bellman equation:
Q t a r g e t = R s + γ Q ,
where γ 0 , 1 is the discount factor, and Q is the estimated Q-value for the next state.

3.4.2. Training of BCN

Let d be the batch size of the training data. The input matrix to the BCN, Z n B C N d , is formed by concatenating d training samples, as expressed by
Z n B C N d = Z n B C N 1 Z n B C N 2 Z n B C N d   .
For the i -th sample ( 1 i d ), the input is Z n B C N i = s i , a i T . The inactivated feature nodes are calculated by
H o f B C N i = w n B C N Z n B C N i + b n B C N
The activated feature nodes are obtained by
H f B C N i = f H o f B C N i ,
where f . is the Leaky ReLU function as expressed in Equation (17).
Subsequently, the enhancement nodes are generated by
H o e n B C N i = w e n B C N H f B C N i + b e n B C N ,
where w e n B C N and b e n B C N are the weights and biases for the enhancement nodes in the BCN, respectively.
H e n B C N i = f H o e n B C N i .
The complete hidden layer vector for the n -th sample combining both feature and enhancement nodes and is shown below:
Z h i = H f B C N i H e n B C N i .
Aggregating all samples in the batch yields the matrix Z m B C N d as expressed by
Z m B C N d = H f B C N 1 H f B C N 2 H f B C N d H e n B C N 1 H e n B C N 2 H e n B C N d .
The target output vector Y consists of the target Q-values for each sample, as expressed by
Y = Q t a r g e t 1 Q t a r g e t 2 Q t a r g e t d ,
where Q t a r g e t 1 , Q t a r g e t 2 , …, Q t a r g e t d represent the target Q-values of the training sets.
The output weight matrix w q is then calculated analytically using the following pseudoinverse:
w q = Y Z m B C N d .
where the pseudoinverse Z m B C N d is defined using ridge regression regularization as shown in
Z m B C N d = lim λ 0 Z m B C N d T Z m B C N d + λ I 1 Z m B C N d T .

3.4.3. Training of EAN

(a)
Training of target actions
Once the BCN structure is fixed, the Q-value depends solely on the action a for a given state s . To maximize the Q-value, a meta-heuristic approach is employed rather than gradient ascent. The cuckoo search (CS) algorithm is selected because of its simplicity and effectiveness [57]. The target action a is optimized to maximize the Q-value while keeping s constant as illustrated in Figure 7.
(b)
Training of EAN parameters
The parameters of the EAN are updated using the ELM method. For a training batch of size d , with the states set as s 1 , s 2 , …, s d and optimized target actions as a 1 t , a 2 t , …, a d t , the aggregate matrices are
s t = s 1 s 2 s d ,   a t = a 1 t a 2 t a d t
The inactivated hidden layer neurons for the i -th sample in s t are
H 1 o i = w 1 s i + b 1 ,
where w 1 and b 1 are the weights and biases in the EAN, respectively.
Applying the activation function yields
H 1 i = f H 1 o i
The hidden layer matrix H t r combines neurons from all training samples, as expressed by
H t r = H 1 1 H 1 2 H 1 d
The output weight matrix w s is then computed by the pseudoinverse:
w s = a t H t r .

3.4.4. Updating Target Networks

The parameters of the t-BCN ( θ t B C N ) and t-EAN ( θ t E A N ) are updated by using a soft update rule after each episode:
{ θ t B C N = λ b θ t B C N + 1 λ b θ B C N θ t E A N = λ e θ t E A N + 1 λ e θ E A N ,
where λ b , λ e 0 , 1 are the soft update coefficients; θ t B C N and θ t E A N represent the parameters of the original t-BCN and t-EAN, respectively; and θ B C N and θ E A N denote the parameters of the updated BCN and EAN.
The stability of the proposed BCEA algorithm is analyzed via a Lyapunov function during and after training. Hardware restrictions (e.g., battery voltage, maximum front wheel steering angle) are considered for training. The BCEA training algorithm is summarized in Algorithm 1.
Algorithm 1: Training of BCEA controller
Randomly initialize θ B C N and θ E A N ;
Initialize target parameters: θ t B C N = θ B C N and θ t E A N = θ E A N ;
Initialize replay buffer D ;
for episode = 1 to M do:
Update θ t E A N and θ t B C N using Equation (42).
for  t = 1 to N do:
  Observe current state s t ;
  Calculate the reward R s using Equation (25);
  Select action a t using EAN;
  Execute action a t and observe next state s ;
  Compute predicted action a using t-EAN;
  Compute predicted Q-value Q s ,   a using t-BCN;
  Calculate Q t a r g e t using Equation (26);
  Store tuple ( s , a , Q t a r g e t ) in replay buffer D ;
  if buffer D has sufficient samples (e.g., when samples = 1000) then:
  Randomly sample batch s 1 , s 2 …, s d , a 1 , a 2 , …, a d and Q t a r g e t 1 , Q t a r g e t 2 , …, Q t a r g e t d from D;
  Update θ B C N using Equations (27)–(36);
  Optimize target actions using CS (as shown in Figure 7).
  Update θ E A N   using Equations (37)–(41);
  end if
 end for
end for

3.5. Design of Lower- Layer Controller

The lower-layer controller translates the stability targets into physical actuator commands [51]. First, the relationship between the front wheel steering angle and the rotor displacement of steering motor is defined as
θ s m = g s m Δ d F , d F
where θ s m is the angular displacement of the steering motor rotor, and g s m . is a kinematic function determined by the specific mechanical design of the steering system. d F represents the current front wheel steering angle, while Δ d F is the compensated steering angle calculated by the BCEA controller.
The desired yaw moment is generated via differential driving of the left- and right- rear independent drive motors, as illustrated in Figure 8.
The relationship between the desired yaw moment M z and the rear-wheel-drive forces is given by
M z = 1 2 W r 1 λ L R F l r 1 λ R R F r r ,
where W r is the wheel track, and F l r and F r r are the drive forces of the left- and right-rear wheels, respectively.
Additionally, the total longitudinal force must satisfy the acceleration demand of the driver:
1 λ L R F l r + 1 λ R R F r r = F x r r e q ,
where F x r r e q is the total desired longitudinal force determined by the accelerator pedal position.
By solving Equations (46) and (47) simultaneously, the required drive forces for each wheel are derived as
F l r F r r = 1 2 1 λ L R + ϵ F x r r e q + M z W r 1 λ L R + ϵ 1 2 1 λ R R + ϵ F x r r e q M z W r 1 λ L R + ϵ ,
where ϵ is a small positive constant added to prevent division by zero during tire blowout (i.e., when λ i 1 ).
Finally, the target output torque for each motor ( T e l r d e s and T e r r d e s ) is calculated by multiplying the required forces by the wheel radius R w , as expressed by
T e l r d e s T e r r d e s = F l r F r r R w = 1 2 1 λ L R + ϵ F x r r e q R w + M z R w W r 1 λ L R + ϵ 1 2 1 λ L R + ϵ F x r r e q R w M z R w W r 1 λ L R + ϵ ,

4. Simulation and Prototype Tests

4.1. Prototype Development

The prototype is regarded as an effective method for testing and verifying the controllers, as indicated by the research of [60,61]. To verify the effectiveness of the proposed controller, a scaled electric vehicle prototype is developed, as shown in Figure 9.
The body of the prototype is constructed using 3D-printed parts. The steering system is actuated by a ZX300 DC servo motor, while the rear wheels are independently driven by two 8883-EX DC motors. The physical parameters of the prototype are measured experimentally and listed in Table 1.
In both the prototype and the model, a tire is considered a blowout tire when the air pressure of the tire falls below a set threshold P s . This definition applies regardless of the speed of pressure decay. Once the tire pressure is detected below P s , the fault detection is activated immediately. In MATLAB 2020a simulation, the dynamic process of a tire blowout is simulated with the Dugoff tire model according to [60,61,62]. Tire parameters are directly adjusted to represent the dynamic changes in the vehicle after a blowout. In prototype vehicle tests, a pre-set blown-out tire is installed to physically simulate a tire blowout because the development of a remote-controlled actuator for sudden tire blowout is challenging.
The electrical parameters and the controllers of the independent drive motors are detailed in [53], as shown in Table 2.
The Pulse Width Modulation (PWM) frequency of the motor controller is set as 30 Hz. The control system is implemented on a Raspberry Pi 5, which processes vehicle motion data captured by a Sense HAT gyroscope. The accuracies of the gyroscope sensors are within ±10%. The 2-DOF reference model and the BCEA control algorithm are executed locally on the Raspberry Pi using Python 3.9.2.
The system operates via a WIFI-based hardware-in-the-loop configuration. A remote laptop captures driver inputs (i.e., desired longitudinal speed and steering angle) and transmits them to the Raspberry Pi 5. The Raspberry Pi 5 calculates the required extra steering angle and yaw moment and then converts them into motor voltages and drives the motors to execute the maneuver. This architecture is illustrated in Figure 10.
The prototype utilizes an Ackermann steering geometry, as depicted in Figure 11a,b.
Equation (43) is designed according to [53]. The training phase is conducted in MATLAB 2020a. To comprehensively evaluate the controller, the validation scenarios are divided into two categories: (a) sudden tire blowout during high speed and (b) restarting the vehicle from a standstill after a tire blowout. Due to the safety risks and limitations of inducing sudden structural tire failure on the scaled prototype, high-speed blowout dynamics is verified via simulation. The capability of the controller to stabilize and restart the vehicle from a standstill under persistent blowout condition is validated through physical road test with the scaled prototype.

4.2. MATLAB-Based Simulation Environment

4.2.1. Vehicle Modeling

The dynamic model of the prototype is established based on the Dugoff tire model proposed in [62,63,64]. The air pressure of a healthy tire is defined as 0.2 bar for the prototype. The threshold pressure P s is consequently defined as 0.1 bar, representing a 50% reduction from the healthy tire pressure [65]. The parameters are acquired directly from the physical prototype with the test rig shown in Figure 12.
The parameters used in the Dugoff tire model for both blown-out and healthy tires are measured through the test rig shown in Figure 12. The Dugoff tire model requires parameters for vertical stiffness, radius, longitudinal slip stiffness and cornering stiffness. According to the research of [66,67], a blown-out tire retains approximately 6.7% of the vertical stiffness, 70% of the radius, 28% of the longitudinal slip stiffness, and 25% of the cornering stiffness of a healthy tire. In this study, the measured vertical stiffness, radius, longitudinal slip stiffness, front wheel lateral stiffness and rear wheel lateral stiffness of a normal wheel are 335.46 N/m, 0.0375 m, 884.3 N/m, 5.58 N/rad, and 9.82 N/rad, respectively. For a blown-out tire, the corresponding parameters are 18.34 N/m, 0.022 m, 206.2 N/m, 1.19 N/rad, and 2.04 N/rad, respectively. Consequently, the ratios of the blown-out parameters to the healthy parameters are 5.47% for vertical stiffness, 58.67% for radius, 23.32% for longitudinal slip stiffness, 21.33% for front wheel cornering stiffness, and 20.77% for rear wheel cornering stiffness. These parameters align closely with the requirements outlined in [66,67]. Thus, the Dugoff tire model utilized in this study can reflect the authenticity of the experimental results.
In the MATLAB simulation, the vertical stiffness decreases from 335.46 N/m to 18.34 N/m. The radius reduces from 0.0375 m to 0.022 m. The longitudinal slip stiffness drops from 884.3 N/m to 206.2 N/m. Additionally, the cornering stiffnesses of the front and rear tires change from 5.58 N/rad and 9.82 N/rad to 1.19 N/rad and 2.04 N/rad, respectively.
To ensure that the model can accurately reflect real failure modes, the prototype is physically configured and tested under two tire blowout scenarios. These configurations include left-front tire blowout (denoted as “LF”) and simultaneous blowouts of the left-front and right-rear tires (denoted as “LF+RR”), as illustrated in Figure 13.
The energy released during a tire blowout is calculated by
δ E = P ¯ Δ V   ,
where δ E represents the energy loss of the blown-out tire, P ¯ represents the average air pressure of the tire during the blowout process, and Δ V represents the volumetric change in the air in the tire. In the prototype, P ¯ equals 0.15 bar and Δ V equals 0.86 × 10−5 m3. Thus, δ E is approximately 0.13 J. This energy value is minimal and can be considered negligible in the context of the lateral dynamics.

4.2.2. Training Process and Stability Analysis

During each training episode, the vehicle is subjected to randomized, continuous steering signals to simulate diverse driving intentions (see Figure 14a for example). The vehicle speed is maintained at a constant 1.5 m/s, which corresponds to a full-scale vehicle speed of approximately 90 km/h. At the beginning of each training episode, the conditions of tire blowouts are randomly set. These conditions include the number of blown-out tires (one or two), the positions of these tires, and the time of blowout. The time of blowout is randomly selected between 0 and 20 s.
To ensure system stability, a Lyapunov function L is defined based on the tracking errors:
L = 1 2 Δ r z 2 + 1 2 Δ β 2
Utilizing Lyapunov’s method, the stability criteria for the training process are twofold:
(a)
For any non-zero errors Δ r z and Δ β , the function L must be non-negative ( L 0 ).
(b)
Throughout the training process, the average value of L must exhibit a decreasing trend, with its standard deviation remaining within a bounded range.
To train the controller for robustness, tire blowout scenarios are introduced stochastically. A maximum of two tires is randomly applied to one episode, with a random timing and wheel locations. The training and testing scenarios encompass two distinct failure categories (“LF” blowout and “LF+RR” blowout).
To verify the training efficiency of the BCEA, three configurations are implemented and evaluated, including BCEA for ASC only (denoted as “BCEA+ASC, No DYC”), BCEA for DYC only (denoted as “BCEA+DYC, No ASC”) and BCEA for both ASC and DYC (denoted as “BCEA+ASC+DYC”).
The training framework is implemented in MATLAB. The key hyperparameters include 100 hidden nodes for the EAN, 36 featured nodes and 1080 enhancement nodes for the BCN. The training duration is set to 2000 episodes.
The input of front wheel steering angle is shown in Figure 13a. The selection of 2000 episodes represents an optimal balance between performance and computational efficiency. As listed in Table 3, insufficient training leads to poor control performance. Conversely, exceeding 2000 episodes yields negligible performance gains.
The training process and results are shown in Figure 13b–d. The average reward demonstrates a continuous upward trend, while the average Lyapunov value consistently decreases. This inverse relationship confirms that the controller is effectively minimizing tracking errors and constraining the system dynamics within a stable region.
During the training process, fluctuations are observed in the Lyapunov function values, and these fluctuations are primarily driven by three factors. First, the location and time of tire blowouts are randomized. Second, the cuckoo search algorithm uses an exploration mechanism to avoid being trapped in local optima. Third, the vehicle system exhibits strong nonlinearity, and the complex interaction between stability and tire dynamics causes small state deviations to produce significant variations in the Lyapunov value. The fluctuations remain small and close to zero, which means that the system tends to have only minor errors. These errors are considered controllable. As a result, the small fluctuations are not critical, and the system is believed to be stable.
As traditional discrete algorithms like Q-learning and DQN are unsuitable for this continuous control tasks, deep deterministic policy gradient (DDPG) is selected as the baseline alongside BCDA algorithm [51]. As detailed in Table 4, the proposed BCEA algorithm achieves maximum reward significantly faster than both DDPG and the BCDA architectures. Specifically, BCEA improves the training time by approximately 83.58% compared to DDPG for the combined ASC+DYC configuration.

4.2.3. Simulation Scenarios

This study exclusively examines the impact of a blown-out tire on the dynamic stability of the vehicle system. It is acknowledged that in real-world scenarios, human driver responses, such as sudden panic steering or emergency braking, can pose greater dangers than tire failure itself. However, to individually exam the performance of the proposed controller, this research focuses on vehicle dynamics and does not consider active driver intervention or panic behavior.
Regarding tire blowout conditions, two typical scenarios are considered. The first scenario is straight-line driving at a high speed when a tire suddenly fails. This scenario tests vehicle stability under the proposed control methods. The second scenario is single lane change (SLC). This is a standard test to evaluate the vehicle steering performance and stability.
The proposed controller is compared with several controllers. To evaluate the effectiveness of the combined ASC and DYC controller, vehicles only with ASC or DYC are tested. Thus, the following vehicle configurations are examined in both simulations and prototype tests:
(1)
Vehicle without any controller (denoted as “No controller”);
(2)
Vehicle with a traditional SMC controller for ASC only (denoted as “SMC+ASC, no DYC”);
(3)
Vehicle with a traditional SMC controller for DYC only (denoted as “SMC+DYC, no ASC”);
(4)
Vehicle with a traditional SMC controller for both ASC and DYC (denoted as “SMC+ASC+DYC”);
(5)
Vehicle with a BCEA controller for ASC only (denoted as “BCEA+ASC, no DYC”);
(6)
Vehicle with a BCEA controller for DYC only (denoted as “BCEA+DYC, no ASC”);
(7)
Vehicle with a BCEA controller for both ASC and DYC (denoted as “BCEA+ASC+DYC”).
The SMC is selected as compared controllers due to its widespread application in vehicle control systems [14]. SMC is also regarded as one of the most effective methods for vehicle control due to its rapid response and strong robustness against uncertainties and disturbances.

4.3. Simulation Results

The initial longitudinal speed is 1.5 m/s, which is regarded as the same as 90 km/h in a real vehicle. Due to the challenges and safety concerns of conducting sudden, high-speed tire blowout tests on physical prototypes, these short-term failure conditions are studied through simulation. The simulations are performed using MATLAB-Simulink. Several driving scenarios are included in the simulation setups. The tire blowout is set to occur at 4 s. The initial vehicle speed is 1.5 m/s, which corresponds to 90 km/h in real vehicle conditions.
(a)
Straight-line driving scenario
The first test case evaluates the vehicle stability during straight-line motion at a constant speed. Figure 15 and Figure 16 illustrate the longitudinal speed, yaw rate, side-slip angle, and driving trajectory under two specific failure conditions. These include LF tire blowout and LF+RR tire blowout.
As shown in Figure 15 and Figure 16, the proposed “BCEA+ASC+DYC” controller outperforms other BCEA-based RL controllers and SMC-based controllers for both vehicles with “LF” and with “LF+RR” for the straight-line scenario.
To quantify the control performance, the root mean square error (RMSE) is adopted as the evaluation metric. The RMSEs for longitudinal speed ( R M S E u ) and trajectory tracking ( R M S E t r a ) are defined as
{ R M S E u = i = 1 N u i u x r e f | i 2 N R M S E t r a = i = 1 N x i x r e f | i 2 + y i y r e f | i 2 N
The comparative results are listed in Table 5.
Table 5 shows that the integrated “BCEA+ASC+DYC” achieves the optimal control performance. Under the “LF” blowout condition, longitudinal speed tracking accuracy improves by 51.52%, and trajectory accuracy improves by 80%. Under the “LF+RR” blowout condition, these improvements reach 54.12% and 94.18%, respectively. This confirms the superior performance of the proposed method in maintaining stability under sudden failure.
(b)
Single lane change (SLC) scenario
SLC maneuver represents a critical dynamic scenario. In this test, the vehicle velocity decelerates from 1.5 m/s to 0.3 m/s (kinematically equivalent to a full-scale deceleration from 90 km/h to 18 km/h). The input steering input profile is illustrated in Figure 17.
Figure 18 and Figure 19 display the simulation results for the “LF” and “LF+RR” blowout conditions during the SLC scenario.
As shown in Figure 18 and Figure 19, the proposed “BCEA+ASC+DYC” controller outperforms other BCEA-based RL controllers and SMC-based controllers for both vehicles with “LF” and with “LF+RR” for the SLC scenario. The RMSEs for the SLC scenario are listed in Table 6.
In the SLC scenario, the collaborative “BCEA+ASC+DYC” controller significantly outperforms conventional SMC method. Under the “LF” blowout condition, improvements of 72.40% in speed tracking and 43.68% in trajectory tracking are observed. Under the “LF+RR” blowout condition, the maximum improvement in longitudinal speed tracking and trajectory tracking are 30.52% and 90.33%, respectively, as compared with no controller, demonstrating enhanced robustness.

4.4. Scaled Prototype Tests

It is necessary to validate the proposed control strategies through a physical test on the prototype. The experiments utilize the high-precision gyroscope within the Sense HAT module to record the real-time trajectory of the vehicle. Due to current experimental constraints, installing a tire blowout simulation device on the prototype vehicle is not feasible. Such a device involves complex mechanical structures, electro-pneumatic and hydraulic systems, presenting significant challenges in design, manufacturing, and calibration. Consequently, the prototype tests are conducted exclusively with tires in a pre-set blown-out state.
The prototype road tests were conducted in a big room. The test area covers approximately 200 m2. The friction coefficient of the road surface is 0.27.
The total delay of the system is about 200 ms. This delay comes mainly from signal transmission, execution of the BCEA algorithm and response of the motor command. The value falls within the typical delay range of the drive-by-wire chassis system and does not exceed the requirements of the maximum delay that are recognized by the automotive industry [68,69]. The sampling period of the tire pressure sensor is set to 40 ms.
The theoretical efficiency for achieving a real-time safety margin is guaranteed through three aspects. First, the controller is trained by a single-layer neural network that has a fast computation speed. The measured computation delay accounts for only 5 percent of the total delay. Second, the control strategy is effective for both normal driving and tire blowout conditions. The proposed controller always provides collaborative active steering and direct yaw moment control, which ensures continuous and reliable control responses in both normal and tire blowout situations. Finally, scaled prototype road tests verify that the overall system delay is only 200 milliseconds, which is significantly lower than the critical requirement of 500 milliseconds in the automotive industry. This demonstrates that the theoretical effectiveness is sufficient, thereby ensuring an adequate real-time safety margin.
Two specific extreme conditions are selected for evaluation, including a “LF” blowout condition and a “LF+RR” blowout condition. The tire blowouts are applied prior to the tests, as shown in Figure 13, because instant tire failure cannot be achieved during actual road tests. These faults are tested under two dynamic maneuvers: (a) J-turning and (b) SLC, which represent typical driving scenarios where a vehicle must be guided to a maintenance area for repair. The comparative trajectory results for these tests are presented in Figure 20.
As shown in Figure 20, the proposed “BCEA+ASC+DYC” controller outperforms the other BCEA-based and SMC-based controllers in trajectory tracking for both the J-turning and SLC maneuvers.
The quantitative performance is evaluated using the RMSE of the trajectory R M S E t r a as shown in Equation (50). The results for the J-turning and SLC scenarios are listed in Table 7 and Table 8.
In the J-turning scenarios, the proposed “BCEA+ASC+DYC” controller demonstrates superior performance compared to all other controllers. As compared with no controller, the proposed controller achieves trajectory tracking improvements of 96.51% under the “LF” blowout condition and 96.78% under the extreme “LF+RR” blowout condition, respectively.
Similarly, in the SLC scenarios, the controller significantly enhances vehicle stability. Performance gains of 75.65% and 97.67%, as compared with no controller, are recorded for the “LF” and “LF+RR” blowout conditions, respectively. These consistent results confirm the adaptability and effectiveness of the proposed control strategy in real-world environments.

4.5. Discussion

The proposed controller demonstrates superior performance across all tested conditions. Compared to the uncontrolled case, the longitudinal speed tracking accuracy has improved significantly, with gains ranging from 27.38% to 80.25% depending on the specific tire blowout scenarios. Similarly, the trajectory tracking accuracy shows substantial enhancements, ranging from 68.9% to 94.18%. This indicates that the “BCEA+ASC+DYC” controller can swiftly stabilize the vehicle and correct its heading immediately after tire blowouts.
The robustness of the system is particularly evident in the most challenging scenarios: the single “LF” blowout and the diagonal “LF+RR” blowout. In the straight-line driving scenario, the controller improves the trajectory tracking by 43.68% for “LF” and 90.33% for “LF+RR”. In the J-turning scenarios, the performance is improved by 96.51% for “LF” and 96.78% for “LF+RR”. In the SLC scenarios, the improvements of 75.65% and 97.67% are recorded. These consistent results under extreme conditions strongly support the practical engineering applicability of the BCEA algorithm.

5. Conclusions

Tire blowouts are hazardous faults that traditional steering controllers often fail to manage effectively. While ASC and DYC offers a solution, the strong coupling between these systems makes controller design challenging. To address this challenge, this study introduces an adaptive controller named BCEA based on RL.
The BCEA controller offers a simpler design architecture and superior robustness compared to traditional methods. Furthermore, it achieves significantly faster training speed, improving the training efficiency by over 80% compared to traditional DDPG. Both simulation and scaled prototype tests verify the ability of controller to maintain vehicle stability during high-speed blowouts
Despite these promising results, this study has certain limitations. In this study, a simplified suspension model is used in the prototype, and road tests are limited to a few specific conditions. Moreover, the scaled prototype introduces practical errors, including sensor zero-offset accumulative error and time delays. Additionally, the current speed estimation assumes “no wheel slip”, which decreases in accuracy during acceleration or on low-friction surfaces. It is important to note that this study focuses solely on the impact of blown-out tire on vehicle dynamic stability. While real-world driver reactions (such as sudden steering or braking) are critical elements of road safety that can exacerbate dangerous conditions, modeling of human driver behavior is beyond the scope of this paper and is not considered in the current control strategy. In addition, this study only considers the rapid loss of tire pressure. Tire explosion conditions are not considered in this study. Finally, offline training may not fully capture the stochastic nature of real-world driving environments.
Future research will focus on several key areas. First, current hardware will be updated including higher-precision sensors and a more advanced vehicle chassis. Furthermore, algorithm refinements will be investigated considering adaptive control errors to further fine-tune the performance. In addition, the training efficiency can be enhanced further with other methods, such as utilizing more efficient meta-heuristic methods. Future work will involve the development and installation of a controllable tire blowout actuator on the experimental prototype. This actuator will be designed to emulate sudden tire pressure decompression scenarios under real vehicle testing conditions. Subsequently, a remote-control test platform with a first-person view will be established. This platform will serve to evaluate the influence of human driver reactions on the control performance. Tire explosion conditions will also be considered in the future. Finally, a full-size vehicle platform will be constructed to validate the effectiveness of the proposed algorithm in real-world settings and to investigate the impact of environmental time delays on control performance. Ultimately, this study establishes a foundation for robust vehicle stability under critical failure modes, directly enhancing road safety and reducing the potential for economic losses and human injuries.

Author Contributions

Conceptualization, X.W. and P.K.W.; Methodology, X.W. and S.T.; Software, X.W. and H.Q.; Validation, X.W., H.Q., S.T. and J.L.; Formal analysis, X.W., H.Q., S.T. and Z.Y.; Investigation, X.W. and J.L.; Resources, X.W., P.K.W., Z.Y. and J.L.; Data curation, X.W., H.Q. and S.T.; Writing—original draft, X.W.; Writing—review & editing, X.W., P.K.W., H.Q., S.T., Z.Y., J.L. and W.H.; Visualization, X.W., H.Q., S.T., Z.Y. and W.H.; Supervision, P.K.W. and W.H.; Project administration, X.W., P.K.W. and W.H.; Funding acquisition, P.K.W. All authors have read and agreed to the published version of the manuscript.

Funding

This study is funded by the research grant of the University of Macau (Grant No: MYRG-CRG2025-00048-FST).

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding author.

Conflicts of Interest

The authors declare no conflict of interest.

References

  1. Ou, Y.; Kittner, N.; Babaee, S.; Smith, S.J.; Nolte, C.G.; Loughlin, D.H. Evaluating long-term emission impacts of large-scale electric vehicle deployment in the US using a human-Earth systems model. Appl. Energy 2021, 300, 117364. [Google Scholar]
  2. Zang, L.; Sun, J.; Peng, X.; Lin, F.; Deng, Y.; Bai, Y. A Comprehensive Review of Safety Tire Research. Lubricants 2025, 13, 357. [Google Scholar] [CrossRef] [Scilit]
  3. Ismail, A.Y.; Zuntion, R.B.R. Comparative Study on the Kinematic Performance of Front Wheel Drive (FWD) and Rear Wheel Drive (RWD) Vehicles. Media Komunikasi Teknologi. J. IPTEK 2024, 28, 169–177. [Google Scholar] [CrossRef] [Scilit]
  4. Zhao, Y.-e.; Zhang, J.; Han, X. Development of a High Performance Electric Vehicle with Four-Independent-Wheel Drive; Technical Paper No. 2008-01-1829; SAE: Warrendale, PA, USA, 2008. [Google Scholar]
  5. Kusuma, D.H.; Ali, M.; Sutantra, N. The comparison of optimization for active steering control on vehicle using PID controller based on artificial intelligence techniques. In Proceedings of the 2016 International Seminar on Application for Technology of Information and Communication (ISemantic), Semarang, Indonesia, 5–6 August 2016; pp. 18–22. [Google Scholar]
  6. Elmi, N.; Ohadi, A.; Samadi, B. Active front-steering control of a sport utility vehicle using a robust linear quadratic regulator method, with emphasis on the roll dynamics. Proc. Inst. Mech. Eng. Part D J. Automob. Eng. 2013, 227, 1636–1649. [Google Scholar] [CrossRef] [Scilit]
  7. Tian, J.; Zeng, Q.; Wang, P.; Wang, X. Active steering control based on preview theory for articulated heavy vehicles. PLoS ONE 2021, 16, e0252098. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  8. Deng, Z.; Zhao, Q.; Zhao, Y.; Wang, B.; Gao, W.; Kong, X. Active LQR multi-axle-steering method for improving maneu-verability and stability of multi-trailer articulated heavy vehicles. Int. J. Automot. Technol. 2022, 23, 939–955. [Google Scholar] [CrossRef] [Scilit]
  9. Qi, G.; Fan, X.; Zhao, Z. Fuzzy and sliding mode variable structure control of vehicle active steering system. Recent Pat. Mech. Eng. 2021, 14, 226–241. [Google Scholar] [CrossRef] [Scilit]
  10. Li, F.; Wu, Z.; Pal, N.R.; Yang, C.; Peng, H.; Kaynak, O.; Huang, T. Lane-keeping control of automatic steering systems via adaptive fuzzy sliding-mode approach. IEEE Trans. Syst. Man Cybern. Syst. 2023, 54, 1683–1693. [Google Scholar] [CrossRef] [Scilit]
  11. Wu, J.; Zhang, J.; Tian, Y.; Li, L. A novel adaptive steering torque control approach for human–machine cooperation auton-omous vehicles. IEEE Trans. Transp. Electrif. 2021, 7, 2516–2529. [Google Scholar] [CrossRef] [Scilit]
  12. Yang, H.; Liu, W.; Chen, L.; Yu, F. An adaptive hierarchical control approach of vehicle handling stability improvement based on Steer-by-Wire Systems. Mechatronics 2021, 77, 102583. [Google Scholar] [CrossRef] [Scilit]
  13. Zhang, C.; Gao, P.; Wang, J.; Dang, M.; Yang, X.; Feng, Y. Research on active rear-wheel steering control method with sliding mode control optimized by model predictive. IEEE Access 2023, 11, 57228–57239. [Google Scholar] [CrossRef] [Scilit]
  14. Gambhire, S.J.; Kishore, D.R.; Londhe, P.S.; Pawar, S.N. Review of sliding mode based control techniques for control system applications. Int. J. Dyn. Control. 2021, 9, 363–378. [Google Scholar] [CrossRef] [Scilit]
  15. Song, Z.; Wang, L.; Liu, Y.; Wang, K.; He, Z.; Zhu, Z.; Qin, J.; Li, Z. Actively steering a wheeled tractor against potential rollover using a sliding-mode control algorithm: Scaled physical test. Biosyst. Eng. 2022, 213, 13–29. [Google Scholar] [CrossRef] [Scilit]
  16. Liu, J.; Song, J.; Li, H.; Huang, H. Direct yaw-moment control of vehicles based on phase plane analysis. Proc. Inst. Mech. Eng. Part D J. Automob. Eng. 2022, 236, 2459–2474. [Google Scholar] [CrossRef] [Scilit]
  17. Zhang, Z.; Huang, M.; Ji, M.; Zhu, S. Design of the linear quadratic control strategy and the closed-loop system for the active four-wheel-steering vehicle. SAE Int. J. Passeng. Cars-Mech. Syst. 2015, 8, 354–363. [Google Scholar] [CrossRef] [Scilit]
  18. Wang, B.; Zhang, J.; Zhang, Y.; Wang, W. The direct yaw-moment control based on adaptive fuzzy LQR for distributed drive electric vehicles. Adv. Mech. Eng. 2024, 16, 16878132241273524. [Google Scholar] [CrossRef] [Scilit]
  19. Xie, X.; Jin, L.; Baicang, G.; Shi, J. Vehicle direct yaw moment control system based on the improved linear quadratic regulator. Ind. Robot. Int. J. Robot. Res. Appl. 2021, 48, 378–387. [Google Scholar] [CrossRef] [Scilit]
  20. Liu, H.; Yan, S.; Shen, Y.; Li, C.; Zhang, Y.; Hussain, F. Model predictive control system based on direct yaw moment control for 4WID self-steering agriculture vehicle. Int. J. Agric. Biol. Eng. 2021, 14, 175–181. [Google Scholar] [CrossRef] [Scilit]
  21. Stano, P.; Tavernini, D.; Montanaro, U.; Tufo, M.; Fiengo, G.; Novella, L.; Sorniotti, A. Enhanced active safety through inte-grated autonomous drifting and direct yaw moment control via nonlinear model predictive control. IEEE Trans. Intell. Veh. 2023, 9, 4172–4190. [Google Scholar] [CrossRef] [Scilit]
  22. Ma, L.; Cheng, C.; Guo, J.; Shi, B.; Ding, S.; Mei, K. Direct yaw-moment control of electric vehicles based on adaptive sliding mode. Math. Biosci. Eng. 2023, 20, 13334–13355. [Google Scholar] [CrossRef] [Scilit]
  23. Lee, J.E.; Kim, B.W. Research on direct yaw moment control based on neural sliding mode control for four-wheel actuated electric vehicles. In Proceedings of the 2023 IEEE 6th International Conference on Knowledge Innovation and Invention (ICKII), Sapporo, Japan, 11–13 August 2023; pp. 736–740. [Google Scholar]
  24. Liang, J.; Feng, J.; Lu, Y.; Yin, G.; Zhuang, W.; Mao, X. A direct yaw moment control framework through robust TS fuzzy approach considering vehicle stability margin. IEEE/ASME Trans. Mechatron. 2023, 29, 166–178. [Google Scholar] [CrossRef] [Scilit]
  25. Cao, Y.; Xie, Z.; Li, W.; Wang, X.; Wong, P.K.; Zhao, J. Combined path following and direct yaw-moment control for un-manned electric vehicles based on event-triggered T–S fuzzy method. Int. J. Fuzzy Syst. 2024, 26, 2433–2448. [Google Scholar] [CrossRef] [Scilit]
  26. Deng, Z.; Hu, W.; Gao, W.; Guo, Y.; Fan, Q. Radial basis function neural network sliding mode control-based direct yaw moment control for electrically driven vehicles. Int. J. Veh. Perform. 2025, 11, 125–158. [Google Scholar] [CrossRef] [Scilit]
  27. Lee, J.E.; Kim, B.W. Improving direct yaw-moment control via neural-network-based non-singular fast terminal sliding mode control for electric vehicles. Sensors 2024, 24, 4079. [Google Scholar] [CrossRef] [Scilit]
  28. Lee, J.E.; Kim, B.W. A Novel Adaptive Non-Singular Fast Terminal Sliding Mode Control for Direct Yaw Moment Control in 4WID Electric Vehicles. Sensors 2025, 25, 941. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  29. Yang, X.; Wang, Z.; Peng, W. Coordinated control of AFS and DYC for vehicle handling and stability based on optimal guaranteed cost theory. Veh. Syst. Dyn. 2009, 47, 57–79. [Google Scholar] [CrossRef] [Scilit]
  30. Chen, W.; Liang, X.; Wang, Q.; Zhao, L.; Wang, X. Extension coordinated control of four wheel independent drive electric vehicles by AFS and DYC. Control Eng. Pract. 2020, 101, 104504. [Google Scholar] [CrossRef] [Scilit]
  31. Fan, Z.; Wu, Z.; Zhang, R. Coordinated Control Strategy of AFS and DYC Based on Stability Category Recognition. Results Eng. 2025, 27, 106316. [Google Scholar] [CrossRef] [Scilit]
  32. Zhang, N.; Wang, J.; Li, Z.; Xu, N.; Ding, H.; Zhang, Z.; Xu, H. Coordinated optimal control of AFS and DYC for four-wheel independent drive electric vehicles based on MAS model. Sensors 2023, 23, 3505. [Google Scholar] [CrossRef] [Scilit]
  33. Zhang, Y.; Wong, P.K.; Li, W.; Cao, Y.; Xie, Z.; Zhao, J. Fault diagnosis and fault tolerant control for distributed drive electric vehicles through integration of active front steering and direct yaw moment control. Mechatronics 2024, 97, 103116. [Google Scholar] [CrossRef] [Scilit]
  34. Ji, Y.; Zhang, J.; Lv, C.; He, C.; Hou, X.; Han, J. Fault-tolerant vehicle stability control based on active steering and direct yaw moment with finite-time constraint performance recovery. IEEE Trans. Veh. Technol. 2023, 72, 15317–15329. [Google Scholar] [CrossRef] [Scilit]
  35. Zhu, S.; Li, H.; Wang, G.; Kuang, C.; Chen, H.; Gao, J.; Xie, W. Research on fault-tolerant control of distributed-drive electric vehicles based on fuzzy fault diagnosis. Actuators 2023, 12, 246. [Google Scholar] [CrossRef] [Scilit]
  36. Lu, J.; Wang, R.; Wan, C. Intelligent vehicle path tracking and stability control method based on extension coordinated control of AFS and DYC. IEEE Access 2025, 13, 7353–7365. [Google Scholar] [CrossRef] [Scilit]
  37. Kostrikov, I.; Nair, A.; Levine, S. Offline reinforcement learning with implicit q-learning. arXiv 2021, arXiv:2110.06169. [Google Scholar] [CrossRef] [Scilit]
  38. Zhou, Q.; Lian, Y.; Wu, J.; Zhu, M.; Wang, H.; Cao, J. An optimized Q-Learning algorithm for mobile robot local path planning. Knowl. Based Syst. 2024, 286, 111400. [Google Scholar] [CrossRef] [Scilit]
  39. Li, J.; Chen, Y.; Zhao, X.; Huang, J. An improved DQN path planning algorithm. J. Supercomput. 2022, 78, 616–639. [Google Scholar] [CrossRef] [Scilit]
  40. Carta, S.; Ferreira, A.; Podda, A.S.; Recupero, D.R.; Sanna, A. Multi-DQN: An ensemble of Deep Q-learning agents for stock market forecasting. Expert Syst. Appl. 2021, 164, 113820. [Google Scholar] [CrossRef] [Scilit]
  41. Hazem, Z.B.; Saidi, F.; Guler, N.; Altaif, A.H. Reinforcement learning-based intelligent trajectory tracking for a 5-DOF Mitsubishi robotic arm: Comparative evaluation of DDPG, LC-DDPG, and TD3-ADX. Int. J. Intell. Robot. Appl. 2025, 9, 1–21. [Google Scholar] [CrossRef] [Scilit]
  42. He, N.; Yang, S.; Li, F.; Trajanovski, S.; Kuipers, F.A.; Fu, X. A-DDPG: Attention mechanism-based deep reinforcement learning for NFV. In Proceedings of the 2021 IEEE/ACM 29th International Symposium on Quality of Service (IWQOS), Tokyo, Japan, 25–28 June 2021; pp. 1–10. [Google Scholar]
  43. Zhou, J.; Xue, S.; Xue, Y.; Liao, Y.; Liu, J.; Zhao, W. A novel energy management strategy of hybrid electric vehicle via an improved TD3 deep reinforcement learning. Energy 2021, 224, 120118. [Google Scholar] [CrossRef] [Scilit]
  44. Zhang, F.; Li, J.; Li, Z. A TD3-based multi-agent deep reinforcement learning method in mixed cooperation-competition environment. Neurocomputing 2020, 411, 206–215. [Google Scholar] [CrossRef] [Scilit]
  45. Yu, C.; Velu, A.; Vinitsky, E.; Gao, J.; Wang, Y.; Bayen, A.; Wu, Y. The surprising effectiveness of ppo in cooperative multi-agent games. Adv. Neural Inf. Process. Syst. 2022, 35, 24611–24624. [Google Scholar]
  46. Engstrom, L.; Ilyas, A.; Santurkar, S.; Tsipras, D.; Janoos, F.; Rudolph, L.; Madry, A. Implementation matters in deep RL: A case study on PPO and TRPO. arXiv 2020, arXiv:2005.12729. [Google Scholar]
  47. Jafari, R.; Sarhadi, P.; Paykani, A.; Refaat, S.S.; Asef, P. A TD3-based reinforcement learning algorithm with curriculum learning for adaptive yaw stability control in all-wheel-drive electric vehicles. IEEE Access 2025, 13, 127150–127169. [Google Scholar] [CrossRef] [Scilit]
  48. Pang, H.; Huang, H.; Fan, Y.; Yao, L.; Chen, Y. Research on vehicle lateral stability control under low-adhesion road conditions using proximal policy optimization algorithm. PLoS ONE 2025, 20, e0335686. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  49. Bari, G.; Palkovics, L. Self driving algorithm for an active four wheel drive racecar. arXiv 2025, arXiv:2506.06077. [Google Scholar] [CrossRef] [Scilit]
  50. Wang, H.; Wang, C.; Chen, W.; Zheng, W. Path Tracking Coordination Control with TD3 Reinforcement Learning Based on Danger Level Recognition Considering Generalized Functional Safety. IEEE Trans. Veh. Technol. 2024, 73, 18498–18511. [Google Scholar] [CrossRef] [Scilit]
  51. Thalagala, S.; Wong, P.K.; Wang, X.; Sun, T. Broad Critic Deep Actor Reinforcement Learning for Continuous Control. IEEE Trans. Neural Netw. Learn. Syst. 2025, 36, 17508–17515. [Google Scholar] [CrossRef] [Scilit]
  52. Chen, C.P.; Liu, Z. Broad learning system: An effective and efficient incremental learning system without the need for deep architecture. IEEE Trans. Neural Netw. Learn. Syst. 2017, 29, 10–24. [Google Scholar] [CrossRef] [Scilit]
  53. Wong, P.K.; Wang, X.; Thalagala, S. Design and Experimental Evaluation of a Novel Multi-axle Electric Truck with Broad-Deep Learning-Based Collaborative Controllers for Active Steering and Direct Yaw Moment Control. Int. J. Automot. Technol. 2025, 1–38. [Google Scholar] [CrossRef] [Scilit]
  54. Gong, X.; Zhang, T.; Chen, C.P.; Liu, Z. Research review for broad learning system: Algorithms, theory, and applications. IEEE Trans. Cybern. 2021, 52, 8922–8950. [Google Scholar] [CrossRef] [Scilit]
  55. Alhijawi, B.; Awajan, A. Genetic algorithms: Theory, genetic operators, solutions, and applications. Evol. Intell. 2024, 17, 1245–1256. [Google Scholar] [CrossRef] [Scilit]
  56. Yang, X.S.; Deb, S. Cuckoo search: Recent advances and applications. Neural Comput. Appl. 2014, 24, 169–174. [Google Scholar] [CrossRef] [Scilit]
  57. Guerrero-Luis, M.; Valdez, F.; Castillo, O. A review on the cuckoo search algorithm. In Fuzzy Logic Hybrid Extensions of Neural and Optimization Algorithms: Theory and Applications; Springer: Berlin/Heidelberg, Germany, 2021; pp. 113–124. [Google Scholar]
  58. Wang, J.; Lu, S.; Wang, S.H.; Zhang, Y.D. A review on extreme learning machine. Multimed. Tools Appl. 2022, 81, 41611–41660. [Google Scholar] [CrossRef] [Scilit]
  59. Wang, X.; Zang, L.; Wang, Z.; Lin, F.; Zhao, Z. Study on the stability control of vehicle tire blowout based on run-flat tire. World Electr. Veh. J. 2021, 12, 128. [Google Scholar] [CrossRef] [Scilit]
  60. Verma, R.; Del Vecchio, D.; Fathy, H.K. Development of a scaled vehicle with longitudinal dynamics of an HMMWV for an ITS testbed. IEEE/ASME Trans. Mechatron. 2008, 13, 46–57. [Google Scholar] [CrossRef] [Scilit]
  61. Hasnain, S.G. Design and Construction of a Dynamically Scaled Vehicle for Emergency Scenario Algorithm Development. Master’s Thesis, Technische Hochschule Ingolstadt, Ingolstadt, Germany, 2021. [Google Scholar]
  62. Li, A.; Chen, Y.; Du, X.; Lin, W.C. Enhanced tire blowout modeling using vertical load redistribution and self-alignment torque. ASME Trans. J. Dyn. Syst. Control. 2021, 1, 011001. [Google Scholar] [CrossRef] [Scilit]
  63. Li, A.; Chen, Y.; Du, X.; Lin, W.C. Should a vehicle always deviate to the tire blowout side?—A new tire blowout model with toe angle effects. J. Dyn. Syst. Meas. Control. ASME Trans. 2021, 143, 101008. [Google Scholar] [CrossRef] [Scilit]
  64. Belrzaeg, M.; Ahmed, A.A.; Almabrouk, A.Q.; Khaleel, M.M.; Ahmed, A.A.; Almukhtar, M. Vehicle dynamics and tire models: An overview. World J. Adv. Res. Rev. 2021, 12, 331–348. [Google Scholar] [CrossRef] [Scilit]
  65. Elfasakhany, A. Tire pressure checking framework: A review study. Reliab. Eng. Resil. 2019, 1, 12–28. [Google Scholar]
  66. Liu, H.; Deng, W.; Zong, C.; Wu, J. Development of Active Control Strategy for Flat Tire Vehicles (No. 2014-01-0859); SAE Technical Paper; SAE: Warrendale, PA, USA, 2014. [Google Scholar]
  67. Guo, H.; Wang, F.; Chen, H.; Guo, D. Stability control of vehicle with tire blowout using differential flatness based mpc method. In Proceedings of the 10th World Congress on Intelligent Control and Automation, Beijing, China, 6–8 July 2012; IEEE: New York, NY, USA; pp. 2066–2071.
  68. Kim, E.; Hwang, M.; Lim, T.; Jeong, C.; Yoon, S.; Cha, H. Communication delay outlier detection and compensation for teleoperation using stochastic state estimation. Sensors 2024, 24, 1241. [Google Scholar] [CrossRef] [Scilit]
  69. Noomwongs, N.; Siriwattana, K.T.; Chantranuwathana, S.; Phanomchoeng, G. The Development of Teleoperated Driving to Cooperate with the Autonomous Driving Experience. Automation 2025, 6, 26. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Structure of proposed controller.
Figure 1. Structure of proposed controller.
Vehicles 08 00040 g001
Figure 2. 2-DOF reference model.
Figure 2. 2-DOF reference model.
Vehicles 08 00040 g002
Figure 3. Fundamental process of RL.
Figure 3. Fundamental process of RL.
Vehicles 08 00040 g003
Figure 4. Flowchart of BCEA algorithm.
Figure 4. Flowchart of BCEA algorithm.
Vehicles 08 00040 g004
Figure 5. Structure of EAN and t-EAN.
Figure 5. Structure of EAN and t-EAN.
Vehicles 08 00040 g005
Figure 6. Architecture of improved BCN.
Figure 6. Architecture of improved BCN.
Vehicles 08 00040 g006
Figure 7. Flowchart of cuckoo search for target action optimization.
Figure 7. Flowchart of cuckoo search for target action optimization.
Vehicles 08 00040 g007
Figure 8. Yaw moment generation via rear independent drive.
Figure 8. Yaw moment generation via rear independent drive.
Vehicles 08 00040 g008
Figure 9. Structure of prototype vehicle.
Figure 9. Structure of prototype vehicle.
Vehicles 08 00040 g009
Figure 10. Control system architecture of prototype.
Figure 10. Control system architecture of prototype.
Vehicles 08 00040 g010
Figure 11. Structure of front wheel steering system.
Figure 11. Structure of front wheel steering system.
Vehicles 08 00040 g011aVehicles 08 00040 g011b
Figure 12. Test rigs for tires.
Figure 12. Test rigs for tires.
Vehicles 08 00040 g012
Figure 13. Prototype configurations under various tire blowout scenarios. (a) Left-front tire blowout (side view); (b) Left-front tire blowout (front view); (c) Left-front and right-rear blowout (left side view); (d) Left-front and right-rear blowout (right side view).
Figure 13. Prototype configurations under various tire blowout scenarios. (a) Left-front tire blowout (side view); (b) Left-front tire blowout (front view); (c) Left-front and right-rear blowout (left side view); (d) Left-front and right-rear blowout (right side view).
Vehicles 08 00040 g013
Figure 14. Training input and results of BCEA. (a) Example of input front wheel steering angle; (b) Average Reward; (c) Average Lyapunov’s value; (d) Standard deviation of Lyapunov’s Value.
Figure 14. Training input and results of BCEA. (a) Example of input front wheel steering angle; (b) Average Reward; (c) Average Lyapunov’s value; (d) Standard deviation of Lyapunov’s Value.
Vehicles 08 00040 g014
Figure 15. Simulation results for straight-line driving under LF tire blowout.
Figure 15. Simulation results for straight-line driving under LF tire blowout.
Vehicles 08 00040 g015
Figure 16. Simulation results for straight-line driving under LF+RR tire blowout.
Figure 16. Simulation results for straight-line driving under LF+RR tire blowout.
Vehicles 08 00040 g016aVehicles 08 00040 g016b
Figure 17. Front wheel steering angle profile for the SLC maneuver.
Figure 17. Front wheel steering angle profile for the SLC maneuver.
Vehicles 08 00040 g017
Figure 18. Simulation results for SLC maneuver under “LF” blowout.
Figure 18. Simulation results for SLC maneuver under “LF” blowout.
Vehicles 08 00040 g018
Figure 19. Simulation results for SLC maneuver under “LF+RR” blowout.
Figure 19. Simulation results for SLC maneuver under “LF+RR” blowout.
Vehicles 08 00040 g019
Figure 20. Trajectories of prototype.
Figure 20. Trajectories of prototype.
Vehicles 08 00040 g020aVehicles 08 00040 g020b
Table 1. Parameters of prototype vehicle.
Table 1. Parameters of prototype vehicle.
ParameterSymbolValue
Mass m s 2.06 k g
Moment of inertia I z z 0.017 k g · m 2
Distance from front axle to C.G. L f 0.114   m
Distance from rear axle to C.G. L r 0.066   m
Wheel radius R w Before tire blowout: 0.0375 m
After tire blowout: 0.022 m
Cornering stiffness of front wheels C f Before tire blowout: 5.58 N / rad
After tire blowout: 1.19 N / rad
Cornering stiffness of rear wheels C r Before tire blowout: 9.82 N / rad
After tire blowout: 2.04 N / rad
Wheel track W r 0.164 m
Table 2. Parameters of electric motors.
Table 2. Parameters of electric motors.
ParameterValue
Winding resistance29.59 Ω
Winding inductance2.78 H
Back EMF coefficient0.091 V · s
Torque constant0.175 N · m · A 1
Damping coefficient0.0005 N · m · s
Table 3. Training performance metrics for BCEA at various stages.
Table 3. Training performance metrics for BCEA at various stages.
No. of EpisodesController ConfigurationAverage RewardAverage Lyapunov ValueStandard Deviation of Lyapunov Value
1BCEA+ASC+DYC−16.1211.13669.62
BCEA+ASC, no DYC−18.312.78597.27
BCEA+DYC, no ASC−1.4936.66407.51
500BCEA+ASC+DYC−3.930.267136.53
BCEA+ASC, no DYC−12.330.66321.84
BCEA+DYC, no ASC14.633.565103.75
1000BCEA+ASC+DYC14.230.15245.729
BCEA+ASC, no DYC4.7430.2173.409
BCEA+DYC, no ASC17.210.61345.723
2000BCEA+ASC+DYC45.530.25115.69
BCEA+ASC, no DYC12.4670.1844.778
BCEA+DYC, no ASC35.910.31626.617
Table 4. Comparison of training time to reach maximum reward with DDPG algorithm.
Table 4. Comparison of training time to reach maximum reward with DDPG algorithm.
MethodController ConfigurationTraining Time (s)Improvement Over DDPG (%)
DDPGASC+DYC5497.99/
ASC4556.54/
DYC5030.25/
BCDAASC+DYC1386.4474.78%
ASC1100.9275.84%
DYC1011.1479.90%
BCEAASC+DYC903.0383.58%
ASC909.5880.04%
DYC930.2381.51%
Table 5. Performance comparison for straight-line scenario ( R M S E u s and R M S E t r a s).
Table 5. Performance comparison for straight-line scenario ( R M S E u s and R M S E t r a s).
Blowout ScenarioController R M S E u Improvement of
R M S E u Compared to “No Controller”
R M S E t r a Improvement of
R M S E t r a Compared to “No Controller”
LFNo controller0.033N/A0.875N/A
SMC+ASC, No DYC0.042−27.27%0.32562.86%
SMC+DYC, No ASC0.121−266.67%0.31863.66%
SMC+ASC+DYC0.185−460.61%0.57534.29%
BCEA+ASC, No DYC0.081−145.45%0.28167.89%
BCEA+DYC, No ASC0.146−342.42%0.34161.03%
BCEA+ASC+DYC0.01651.52%0.17580.00%
LF + RRNo controller0.085N/A5.868N/A
SMC+ASC, No DYC0.0841.18%0.48591.73%
SMC+DYC, No ASC0.196−130.59%10.555−40.15%
SMC+ASC+DYC0.0823.53%0.48393.59%
BCEA+ASC, No DYC0.191−124.71%2.41167.99%
BCEA+DYC, No ASC0.098−15.29%1.12585.06%
BCEA+ASC+DYC0.03954.12%0.43894.18%
Table 6. Performance comparison for SLC scenario ( R M S E u and R M S E t r a ).
Table 6. Performance comparison for SLC scenario ( R M S E u and R M S E t r a ).
Blowout ScenarioController R M S E u Improvement of
R M S E u Compared to “No Controller”
R M S E t r a Improvement of
R M S E t r a Compared to “No Controller”
LFNo controller0.366N/A0.261N/A
SMC+ASC, No DYC0.18150.55%0.20122.99%
SMC+DYC, No ASC0.17352.73%0.318−21.84%
SMC+ASC+DYC0.30117.76%0.426−63.22%
BCEA+ASC, No DYC0.16754.37%0.18329.89%
BCEA+DYC, No ASC0.15358.20%0.18628.74%
BCEA+ASC+DYC0.10172.40%0.14743.68%
LF + RRNo control0.213N/A2.058N/A
SMC+ASC, No DYC0.2034.69%0.31784.60%
SMC+DYC, No ASC0.246−15.49%0.46877.26%
SMC+ASC+DYC0.262−23.00%0.48976.24%
BCEA+ASC, No DYC0.1977.51%0.27886.49%
BCEA+DYC, No ASC0.269−26.29%0.53973.81%
BCEA+ASC+DYC0.14830.52%0.19990.33%
Table 7. Experimental results for J-turning scenario ( R M S E t r a ).
Table 7. Experimental results for J-turning scenario ( R M S E t r a ).
Blowout ScenarioController R M S E t r a Improvement of
R M S E t r a Compared to “No Controller”
LFNo controller9.275N/A
SMC+ASC, No DYC0.45995.05%
SMC+DYC, No ASC0.71692.29%
SMC+ASC+DYC2.76870.16%
BCEA+ASC, No DYC0.47394.91%
BCEA+DYC, No ASC0.58193.74%
BCEA+ASC+DYC0.32496.51%
LF and RRNo controller6.498N/A
SMC+ASC, No DYC2.59860.01%
SMC+DYC, No ASC0.49992.32%
SMC+ASC+DYC1.07983.39%
BCEA+ASC, No DYC0.85986.79%
BCEA+DYC, No ASC3.95339.17%
BCEA+ASC+DYC0.20996.78%
Table 8. Experimental results for SLC scenario ( R M S E t r a ).
Table 8. Experimental results for SLC scenario ( R M S E t r a ).
Blowout ScenarioController R M S E t r a Improvement of
R M S E t r a Compared to “No Controller”
LFNo controller0.918N/A
SMC+ASC, No DYC0.52143.32%
SMC+DYC, No ASC0.23974.04%
SMC+ASC+DYC1.381−50.43%
BCEA+ASC, No DYC0.33463.66%
BCEA+DYC, No ASC0.79113.87%
BCEA+ASC+DYC0.22375.65%
LF and RRNo controller7.215N/A
SMC+ASC, No DYC0.75989.48%
SMC+DYC, No ASC0.66290.83%
SMC+ASC+DYC0.61691.46%
BCEA+ASC, No DYC0.74589.67%
BCEA+DYC, No ASC0.93387.06%
BCEA+ASC+DYC0.16997.67%
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Wang, X.; Wong, P.K.; Qi, H.; Thalagala, S.; Yang, Z.; Lu, J.; Huang, W. Collaborative Control of Rear-Wheel Independent Drive Electric Vehicles During Tire Blowouts Using Broad-Extreme Reinforcement Learning: Simulation and Scaled Prototype Verification. Vehicles 2026, 8, 40. https://doi.org/10.3390/vehicles8020040

AMA Style

Wang X, Wong PK, Qi H, Thalagala S, Yang Z, Lu J, Huang W. Collaborative Control of Rear-Wheel Independent Drive Electric Vehicles During Tire Blowouts Using Broad-Extreme Reinforcement Learning: Simulation and Scaled Prototype Verification. Vehicles. 2026; 8(2):40. https://doi.org/10.3390/vehicles8020040

Chicago/Turabian Style

Wang, Xiaozheng, Pak Kin Wong, Hengli Qi, Shiron Thalagala, Ziqi Yang, Jingyu Lu, and Wei Huang. 2026. "Collaborative Control of Rear-Wheel Independent Drive Electric Vehicles During Tire Blowouts Using Broad-Extreme Reinforcement Learning: Simulation and Scaled Prototype Verification" Vehicles 8, no. 2: 40. https://doi.org/10.3390/vehicles8020040

APA Style

Wang, X., Wong, P. K., Qi, H., Thalagala, S., Yang, Z., Lu, J., & Huang, W. (2026). Collaborative Control of Rear-Wheel Independent Drive Electric Vehicles During Tire Blowouts Using Broad-Extreme Reinforcement Learning: Simulation and Scaled Prototype Verification. Vehicles, 8(2), 40. https://doi.org/10.3390/vehicles8020040

Article Metrics

Back to TopTop