1. Introduction
Cold standby voting systems are fault-tolerant architectures featuring high reliability, flexibility, and scalability. They are widely employed in aerospace systems, nuclear power plant control, and engineering equipment [
1,
2,
3]. The efficient calculation of reliability indices for these systems remains a key area of research. The cold standby voting system
comprises
operational and
cold standby units, where normal operation is maintained when at least
units are functioning properly. When the number of normally working units decreases from
to
due to component failures, cold standby units can be activated and replace the failed ones, ensuring continued system reliability.
allows for the lifetime distribution of the units to follow any distribution, not limited to exponential distribution, enhancing its relevance to real-world engineering applications.
Extensive studies have been conducted on the reliability of cold standby voting systems. Yang Lechang [
4] developed a survival signature-based reliability approach for a multi-state system with both dependence and imprecision developed. The survival function was derived through structural reliability treatments and calculated by two numerical simulation algorithms. Zong Xiaofei [
5] analyzed the mean time to failure of such systems using order statistics with a single cold standby, obtaining the system’s mean life under the Weibull distribution via numerical integration. Hu Xinchun [
6] derived an analytical formula for the reliability of heterogeneous cold standby voting systems with a single spare part, where the working components follow any distribution. Gu Ruoxing [
7] examined typical series, parallel, and voting systems, applying order statistics theory to directly establish a system reliability model that accounts for failure correlations in common load failures. Eryilmaz [
8] introduced a weighted reliability calculation method for cold standby voting systems, incorporating components with mixed lifetimes, including two types of active components and one cold standby unit. Zhu Qiyue [
9] explored the activation sequence of cold standby parts in the
cold standby voting system, concluding that in the homogeneous
cold standby voting system, the delayed activation of cold standby parts improves overall system reliability, though no formal proof was provided. Achintya Roy [
10] analyzed the reliability of k-out-of-n systems with two cold standby components, building upon the work of Eryilmaz [
11]. Anna [
12] developed a Monte Carlo simulation algorithm using stochastic process theory to calculate the mean life of weighted
voting systems and continuous weighted
voting systems. Ioannis [
13] derived exact formulas for the reliability indices of combined systems consisting of
consecutive
voting systems and consecutive
voting systems and analyzed the impact of design parameters
,
,
and
on system performance through extensive numerical experiments. Eryilmaz [
14] also derived expressions for the survival function and mean time to failure of coherent systems with cold standby components, considering both exponential and non-exponential distributions (e.g., Pareto distribution). Regarding spare part activation strategies, Zhang Bowei [
15] proposed a maintenance decision-making method that optimizes spare part ordering and component replacement by predicting the remaining life to trigger spare part deployment, based on intelligent predictions using the C-MAPSS dataset.
Previous studies have provided important insights into reliability modeling and evaluation for cold standby voting systems. However, analytical solutions remain difficult to obtain when lifetimes follow the Weibull distribution, as obtaining analytical expressions for system reliability indices proves difficult. This limitation restricts the practical applicability of the Weibull distribution in such systems. Therefore, this study focuses on determining the optimal spare part activation strategy and calculating reliability indices under the assumption that both component and spare part lifetimes follow the Weibull distribution, aiming to offer valuable theoretical insights and methodological references for the design and application of high-reliability systems.
2. Research on Spare Part Activation Strategy of Cold Standby Systems
2.1. Description of Cold Spare Systems
The cold standby voting system comprises operational and cold standby units, where normal operation is maintained when at least units are functioning properly. Let represent the critical number of failures for spare activation and , and use (i.i.d.) to denote the independent and identically distributed inherent lifetimes of units; for the same sample path, a fixed set of realizations is used.
Define two spare part activation strategies:
Early Activation Strategy (denoted as ): The system starts at time , all units operate simultaneously, and aging timing begins.
The failure time of unit
is as follows:
The moment of system failure, that is, the time of the
th failed unit, is as follows:
Late Activation Strategy (denoted as ): Initially, only units operate, and the remaining cold spares are not timed; when the number of original component failures is within the first times, only replacement is performed without activating spares; when the cumulative failures reach the critical value , the first cold spare is activated immediately, and timing starts from this moment; thereafter, each subsequent unit failure triggers the immediate activation of the next cold spare, which replaces the failed unit one by one.
The operating time of working units remains unchanged, and the failure time of the standby unit is equal to the activation time plus its own lifetime.
2.2. A Point-by-Point Comparison of Failure Times for the Same Sample Path
For any fixed sample path , for any unit , we obtain the following.
Early activation: Starting from the moment
, the failure time is
Late activation: If the unit had been part of the working group initially, the failure time is
If we consider a cold standby unit activated later, and activation time is
, then
Therefore, the following holds unit by unit and path by path:
2.3. Pointwise Dominance of System Lifetime
The system failure logic is identical for both strategies: the system fails when the number of failed units reaches , i.e., the system lifetime is equal to the time of the th failure.
All failure times are then sorted:
By the order statistic inequality preservation property, the following is obtained:
Take
; thus
This holds for every sample path; thus pointwise dominance is achieved:
2.4. Conclusion on Optimal Activation Strategy
Under the conditions of ideal cold standby (failure rate is 0 during standby), instantaneous and reliable switching, and i.i.d. unit lifetimes, for any sample path, the following hold.
(1) Under the late activation strategy, the failure time of each unit is pointwise no less than that under the early activation strategy.
(2) The system lifetime is the th-order statistic.
(3) Therefore, the system lifetime satisfies pointwise dominance:
Further, first-order stochastic dominance holds, and reliability always satisfies the following:
The mean time to failure (MTTF) satisfies the following:
For cold standby voting systems under any lifetime distribution, under ideal conditions without considering switching delay and activation failure, the later the activation timing, the better the reliability and mean lifetime performance of the cold standby voting system. Thus, the late activation strategy is superior to the early activation strategy, and this conclusion is independent of the unit lifetime distribution form.
2.5. Hazard Behavior Plot Based on Failure Rate Curve
To intuitively demonstrate the suppression effect of the late activation strategy on the system failure rate,
Figure 1 presents the curves of the system instantaneous failure rate h(t) over time under different activation strategies with the same Weibull distribution parameters (
).
It can be seen that under the early activation strategy, all spares age synchronously with the main components, leading to a rapid rise in the system failure rate at an early stage; under the late activation strategy, spares are activated with delay, only original components operate initially with a low failure rate, and as original components fail gradually and spares are put into use sequentially, the failure rate rises much more slowly than that under the early activation strategy. This plot verifies the theoretical conclusion in
Section 2.4 from the perspective of the failure rate: the late activation strategy yields a lower system failure rate over the entire time domain, thereby achieving higher reliability and a longer mean lifetime.
4. Example Simulation
In railway signal systems, the stable operation of signal control equipment is critical to train safety and efficiency. To adapt to complex operating conditions and potential failures, such systems typically adopt a multi-device collaborative working mode. This involves series, parallel, and voting configurations, along with backup equipment to facilitate rapid fault replacement. This design enables continuous and stable system operation.
Assume that the main system consists of three signal control devices, denoted as , , and with one standby signal control device available. The failure time of each signal control device follows a Weibull distribution.
Based on varying reliability requirements, the system is designed with three core operating mechanisms, described as follows. (1) G1: 3/3:1(G)—The system operates normally only if all three main control devices are functional, following a series-type logic with high reliability demands. (2) G2: 1/3:1(G)—The system remains operational if at least one of the three main control devices is functioning, following a parallel-type logic that emphasizes continuity. (3) G3: 2/3:1(G)—The system operates normally if at least two of the three main control devices are functioning, embodying a voting-type logic that balances reliability and fault tolerance. The standby equipment is deployed to replace any failed device in the main system immediately.
Given the parameters , , and , in this simulation, the preliminary experiment yielded a sample standard deviation of and an absolute error of . Calculated according to Equation (27) and rounded off, the result is , which satisfies the requirement of , so it is applicable. The simulation step size is taken as . The operational performance of the system under these different working mechanisms is accurately evaluated using both the non-homogeneous Markov model and the Monte Carlo simulation model. These models are employed to predict and analyze the system’s average life, remaining life after 500 h of operation, and overall reliability. The prediction results of both models are compared and presented through data comparison and graphical visualization.
The comparison of prediction values and prediction rates for different remaining life prediction models, based on the non-homogeneous Markov and Monte Carlo simulation models, is summarized in
Table 1 and
Table 2.
Table 1 demonstrates that the predictions of both models are highly consistent, with deviation rates remaining low (≤2.54%). This suggests that the non-homogeneous Markov model offers prediction accuracy comparable to the traditional Monte Carlo simulation model. Notably, the deviation rate under the G3 (2/3:1(G)) mechanism is the lowest (0.8052%), highlighting that the non-homogeneous Markov model delivers the highest prediction accuracy under this operational configuration.
Table 2 demonstrates that the running time of the non-homogeneous Markov model is extremely short, consistently staying below 0.13 s. In contrast, the Monte Carlo model requires significantly more time, ranging from 60 to 105 s, highlighting a substantial disparity between the two. The time cost increase rate for the Monte Carlo model relative to the non-homogeneous Markov model exceeds 99.8%, meaning that the Monte Carlo model demands hundreds of times more time for the same prediction task.
Further analysis reveals that the Monte Carlo model becomes increasingly time-consuming as the system’s working mechanism places greater demands on device collaboration. In contrast, the running time of the non-homogeneous Markov model remains relatively stable or even slightly decreases. This suggests that the non-homogeneous Markov model exhibits superior adaptability under more complex operational configurations and offers a clear efficiency advantage. Meanwhile, the computational complexity of the Monte Carlo model grows substantially with increasing system complexity, leading to a significant rise in running time.
The reliability and remaining life prediction results obtained from both the non-homogeneous Markov model and the Monte Carlo simulation under the three mechanisms, G1, G2, and G3, are presented in
Figure 3,
Figure 4 and
Figure 5, respectively.
For the G1 (3/3:1(G)) mechanism, reliability declines the most rapidly, with a sharp decrease in the early stage (0–1000 h) and stabilization at a low level thereafter. This behavior is attributed to the series configuration, where all units must remain operational; the failure of any single unit leads to a significant reduction in reliability. The remaining life after 500 h is relatively short, ranging from approximately 300 to 400 h. It also exhibits a rapid decay rate. These results indicate a substantially increased failure risk after prolonged operation under the series configuration.
For the G2 (1/3:1(G)) mechanism, reliability decreases at the slowest rate, remaining at a relatively high level throughout the 5000 h observation period. This behavior reflects the strong fault tolerance of the parallel configuration, where system operation can be sustained even if multiple main devices fail. The remaining life after 500 h is the longest, exceeding 1400 h, and exhibits a gradual decay over time. These results indicate that the parallel mechanism maintains stable performance even after prolonged operation.
For the G3 (2/3:1(G)) mechanism, the reliability decline rate lies between those of the G1 and G2 mechanisms. It remains relatively stable in the early stage and gradually decreases in the later stage. This behavior is consistent with the design principle of the 2/3 voting mechanism, which avoids the high failure risk of the G1 configuration while reducing excessive dependence on a single device, thereby achieving a balance between reliability and fault tolerance. The remaining life after 500 h is approximately 700–800 h and exhibits a moderate decay rate, reflecting the typical lifetime characteristics of voting systems.
Under all three operating mechanisms, the reliability trends predicted by the non-homogeneous Markov model (state transition method) and the Monte Carlo method show a high degree of agreement. The predicted curves of remaining life after 500 h are also closely aligned. This consistency confirms the accuracy of the non-homogeneous Markov model. As shown in
Figure 3,
Figure 4 and
Figure 5, the results of both methods remain highly consistent during key intervals where reliability changes either rapidly or gradually, with no significant deviations observed. These results demonstrate that the non-homogeneous Markov model accurately captures reliability variations and maintains strong predictive stability.
Overall, when component and spare part lifetimes follow the Weibull distribution, the parallel mechanism provides the highest fault tolerance and longest lifetime, whereas the series mechanism imposes the strictest reliability requirements and yields the shortest lifetime. The voting mechanism offers a balanced compromise between these two, allowing for the selection of an appropriate configuration based on practical engineering needs. The non-homogeneous Markov model effectively supports the high-accuracy prediction of residual life and system reliability in cold standby voting systems. Compared with Monte Carlo simulation, it significantly reduces computational cost, with its efficiency advantage becoming more pronounced under complex configurations (e.g., G3), making it suitable for fast engineering prediction.
Considering the reliability time scales of typical engineering systems (such as electromechanical and aerospace electronic systems), when the abscissa time (hours) spans from hundreds to thousands of hours, the commonly used step size ranges from 0.1 to 1 h. Therefore, a step size of 0.1 h is adopted in this simulation.
However, in practical scenarios, the prediction accuracy of the non-homogeneous Markov model is related to the prediction step size. Sensitivity comparison charts for different step sizes are shown in
Figure 6. In future research, the stability and practicability of the model can be further improved by optimizing the step size adaptive algorithm.