Next Article in Journal
Industrial-Scale Valorization of Low-Grade Fruits Through Green Polyphenol Extraction: Integrating Process Simulation, Techno-Economic Analysis, and Environmental Assessment
Previous Article in Journal
DIET-Intensified Dark Fermentation: Magnetite and Nickel–Iron-Doped Activated Carbon for Biohydrogen Production from Food Waste
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

MFD-Mamba: A Mamba-Based Framework for Multiple-Fault Process Monitoring in Industrial Systems

1
School of Information and Control Engineering, China University of Mining and Technology, Xuzhou 221116, China
2
Guixi Dasanyuan Industry (Group) Co., Ltd., Guixi 335419, China
3
China Petroleum Pipeline Engineering Corporation, Shenyang 110000, China
*
Author to whom correspondence should be addressed.
Processes 2026, 14(18), 2973; https://doi.org/10.3390/pr14182973 (registering DOI)
Submission received: 1 September 2026 / Revised: 13 September 2026 / Accepted: 16 September 2026 / Published: 18 September 2026
(This article belongs to the Section Chemical Processes and Systems)

Abstract

Multiple faults in industrial processes may occur simultaneously or successively, leading to abnormal responses distributed across coupled process regions and evolving over time. This paper proposes a Mamba-based multiple-fault detection framework, termed MFD-Mamba, for dynamic industrial processes. A hybrid multiblock decomposition first combines lag-aware data relationships with prior process connections to construct soft variable memberships and a block-relation prior. The resulting blockwise sequences are then modeled by a shared modified Mamba encoder, while dynamic block relation modeling is used to describe interactions among process regions. Blockwise process predictions are further used to construct hybrid residuals. For fault detection, block residuals are converted into probabilistic evidence and fused within a Bayesian framework, where multiple residual-based monitoring statistics, persistence information, and cross-evidence terms are jointly incorporated to construct the final monitoring index. The proposed method is evaluated on the Tennessee Eastman process and the activated sludge process under overlapping and sequential multiple-fault conditions. On the TE process, MFD-Mamba achieves FDR/FAR values of 99.88%/0.13% and 99.81%/0.20% for the overlapping and sequential fault scenarios, respectively. On the ASP, the corresponding FDR/FAR values are 99.79%/0.78% and 99.58%/1.56%. These results demonstrate that MFD-Mamba provides accurate and reliable multiple-fault detection across different dynamic industrial processes.

1. Introduction

Modern industrial processes consist of multiple interacting units in which variables are coupled through material flows, energy transfer, recycle streams, and control loops [1,2,3]. A fault occurring in one part of the process may propagate to other units, leading to distributed abnormal responses across several variables. In practical operation, faults may also occur simultaneously or successively rather than in isolation, giving rise to overlapping and sequential multiple-fault conditions. Under these conditions, the observed process response is influenced by the combined effects of different faults and their temporal evolution [4,5]. The resulting abnormal information may be distributed across different process regions, and may vary over time as the fault condition changes. Consequently, multiple-fault process monitoring requires the detecting model to account for both the structural coupling among process variables and the dynamic evolution of fault-related information [6,7].
Data-driven fault detection methods construct monitoring models directly from process measurements without requiring a complete first-principles description of the process [8,9]. Multivariate statistical methods such as principal component analysis, partial least squares, and canonical correlation analysis characterize normal operating behavior through latent variable representations and corresponding monitoring statistics. With the development of deep learning, models based on autoencoders, recurrent neural networks, convolutional networks, and transformers have been introduced to capture nonlinear relationships and temporal dependencies in industrial process data [9,10]. More recently, selective state-space models have provided another approach to representing long process sequences through input-dependent state updates. These methods have extended the ability to model complex process dynamics; however, many existing approaches still construct a unified representation over the complete variable space, which may not fully preserve the local characteristics of different process regions under multiple-fault conditions [11,12].
Multiple-fault monitoring has received increasing attention in industrial process applications. Early studies addressed this problem through process decomposition and latent variable modeling [13]. Lee et al. [14] combined system decomposition with dynamic PLS for multiple-fault diagnosis in the Tennessee Eastman process. Zhao and Gao [15] considered the combined effects of multiple faults through fault subspace reconstruction, while Ma et al. [16] developed a hierarchical framework that integrated subprocess-level monitoring results at the plant-wide level for KPI-related multiple faults. More recently, Ma et al. [17] considered compound faults occurring simultaneously or successively in coupled manufacturing processes, and Zhou et al. [18] investigated simultaneous faults using a multi-label learning framework. Zhu et al. [19] further studied simultaneous compound faults occurring at different locations of an industrial process and considered their coupled temporal characteristics. In addition, Yang et al. [20] investigated composite faults that evolve from individual fault conditions during system operation. These studies indicate that multiple faults may exhibit both spatial coupling and temporal evolution. A recent review also identifies multiple-fault diagnosis as an ongoing problem in industrial and chemical processes [21].
To preserve local process information, multiblock monitoring methods divide the complete variable space into several process regions and construct local monitoring models for individual blocks. Recent studies have combined multiblock structures with temporal modeling and block-level information fusion for complex process monitoring [22,23,24]. Such decomposition can retain local abnormal responses that may be weakened in a global representation. However, conventional multiblock methods commonly rely on predefined or disjoint variable partitions. In strongly coupled industrial processes, variables associated with recycle streams, shared operating units, or neighboring subsystems may participate in several process regions. A strict assignment of each variable to only one block may discard part of the relationships among coupled regions; moreover, partitions based only on process knowledge may overlook dependencies revealed by process data, whereas data-driven decomposition may not preserve known process connections. These limitations motivate the development of a process representation that integrates data-derived relationships with prior process information while allowing variables to be associated with multiple related blocks.
Beyond process decomposition, the temporal evolution of multiple faults also needs to be considered. Fault responses may develop after fault occurrence, propagate among coupled process regions, or change when a new fault is introduced during an existing abnormal condition. Transformer-based models have been combined with temporal feature extraction for chemical process fault detection [25]. More recently, Mamba-based state-space models have been introduced into fault diagnosis to represent long-range dependencies and interactions in multivariate sequences [26]. For multiblock process monitoring, however, temporal modeling should both characterize the dynamics within individual blocks and consider the interactions among related process regions. Such interactions are particularly important under multiple-fault conditions, where different local regions may exhibit distinct temporal responses as the operating state evolves.

Comparison with Recent Mamba-Based Monitoring Methods

Recent studies have introduced Mamba-based state-space models into multivariate time series anomaly detection and industrial fault diagnosis. For example, MambaAD [27] employs bidirectional multi-view Mamba modules to capture temporal dependencies and inter-variable correlations, while MemMambaAD [28] incorporates sequence decomposition and memory mechanisms to improve its representation of normal patterns. The more recent Dual-Path Mamba [29] combines adaptive signal decomposition with variable-wise state-space modeling for chemical process fault diagnosis.
Despite these advances, existing Mamba-based methods mainly focus on temporal representation or inter-variable dependency modeling within a global sequence framework. In contrast, MFD-Mamba first constructs overlapping process blocks by combining data-driven relations with process prior information, then performs block-wise Mamba modeling and integrates block-level fault evidence through Bayesian fusion. Therefore, the improvement of MFD-Mamba arises not only from the Mamba backbone but also from the proposed process decomposition and evidence fusion strategy, as supported by the ablation results.
The integration of fault information from different process regions is another issue in multiple-fault monitoring. Individual blocks may exhibit different abnormal responses, and their contributions to the final monitoring decision may differ. Recent studies have also considered information fusion strategies for simultaneous multi-fault diagnosis [30]. However, simple averaging may weaken localized fault information, whereas maximum-based fusion may depend excessively on a single block. In addition, prediction residuals and deviations from the normal operating state describe different aspects of abnormal process behavior. Therefore, the final monitoring decision should integrate block-level fault evidence while retaining complementary state deviation information and the temporal continuity of abnormal responses.
To address the above issues, this paper proposes a Mamba-based multiple-fault detection framework for dynamic industrial processes, which we term MFD-Mamba. The framework first constructs a hybrid multiblock representation by combining data-derived variable relationships with prior process information. Soft memberships are introduced to retain associations with multiple coupled blocks rather than being restricted to a strictly disjoint partition. Based on the resulting block structure, a blockwise selective state-space model is developed to characterize local temporal dynamics. The block states are further combined with block relation information to describe interactions among related process regions. Block-level prediction residuals are then transformed into probabilistic fault evidence and integrated through Bayesian fusion. State deviation information and temporal persistence are incorporated into the final monitoring decision. MFD-Mamba combines process decomposition, temporal modeling, and block-level evidence fusion for overlapping and sequential multiple-fault detection.
The main contributions of this work are summarized as follows:
  • A hybrid multiblock decomposition is introduced by combining lag-aware data relationships with prior process connections. Soft memberships are used to describe variables associated with multiple process regions, and the resulting memberships are further used to construct a block relation prior.
  • A blockwise selective state-space modeling scheme is constructed for temporal process representation. Instead of modeling the complete variable space directly, the learned blockwise sequences are processed by a shared modified Mamba encoder to describe the temporal behavior of individual process blocks.
  • Dynamic block relation modeling is introduced to combine the structural block prior with the current block representations. The updated block states are used for blockwise process prediction, from which hybrid residuals are constructed as block-level monitoring information.
  • A Bayesian fusion and multi-channel detection scheme integrates block-level residuals with complementary monitoring information to construct the final monitoring index, with thresholds determined from normal operating data.

2. Related Work

Data-driven fault detection methods mainly construct monitoring models from process measurements. Multivariate statistical methods characterize process variation through latent variables and monitoring statistics. Deep learning methods further extend this framework by learning nonlinear representations from multivariate process data. Autoencoders, recurrent networks, temporal convolutional models, and attention-based architectures have been used to describe normal operating patterns and process dynamics. For dynamic process monitoring, transformer-based models provide long-range dependency modeling through self-attention, while state-space models offer another approach to sequence representation. Mamba employs a selective state-space mechanism with input-dependent state updates, and has been introduced into both multivariate time series modeling and anomaly detection. However, existing temporal monitoring methods are often constructed over the complete variable space. For coupled industrial processes, such a global representation may not explicitly preserve the different temporal responses associated with individual process regions.
Multiple faults may occur simultaneously or successively, resulting in fault information being distributed across several process regions. Multiblock monitoring methods address this problem by dividing process variables into local subsystems and constructing block-level monitoring models. While this representation preserves local process information, a strictly disjoint partition may be restrictive for strongly coupled processes in which individual variables may be associated with multiple process units or recycle streams. A global monitoring decision also requires the fusion of block-level responses. Common approaches include averaging, weighted summation, voting, and maximum statistics. These methods use fixed fusion rules, whereas the contribution of individual blocks may vary under different fault conditions. Probabilistic fusion provides an alternative by combining local evidence according to its support for normal and faulty states. These studies provide the basis for structured process monitoring and temporal fault representation. However, process decomposition, temporal state modeling, and block-level evidence fusion are often considered separately. As described in the next section, the proposed MFD-Mamba framework integrates these components for multiple-fault monitoring.

3. MFD-Mamba for Multiple-Fault Detection

This section describes the proposed MFD-Mamba framework for multiple-fault detection. MFD-Mamba first organizes the process variables into several coupled blocks through hybrid multiblock decomposition. The blockwise sequences are then modeled by the modified Mamba architecture in order to capture their temporal evolution and cross-block interactions. Based on the resulting block-level fault evidence, Bayesian fusion and a multi-channel detection strategy are employed to form the final monitoring decision.

3.1. Hybrid Multiblock Process Decomposition

Let X = [ x 1 , x 2 , , x N ] T R N × m denote the normal operating data, where N and m are the numbers of samples and process variables, respectively. Industrial process variables are usually associated with different operating units while remaining coupled through material flows, recycle streams, and dynamic interactions. Modeling all variables in a single block may obscure local process relationships, whereas a strictly disjoint partition may separate variables shared by several coupled process regions. Therefore, a hybrid multiblock decomposition is constructed by combining data-derived variable dependence with prior process information.
The data-derived relation is first determined from normal operating data. Since process interactions may involve finite transport and response delays, the relation between the ith and jth variables is evaluated over a set of candidate lags
R i j d = max τ L ρ x i ( t ) , x j ( t τ ) ,
where L denotes the considered lag set and ρ ( · , · ) is the Pearson correlation coefficient. The absolute value is used because both positive and negative process dependencies indicate coupling between variables. Applying this calculation to all variable pairs gives R d R m × m . When several normal operating sequences are available, the lagged dependence is evaluated from the available normal sequences before the common relation matrix is constructed.
Data-derived dependence alone does not explicitly retain known process connections. Therefore, a process relation matrix R p R m × m is introduced according to the physical connections and process unit associations among the monitored variables. Its elements indicate whether two variables are connected through the predefined process structure. The data and process relations are combined as
R = ( 1 λ p ) R d + λ p R p ,
where λ p [ 0 , 1 ] controls the contribution of prior process information. The resulting matrix R retains the relationships observed from normal process data while introducing structural information that may not be fully reflected by statistical dependence alone.
Based on R , the process variables are organized into B blocks using a soft symmetric non-negative matrix factorization. Let U R + m × B denote the membership matrix, where U i b represents the membership of variable i in block b. To avoid an excessively dense block structure and discourage memberships inconsistent with the predefined process connections, the membership matrix is obtained by solving
min U 0 1 4 R U U T F 2 + λ s U 1 + λ n 2 Tr U T Q U ,
where λ s controls membership sparsity, λ n controls the penalty on nonprocess relationships, and Q is obtained from the complement of the process relation matrix, with its diagonal elements set to zero. The first term preserves the hybrid variable relation, the second limits unnecessary memberships, and the third penalizes simultaneous assignments of variables that are not supported by the predefined process structure.
The non-negative membership matrix is updated iteratively using the multiplicative rule
U U R U + ϵ U U T U + λ n Q U + λ s + ϵ ,
where division and multiplication are element-wise and ϵ is a small positive constant used for numerical stability. After convergence, the memberships of each variable are normalized by
U i b U i b max 1 k B U i k + ϵ .
Thus, the membership values describe the relative association of each variable with the learned process blocks.
Unlike a hard partition, the resulting representation does not require each variable to belong exclusively to one block. A variable can retain nonzero memberships in several related blocks when it is involved in multiple process regions. This preserves local block structure while retaining cross-unit relationships associated with shared variables. For interpretation and presentation of the learned blocks, the soft memberships can be converted into representative block-variable lists using a membership threshold, together with constraints on the number of blocks associated with each variable and the minimum number of variables in each block. The soft membership matrix U is retained in the subsequent model rather than the displayed hard block lists.
The learned membership matrix is also used to transfer the variable-level relation to the block level. The relation prior among the B process blocks is constructed as
A B = N U T R U ,
where N ( · ) denotes the symmetric normalization of the block-relation matrix. This prior describes the connections among the learned blocks and is subsequently incorporated into the blockwise selective state-space model. Consequently, the decomposition provides both the soft block memberships and the corresponding block relation structure required for temporal process modeling.

3.2. Blockwise Selective State-Space Modeling

The multiblock decomposition provides the structural organization of process variables, while the temporal behavior within each block and the interactions among different blocks remain to be modeled. To this end, a blockwise selective state-space model is developed. Instead of directly applying Mamba to the complete variable space, the process sequence is first organized according to the learned block memberships. The resulting blockwise sequences are then processed by a shared Mamba encoder, followed by block relation modeling and next-step prediction.
For an observation window of length L, the input sequence at time t is defined as
X t = [ x t L , , x t 1 ] R L × m .
According to the membership matrix U , X t is organized into B blockwise sequences { Z t ( 1 ) , , Z t ( B ) } . Since a process variable can contribute to more than one block, the resulting representations retain both local information and the relationships associated with shared variables.
For the bth block, layer normalization and linear projection are first applied to the input, followed by causal depthwise convolution
v k ( b ) = DWConv W x LN ( z k ( b ) ) ,
where DWConv ( · ) denotes causal depthwise convolution. The selective state-space parameters are generated from the current input as
Δ k ( b ) , B k ( b ) , C k ( b ) = G v k ( b ) ,
where G ( · ) denotes the input-dependent projection. The discrete state update is then given by
h k ( b ) = A ¯ k ( b ) h k 1 ( b ) + B ¯ k ( b ) v k ( b ) ,
s k ( b ) = C k ( b ) h k ( b ) ,
where h k ( b ) denotes the latent state, while A ¯ k ( b ) and B ¯ k ( b ) are the discretized state-space parameters. Since Δ k ( b ) , B k ( b ) , and C k ( b ) depend on the current input, the state update can adapt to changes in the process sequence.
The state-space output is further modulated by a gated branch and combined with the input through a residual connection:
o k ( b ) = z k ( b ) + W o s k ( b ) σ W g LN ( z k ( b ) )
where σ ( · ) denotes the gating activation and ⊙ denotes element-wise multiplication. Multiple Mamba blocks are stacked to extract the temporal representation of each process block.
The same temporal encoder is shared among all process blocks:
S t ( b ) = M θ Z t ( b ) , b = 1 , , B
where M θ ( · ) denotes the shared Mamba encoder. This allows different process blocks to be represented in a common temporal feature space while retaining their block-specific inputs.
The temporal representations are subsequently coupled according to the relationships among process blocks.
To account for changes in block interactions under different process states, the fixed prior is combined with the current block representations:
A t = F r S t , A B ; λ r
where F r ( · ) denotes the dynamic relation mapping and λ r controls the contribution of the structural prior. The block states are then updated as
S ˜ t = A t V t + S t ,
where V t is the value projection of the current block representations.
The updated representation of each block is used to predict the next process state:
x ^ t ( b ) = f b s ˜ t ( b ) , x ^ t ( b ) R m
where f b ( · ) denotes the prediction head of the bth block. Each block produces a prediction over the complete process-variable space, which provides a common basis for evaluating blockwise prediction errors.
The model is trained using normal operating sequences. The objective function consists of the prediction loss and a cross-block consistency term:
L = L pred + λ c L con
where λ c controls the contribution of the consistency term.
For online monitoring, both the overall prediction error and the larger variable-wise deviations are retained. For the bth block, the hybrid residual is defined as
r b , t = ( 1 γ ) r b , t mean + γ r b , t tail ,
where r b , t mean denotes the mean prediction residual, r b , t tail represents the residual associated with the largest variable-wise deviations, and γ determines their relative contribution. The residuals of all blocks are finally collected as
r t = [ r 1 , t , r 2 , t , , r B , t ] T .
The vector r t is subsequently used as the block-level fault evidence for Bayesian fusion and final fault detection.

3.3. Bayesian Fusion and Fault Detection

The blockwise selective state-space model produces the residual vector r t for each monitoring sample. Since a multiple fault may affect several process regions with different response magnitudes, directly averaging the block residuals may suppress fault information from locally affected blocks. Therefore, the block residuals are first converted into probabilistic fault evidence and subsequently fused at the process level.
For each block, a control limit H b is determined using only normal operating residuals:
H b = Q α b { r b , t : t T N }
where T N denotes the normal reference set and Q α b ( · ) denotes the control-limit estimator at confidence level α b . Based on the block residual and its normal control limit, the evidence under the faulty and normal hypotheses is denoted by p ( r b , t | F ) and p ( r b , t | N ) , respectively. Given the prior fault probability π F , the posterior fault probability of block b is obtained from Bayes’ rule:
P b , t F = p ( r b , t | F ) π F p ( r b , t | F ) π F + p ( r b , t | N ) ( 1 π F )
where P b , t F represents the instantaneous Bayesian fault evidence of the bth process block. The corresponding normal-state probability is P b , t N = 1 P b , t F .
The block-level Bayesian evidence is then combined to obtain a global fault index
g t = F B P 1 , t F , P 2 , t F , , P B , t F ,
where F B ( · ) denotes the weighted Bayesian fusion operator. In contrast to equal averaging, the fusion result reflects the instantaneous fault support provided by the individual process blocks. A global control limit H g is calibrated from normal operating data, and g t / H g provides the normalized Bayesian evidence.
Prediction-based evidence mainly describes deviations from the temporal behavior learned by MFD-Mamba. To retain information on the current process state, two additional normal-state deviation statistics are introduced. Let μ and Σ ^ denote the mean and shrinkage covariance matrix estimated from normal operating data. The Mahalanobis statistic is defined as
M t = ( x t μ ) T Σ ^ 1 ( x t μ ) .
The statistic quantifies deviations from the correlated normal operating region, while the squared prediction error is calculated as
Q t = x t P P T x t 2 2 ,
where P contains the retained principal directions. The SPE statistic describes deviations that cannot be represented by the normal principal subspace.
To retain persistent changes in the SPE channel, a causal exponentially weighted moving statistic is further computed:
E t = λ e Q t + ( 1 λ e ) E t 1
where λ e ( 0 , 1 ] is the smoothing coefficient and E 0 is initialized from normal operating data. The corresponding limits of the Mahalanobis, SPE, and EWMA channels are also determined from normal reference samples.
A persistence operator is used to reduce isolated threshold crossings. For a normalized evidence sequence z t , define
P K , N ( z t ) = kth K { z t N + 1 , , z t } ,
where kth K ( · ) denotes the Kth largest value within the most recent N samples. Hence, P K , N ( z t ) > 1 indicates that at least K of the previous N samples exceed the corresponding normalized control limit.
The Bayesian and current-state channels are combined through a final normalized monitoring index. In the present detector, the persistent Bayesian evidence and Mahalanobis evidence are defined as
J t g = P 2 , 3 g t H g , J t M = P 2 , 3 M t H M L ,
where H M L denotes the lower Mahalanobis decision level. Direct high-deviation channels are given by
J t M H = M t H M H , J t Q = Q t H Q , J t E = E t H E ,
where H M H , H Q , and H E are the high Mahalanobis, SPE, and EWMA limits, respectively.
A cross-evidence channel is additionally introduced to retain samples for which both the Bayesian residual evidence and the SPE evidence are elevated:
C t = min g t H g E , Q t H Q E
where H g E and H Q E are normal-data-calibrated evidence levels. Its persistent form is
J t C = P 2 , 2 ( C t ) .
The final monitoring statistic is then defined as
J t = max J t g , J t M , J t M H , J t Q , J t E , J t C .
A fault alarm is generated according to
A t = 1 , J t > 1 , 0 , J t 1 .
All control limits and evidence thresholds involved in the final decision are calibrated using normal operating data. No fault samples are used in the threshold determination. The resulting detector combines blockwise temporal prediction evidence with current-state deviations, providing a common monitoring decision for both simultaneous and sequential multiple-fault conditions. The overall framework is illustrated in Figure 1.

3.4. Hyperparameter Selection Strategy

The hyperparameters are selected in a staged manner according to their functional roles in the proposed framework. Parameters related to the multiblock construction are first adjusted to obtain stable and physically meaningful block structures. After the block structure is fixed, the parameters associated with temporal modeling and relation learning are determined based on the prediction behavior and the stability of the monitoring statistics. Finally, the smoothing-related parameter is selected to balance fault sensitivity and the suppression of isolated fluctuations. This staged strategy reduces the coupling among parameters from different modules and avoids simultaneously tuning all hyperparameters.

4. Two Industrial Cases

4.1. Experimental Settings

In this study, all computational efficiency experiments are conducted on the same Windows 11 workstation equipped with a 13th Gen Intel Core i7-13620H CPU, 15.73 GB of system memory, and an NVIDIA GeForce RTX 4060 Laptop GPU with 8 GB of memory. The experiments use Python 3.12.3 and PyTorch 2.13.0. The hyperparameters are selected in a staged manner according to their roles in the proposed framework. Specifically, λ p , λ s , and λ n , which respectively control the contribution of process prior information, membership sparsity, and the penalty on non-mechanistic relations, are first determined to obtain stable and physically meaningful overlapping process blocks. After the block structure is fixed, λ r , λ c , and γ are adjusted by considering the prediction performance and the stability of the resulting monitoring statistics. The EWMA smoothing coefficient λ e is selected to balance detection sensitivity and the suppression of isolated fluctuations, and λ e = 0.40 is used in the experiments. The corresponding detection thresholds are subsequently calibrated using normal operating data. Parameter selection is conducted using the training and validation data, while the test fault scenarios are reserved for final performance evaluation. This staged strategy avoids simultaneously tuning a large number of parameters and reduces the interactions among parameters belonging to different modules. For reproducibility, the final hyperparameter settings are summarized in Table 1.

4.2. Tennessee Eastman Process

The Tennessee Eastman (TE) process is adopted as the benchmark system for evaluating the proposed monitoring framework [31]. As shown in Figure 2, the TE process consists of five major operating units, including a reactor, condenser, compressor, separator, and stripper. The benchmark provides a set of representative fault conditions with different dynamic characteristics, making it suitable for assessing monitoring performance under complex process disturbances. For the TE experiments, 33 variables are used as the input of the monitoring model, comprising 22 process measurements and 11 manipulated variables. Model training and control-limit estimation are both performed using normal-operation samples. The fault sequences are excluded from model construction and are introduced only during the testing stage. In this way, the detection results reflect the response of the trained model to previously unseen abnormal operating conditions.
To further evaluate the monitoring performance under multiple-fault conditions, two scenarios are considered, namely, overlapping and sequential faults. For the overlapping-fault scenario, the first 160 samples represent normal operation. Fault 4 and Fault 6 are then introduced simultaneously from the 161st sample onward and remain active throughout the subsequent fault period. Concurrent-fault detection performance is evaluated under conditions where multiple fault effects coexist. For the sequential-fault scenario, the first 160 samples correspond to normal operation, followed by 800 samples affected by Fault 1. The process then returns to the normal condition for 200 samples to form a recovery interval, after which Fault 14 is introduced for the remaining 800 samples. Sequential-fault monitoring performance is evaluated in terms of the ability to track the transition from normal operation to a fault condition, recover after fault removal, and respond promptly to a subsequent fault of a different type.
For quantitative evaluation, the fault detection rate (FDR), false alarm rate (FAR), and detection delay are adopted as the performance metrics. The proposed MFD-Mamba is compared with MambaAD [27]; Mola [22], a multi-view Mamba-based method for multivariate time-series anomaly detection; and the incipient fault detection (IFD) method [32]. All methods are evaluated using the same test sequences and fault occurrence settings. The method-specific hyperparameters are configured according to their respective implementations, while the data partition and evaluation criteria are kept consistent across all methods to ensure a fair comparison. For the TE experiments, the input sequence length is set to 10 and the batch size is 32. The proposed model is trained for 50 epochs using the Adam optimizer with an initial learning rate of 3 × 10 4 . The control limit is determined from the normal training data with a confidence level of 0.99. A moving-average window of five samples is further employed to suppress isolated fluctuations in the monitoring statistic. All experiments are implemented under the same data partition and fault setting protocol.
Figure 3, Figure 4, Figure 5 and Figure 6 compare the detection performance of MFD-Mamba, MambaAD, IFD, and MOLA under the overlapping-fault scenario. As shown in Figure 3, the monitoring index of MFD-Mamba remains below the control limit during normal operation and rapidly enters the abnormal region after the overlapping faults occur, where it remains stable throughout the fault period. The detection result of MambaAD is shown in Figure 4. Its anomaly score increases gradually after fault occurrence and crosses the control limit after a transition period, resulting in a longer detection delay. As shown in Figure 5, IFD also responds to the overlapping faults, but several threshold exceedances occur during normal operation, leading to a higher FAR. MOLA, shown in Figure 6, exhibits a more fluctuating and intermittent response after fault occurrence, particularly during the early fault stage, indicating less stable fault detection.
Figure 7, Figure 8, Figure 9 and Figure 10 compare the detection performance of MFD-Mamba, MambaAD, IFD, and MOLA under the sequential-fault scenario. MFD-Mamba clearly tracks the successive transitions from normal operation to the first fault, recovery, and the second fault, with distinct changes in the monitoring index. In Figure 8, MambaAD also identifies the two fault stages and returns to the normal region during the recovery interval, although its response to the fault transitions is relatively slower. In Figure 9, IFD responds to both faults, but its monitoring statistic exhibits larger fluctuations and several pronounced peaks during the first fault stage.
In contrast, MOLA shows a delayed increase after the first fault occurs in Figure 10. Its monitoring index remains at a high level during the recovery interval, making the recovery and subsequent fault transition less distinguishable. MFD-Mamba provides clearer and more stable tracking of the successive operating-state transitions under the sequential-fault condition.
To examine the effect of training randomness, each method is independently trained and evaluated over six runs using six different random seeds. The same set of random seeds is used for all compared methods to enable paired statistical testing. The FDR, FAR, and detection delay are reported as the mean ± standard deviation, and a paired t-test is conducted on the seed-matched FDR results.
The quantitative results are summarized in Table 2. MFD-Mamba provides the best overall performance under both overlapping-fault and sequential-fault conditions, showing consistently high detection accuracy, low false alarm rates, and short detection delays. Compared with single-block Mamba, the proposed method achieves more reliable monitoring, particularly in terms of false alarm suppression and fault-response speed. This comparison indicates that the performance improvement does not arise from the Mamba backbone alone but also benefits from the proposed multiblock decomposition and fusion strategy. As shown in Table 3, the FDR improvements over MambaAD, IFD, and MOLA are statistically significant under the overlapping-fault condition. Under the sequential-fault condition, the difference between MFD-Mamba and MambaAD is not statistically significant, whereas significant differences are observed relative to IFD and MOLA. Overall, the results demonstrate that MFD-Mamba improves the selectivity and reliability of fault evidence while maintaining high sensitivity to multiple-fault transitions.

4.3. Activated Sludge Process

The activated sludge process (ASP) is used to evaluate the monitoring performance of the proposed method. The ASP plant consists of five biological reactors connected in series, followed by a secondary settler [33]. The first two reactors are operated under anoxic conditions for denitrification, whereas the remaining three reactors are aerated for carbon removal and nitrification. The secondary settler separates the treated effluent from activated sludge and includes return-sludge and waste-sludge streams. An internal recycle stream connects the aerobic and anoxic sections to support nitrogen removal. The biological reactions are described by ASM1, while the secondary settler is represented by a ten-layer settling model. A total of 25 process variables are selected as model inputs, and the effluent total nitrogen is retained as the quality variable. Normal operating data under dry, rainy, and storm conditions are used for model construction. The sampling interval is 15 min, and each test sequence covers Day 7 to Day 14, giving 672 samples.
For all methods, the normal samples under dry, rainy, and storm conditions are used for model construction, with 80% of each sequence used for training and the remaining 20% for validation. For MFD-Mamba, the window length is set to 32. The Mamba feature dimension and state dimension are 32 and 8, respectively, with two Mamba layers, a convolution width of 3, an expansion factor of 2, and a dropout rate of 0.1. Four overlapping process blocks are used. The model is trained using AdamW with a learning rate of 3 × 10 4 , a batch size of 64, and a maximum of 80 epochs. For MambaAD, the window length is 64, with a patch length of 8 and an overlap of 4. The model and state dimensions are 64 and 32, respectively. The learning rate is 1 × 10 4 , the batch size is 64, and the number of training epochs is 80. The temporal and signal masking ratios are both set to 0.1. For IFD, the window length and latent dimension are set to 10 and 32, respectively. The neighborhood sizes of the multiscale graph are 8, 16, and 24. The graph and stability regularization coefficients are set to 0.02 and 0.001. The model is trained for 200 epochs using Adam with a learning rate of 1 × 10 3 . The control limits are estimated from normal validation data with a confidence level of 0.99.
Two multiple-fault scenarios are considered, including overlapping and sequential faults. In the overlapping-fault scenario, Fault a and Fault b are introduced simultaneously. Fault a represents an aeration loss condition in which the oxygen transfer coefficient in the third aeration tank, K L a 3 , is reduced from 240 to 120 d−1. Fault b represents an influent nitrate shock, where the influent S N O is increased by 15 mgN/L. For the sequential-fault scenario, Fault c is followed by Fault d. Fault c represents an influent organic load shock, where the influent S S and X S are increased to twice their normal values. Fault d represents an influent nitrate shock, with the influent S N O increased by 15 mgN/L. No recovery interval is introduced between the two fault stages. Figure 11 presents the schematic structure of the ASP, including the biological reaction units, secondary clarifier, and internal recycle streams.
Figure 12, Figure 13, Figure 14 and Figure 15 compare the detection performance of MFD-Mamba, MambaAD, IFD, and MOLA under the overlapping-fault condition. MFD-Mamba responds rapidly after fault onset and maintains a clear separation between the normal and faulty operating regions. In Figure 13, MambaAD also detects the overlapping faults, although a short threshold crossing occurs during normal operation. IFD exhibits larger fluctuations and several threshold exceedances before fault occurrence, indicating a higher tendency toward false alarms in Figure 14. In Figure 15, MOLA shows a relatively slow increase after fault onset and crosses the control limit only after a noticeable transition period, resulting in a longer detection delay.
Figure 16, Figure 17, Figure 18 and Figure 19 compare the detection performance of MFD-Mamba, MambaAD, IFD, and MOLA under the sequential-fault condition. MFD-Mamba responds promptly to the first fault and maintains a clear abnormal response after the transition to the second fault in Figure 17. MambaAD also detects both fault stages, but its monitoring index exhibits larger variations around the fault transitions. IFD shows several pronounced peaks and stronger fluctuations during the sequential-fault period in Figure 18. MOLA responds to the faults, but exhibits a more irregular monitoring trajectory, particularly around the transition between the two fault stages.
The quantitative results are summarized in Table 4. MFD-Mamba provides the best overall performance under both overlapping-fault and sequential-fault conditions, with consistently high detection accuracy, low false alarm rates, and short detection delays. As shown in Table 5, the FDR improvements over MambaAD and MOLA are statistically significant under the overlapping-fault condition, whereas the difference relative to IFD is not significant. The comparison with single-block Mamba also shows a significant improvement, supporting the contribution of the proposed multiblock decomposition and fusion strategy.
Under the sequential-fault condition, significant FDR improvements are observed over all comparison methods, including single-block Mamba. Although some competing methods achieve relatively high FDRs, they generally exhibit higher false alarm rates or longer detection delays. The ASP results indicate that the performance gain does not arise from the Mamba backbone alone. The proposed multiblock decomposition and fusion strategy improves fault evidence selectivity and monitoring reliability under multiple-fault conditions.

4.4. Computational Efficiency Analysis

To evaluate the computational cost of the different methods, the number of trainable parameters, model size, total training time, and per-sample inference time are measured for MFD-Mamba, Transformer, MambaAD, IFD, and MOLA. Training time is measured as the wall-clock time required to complete the corresponding training procedure. For inference-time measurement, the batch size is set to one. After 50 warm-up runs, the average latency is calculated over 500 forward passes. Data loading and preprocessing are excluded from the inference-time measurement. The reported computational costs correspond to the sequence length adopted by each method in the fault detection experiments. For the TE process, the sequence lengths of MFD-Mamba, Transformer, MambaAD, IFD, and MOLA are 32, 32, 64, 10, and 10, respectively. For the ASP, the corresponding sequence lengths are 32, 32, 64, 10, and 32.
In addition to the empirical measurements, the computational complexity with respect to sequence length is considered. Let L denote the sequence length, d the hidden dimension, N the state dimension, and B the number of process blocks. The self-attention operation in a standard transformer has a sequence-dependent computational complexity of approximately O ( L 2 d ) due to the pairwise attention computation. In contrast, the selective state-space operation used by Mamba scales linearly with L, with an approximate complexity of O ( L d N ) . Therefore, for MFD-Mamba, the dominant temporal modeling term over multiple process blocks is expressed approximately as O ( B L d N ) . The subsequent block relation and evidence fusion operations depend mainly on the number of process blocks, and do not introduce a quadratic dependence on the sequence length.
As shown in Table 6, MFD-Mamba remains compact in terms of parameter count and model storage. Its parameter count and model size are considerably lower than those of MambaAD on both datasets, and it also requires less training and inference time than MambaAD. The Transformer baseline exhibits lower empirical latency under the relatively short input windows used in the present experiments. MOLA also requires lower inference latency than MFD-Mamba. Therefore, the present results do not indicate that MFD-Mamba is the fastest method under short-sequence conditions; instead, the computational motivation for the Mamba architecture mainly lies in its linear dependence on sequence length, in contrast to the quadratic dependence of standard self-attention (see Table 7).
These results demonstrate that MFD-Mamba achieves the reported multiple-fault detection performance with a compact model and millisecond-level CPU inference. Although its empirical runtime is higher than that of the Transformer baseline, IFD, and MOLA under the present short-window settings, its state-space sequence modeling avoids the quadratic sequence-length dependence of self-attention. The additional computational cost mainly arises from multiblock temporal modeling, dynamic block relation modeling, and the subsequent evidence fusion procedure.

4.5. Interpretability Analysis

The block-wise Bayesian structure of MFD-Mamba provides a direct form of model interpretability. Prior to global evidence fusion, each process block generates posterior fault evidence P b , t F , the magnitude of which reflects the contribution of the corresponding block to the detected abnormal condition. For a given operating stage T s , the mean Bayesian fault evidence is calculated as
P ¯ b , s F = 1 | T s | t T s P b , t F ,
where the block with the largest P ¯ b , s F is regarded as the dominant contributing block.
As shown in Table 8, overlapping-fault conditions activate multiple process blocks, while the dominant contributing block can still be identified from the Bayesian evidence. Block 3 provides the strongest contribution in the TE process, whereas Block 2 is dominant in the ASP process.
Under the sequential-fault condition, the dominant Bayesian evidence changes with the operating stage. In the TE process, the dominant contribution shifts from Block 1 during the first fault to Block 4 during the second fault, while the block evidences decrease to near-zero levels during recovery. A similar transition from Block 1 to Block 4 is observed in the ASP process. These stage-wise changes illustrate how the Bayesian fault evidence evolves as the operating condition changes.
The block-wise evidence provides an interpretable indication of which process region contributes most strongly to the monitoring decision and supports block-level fault localization; however, variable-level localization requires an additional within-block contribution analysis, and is left for future work.

4.6. Sensitivity to the Number of Process Blocks

To evaluate the sensitivity of MFD-Mamba to the number of process blocks, B is varied from 2 to 6 on both the TE and ASP processes, while the remaining settings are kept unchanged. As shown in Table 9, the monitoring performance generally improves as B increases from 2 to 4.
On the TE process, B = 4 and B = 5 provide the best overall monitoring performance, whereas a further increase to B = 6 leads to noticeable deterioration, particularly under the overlapping-fault condition. On the ASP, the performance improves markedly from B = 2 to B = 4 , then remains relatively stable for larger values of B. These results indicate that an excessively small number of blocks may merge process regions with different dynamic characteristics, whereas an excessively large number of blocks may fragment correlated process information without providing additional monitoring benefits. Therefore, B = 4 is adopted in the final configuration as a balance between monitoring performance and model simplicity.

4.7. Ablation Study

Ablation experiments are conducted on the TE process and the ASP to evaluate the contributions of the main components of MFD-Mamba. The Mamba module and Bayesian fusion strategy are examined on both processes, while the final detector and block fusion mechanism are further evaluated on the TE and ASP processes, respectively. The results are summarized in Table 10 and Table 11.
For the TE process, removing Mamba mainly degrades the fault detection rate and response speed, whereas removing Bayesian fusion increases the false alarm rate and detection delay. The final detector also plays an important role in alarm reliability. Similar trends are observed on the ASP, where both the removal of Mamba and the replacement of the structured block fusion mechanism lead to clear performance degradation. The results indicate that the performance gain of MFD-Mamba does not arise from the Mamba backbone alone. Mamba-based temporal modeling, multiblock organization, and evidence fusion provide complementary contributions to fault detection, response speed, and alarm reliability.

5. Conclusions

This paper has presented MFD-Mamba for multiple-fault process monitoring in dynamic industrial processes. The proposed framework combines hybrid process decomposition, blockwise selective state-space modeling, dynamic block relation modeling, and fault evidence fusion. Data-derived lagged relationships and prior process connections are used to construct soft variable memberships and block relation information. The resulting blockwise sequences are processed by a shared modified Mamba encoder, and the block representations are subsequently updated according to their current interactions. Blockwise prediction residuals are then integrated through Bayesian fusion and combined with state deviation and persistence information to form the final monitoring decision. The proposed method was evaluated on the TE and the ASP under overlapping and sequential multiple-fault conditions. Comparative experiments showed that MFD-Mamba maintained high fault detection rates while controlling false alarm rates and detection delays across the considered fault scenarios. The ablation results further showed that temporal modeling, block relation modeling, block-level fusion, and the final detection components contribute to different aspects of monitoring performance.

Author Contributions

Conceptualization, K.L. and S.W.; methodology, K.L.; software, G.C.; validation, K.L.; formal analysis, K.L.; investigation, K.L.; resources, C.S.; data curation, K.L.; writing—original draft preparation, C.S.; writing—review and editing, G.C., S.W., and C.S.; visualization, K.L.; supervision, C.S. and S.W.; project administration, S.W.; funding acquisition, C.S. and S.W. All authors have read and agreed to the published version of the manuscript.

Funding

This work was funded in part by the National Natural Science Foundation of China under Grant 62403468 and in part by the Postdoctoral Fellowship Program of CPSF under Grant GZC20241921.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The data used in this study are publicly available. The Tennessee Eastman (TE) process dataset is available at https://github.com/camaramm/tennessee-eastman-profBraatz (accessed on 31 August 2026), and the Benchmark Simulation Model No. 1 (BSM1) data and simulation resources are available at https://iwa-mia.org/benchmarking/ (accessed on 31 August 2026). No new datasets were generated in this study.

Acknowledgments

During the preparation of this manuscript, the authors used ChatGPT (GPT-5.5), developed by OpenAI, only for language editing, including grammar, spelling, punctuation, and formatting improvement. The authors have reviewed and edited the output and take full responsibility for the content of this publication.

Conflicts of Interest

Author Shijun Wu was employed by Guixi Dasanyuan Industry (Group) Co., Ltd. Author Guangbo Chen was employed by China Petroleum Pipeline Engineering Corporation. The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

References

  1. Gao, L.; Chen, Z.; Ji, H.; Niu, Y.; Dong, H. Correlation-Aided Neural Network for Distributed Process Monitoring of Large-Scale Industrial Automation Systems. IEEE Trans. Ind. Inform. 2025, 21, 9526–9537. [Google Scholar] [CrossRef] [Scilit]
  2. Chen, Y.; Ma, H.; Wang, Y.; Liu, X.; Zhang, X. Deep Feature Correlation Analysis for KPI-Related Fault Detection with Missing Data. IEEE Trans. Autom. Sci. Eng. 2026, 23, 11340–11357. [Google Scholar] [CrossRef] [Scilit]
  3. Huo, M.; Yin, S. Interpretable Fault Diagnosis for Industrial Processes: A Unified Perspective. Artif. Intell. Sci. Eng. 2026, 2, 101–120. [Google Scholar] [CrossRef] [Scilit]
  4. Song, B.; Zhang, J.; Wang, L.; Shi, H.; Tao, Y.; Li, Z.; Zheng, C. Fault Path Tracing and Root Cause Diagnosis Based on Semifrozen Autoencoder. IEEE Trans. Instrum. Meas. 2026, 75, 3513310. [Google Scholar] [CrossRef] [Scilit]
  5. Si, Y.; Wang, Y.; Zhou, D. Key-performance-indicator-related process monitoring based on improved kernel partial least squares. IEEE Trans. Ind. Electron. 2020, 68, 2626–2636. [Google Scholar] [CrossRef] [Scilit]
  6. Song, Y.; Zhang, A. From black box to physically interpretable: Trustworthy computing for AI-driven decision-making and control. J. Autom. Intell. 2026, 5, 91–111. [Google Scholar] [CrossRef] [Scilit]
  7. Luo, H.; Wu, S.; Li, L. A time-domain estimation of the ν-gap metric and stability margin with its application to attack defense in cyber-physical systems. Automatica 2025, 182, 112549. [Google Scholar] [CrossRef] [Scilit]
  8. Sun, C.Y.; Yang, G.H. A Quality-Relevant Fault Diagnosis Scheme Aided by Enhanced Dynamic Just-in-Time Learning for Nonlinear Industrial Systems. IEEE Trans. Ind. Inform. 2024, 20, 7471–7480. [Google Scholar] [CrossRef] [Scilit]
  9. Wu, T.; Wang, L.; Xu, X.; Su, L.; He, W.; Wang, X. An intelligent fault detection algorithm for power transmission lines based on multi-scale fusion. Intell. Robot. 2025, 5, 474–487. [Google Scholar] [CrossRef] [Scilit]
  10. Li, K.; Feng, J.; Sun, C.; Jiang, Y. TF-LGNet: A Time-Frequency Guided Local-Global Feature Network for Nonintrusive Load Monitoring. IEEE Sens. J. 2026, 26, 17262–17272. [Google Scholar] [CrossRef] [Scilit]
  11. Mishra, R.K.; Choudhary, A.; Fatima, S.; Mohanty, A.R.; Panigrahi, B.K. A Systematic Review on Advancement and Challenges in Multi-Fault Diagnosis of Rotating Machines. Eng. Appl. Artif. Intell. 2025, 156, 111306. [Google Scholar] [CrossRef] [Scilit]
  12. Jiang, Y.; Zhang, Z.; Mei, Y.; Zhang, M.; Mo, J.; Zio, E. A Bias-Variance Balanced Solution for Vibration-Based Fault Diagnosis of Tunnel Boring Machine Main Bearing. IEEE Sens. J. 2026, 26, 21360–21368. [Google Scholar] [CrossRef] [Scilit]
  13. Sun, R.; Yin, X.; Wang, R. Hybrid kernel deep orthogonal subspace analysis-based anomaly detection for key performance indicators. Chemom. Intell. Lab. Syst. 2026, 278, 105874. [Google Scholar] [CrossRef] [Scilit]
  14. Lee, G.; Han, C.; Yoon, E.S. Multiple-Fault Diagnosis of the Tennessee Eastman Process Based on System Decomposition and Dynamic PLS. Ind. Eng. Chem. Res. 2004, 43, 8037–8048. [Google Scholar] [CrossRef] [Scilit]
  15. Zhao, C.; Gao, F. Fault Subspace Selection Approach Combined with Analysis of Relative Changes for Reconstruction Modeling and Multifault Diagnosis. IEEE Trans. Control Syst. Technol. 2016, 24, 928–939. [Google Scholar] [CrossRef] [Scilit]
  16. Ma, L.; Dong, J.; Peng, K.; Zhang, C. Hierarchical Monitoring and Root-Cause Diagnosis Framework for Key Performance Indicator-Related Multiple Faults in Process Industries. IEEE Trans. Ind. Inform. 2019, 15, 2091–2100. [Google Scholar] [CrossRef] [Scilit]
  17. Ma, L.; Yang, P.; Peng, K. Multitask Learning Based Collaborative Modeling of Heterogeneous Data for Compound Fault Diagnosis in Manufacturing Processes. IEEE Trans. Ind. Inform. 2024, 20, 14174–14183. [Google Scholar] [CrossRef] [Scilit]
  18. Zhou, K.; Tong, Y.; Wei, X.; Song, K.; Chen, X. A Novel Multi-Label Classification Deep Learning Method for Hybrid Fault Diagnosis in Complex Industrial Processes. Measurement 2025, 242, 115804. [Google Scholar] [CrossRef] [Scilit]
  19. Zhu, C.; Liu, N.; Li, Z.; Li, Y.; Guo, J.; Ji, L.; Zhao, Y.; Shi, X.; Lan, X. A Multi-Label Domain Adaptation Method for Unseen Compound Fault Recognition in FCC Settler Catalyst Loss Processes. Process Saf. Environ. Prot. 2025, 203, 107885. [Google Scholar] [CrossRef] [Scilit]
  20. Yang, C.; Cai, B.; Kong, X.; Zhang, R.; Gao, C.; Liu, X.; Shao, H.; Zhao, X. Fast and Stable Fault Diagnosis Method for Composite Fault of Subsea Production System. Mech. Syst. Signal Process. 2025, 226, 112373. [Google Scholar] [CrossRef] [Scilit]
  21. Eydam, L.; Lange, H.; Urbas, L. Hybrid Fault Diagnosis in the Process and Chemical Industries: A Review. Annu. Rev. Control 2026, 62, 101064. [Google Scholar] [CrossRef] [Scilit]
  22. Ma, F.; Ji, C.; Wang, J.; Sun, W.; Tang, X.; Jiang, Z. MOLA: Enhancing Industrial Process Monitoring Using a Multi-Block Orthogonal Long Short-Term Memory Autoencoder. Processes 2024, 12, 2824. [Google Scholar] [CrossRef] [Scilit]
  23. Xue, C.; Zhang, T.; Xiao, D. Output Related Fault Detection and Diagnosis Based on Multi Block Modified Orthogonal Broyden-Fletcher-Goldfarb-Shanno Algorithm. Neurocomputing 2024, 607, 128350. [Google Scholar] [CrossRef] [Scilit]
  24. Yan, S.; Yang, Y.; Ran, J.; Xiong, T.; Huang, H.; Zhu, J.; Yang, X.; Zhai, C.; Zhang, G.; Zhang, H. Fast and interpretable just-in-time learning via information-weighted similarity search: Probabilistic soft sensing of catalyst coke in MTO process. Chem. Eng. Sci. 2026, 330, 123911. [Google Scholar] [CrossRef] [Scilit]
  25. Zhu, Z.; Chen, F.; Ni, L.; Bian, H.; Jiang, J.; Chen, Z. A Novel Transformer-Based Model with Large Kernel Temporal Convolution for Chemical Process Fault Detection. Comput. Chem. Eng. 2024, 188, 108762. [Google Scholar] [CrossRef] [Scilit]
  26. Zhang, H.; Li, Y.; Zhang, X.; Feng, H. Multi-Scale Group Mamba Network with Structural Attention for Rotating Machinery Fault Diagnosis Using Multisensor Data. Adv. Eng. Inform. 2025, 67, 103521. [Google Scholar] [CrossRef] [Scilit]
  27. Qin, S.; Zhu, J.; Guo, A.; Yang, Y.; Wang, L.; Tao, G. MambaAD: Multivariate time series anomaly detection in IoT via multi-view Mamba. Neurocomputing 2025, 655, 131385. [Google Scholar] [CrossRef] [Scilit]
  28. Li, G.; Ge, M.; Wan, J.; Han, D.; Li, M.; Zhou, M. MemMambaAD: Memory-augmented state space model for multivariate time series anomaly detection. Eng. Appl. Artif. Intell. 2025, 158, 111308. [Google Scholar] [CrossRef] [Scilit]
  29. Yang, P.; Xie, Z.; Sheng, S. Dual-path mamba for chemical process fault diagnosis with adaptive signal decomposition. Chin. J. Chem. Eng. 2026; in press. [CrossRef] [Scilit]
  30. Mishra, R.K.; Choudhary, A.; Fatima, S.; Mohanty, A.R.; Panigrahi, B.K. Adaptive Information Fusion for Multi-Fault Diagnosis Using Multi-Stream Convolutional Neural Network. Measurement 2026, 272, 121043. [Google Scholar] [CrossRef] [Scilit]
  31. Wang, G.; Yang, J.; Qian, Y.; Han, J.; Jiao, J. KPCA-CCA-Based Quality-Related Fault Detection and Diagnosis Method for Nonlinear Process Monitoring. IEEE Trans. Ind. Inform. 2023, 19, 6492–6501. [Google Scholar] [CrossRef] [Scilit]
  32. Sun, C.; Wang, X.; Cheng, Y. A Quality-Related Incipient Fault Detection and Location for Industrial IoT with Incomplete Data. IEEE Internet Things J. 2026, 13, 23524–23536. [Google Scholar] [CrossRef] [Scilit]
  33. Alex, J.; Benedetti, L.; Copp, J.; Gernaey, K.V.; Jeppsson, U.; Nopens, I.; Pons, M.N.; Steyer, J.P.; Vanrolleghem, P.A. Benchmark Simulation Model No. 1 (BSM1); Technical Report TEIE-7229; Lund University: Lund, Sweden, 2008. [Google Scholar]
Figure 1. Framework of the proposed MFD-Mamba method.
Figure 1. Framework of the proposed MFD-Mamba method.
Processes 14 02973 g001
Figure 2. Schematic diagram of the Tennessee Eastman process.
Figure 2. Schematic diagram of the Tennessee Eastman process.
Processes 14 02973 g002
Figure 3. Overlapping fault detection results of MFD-Mamba for TE.
Figure 3. Overlapping fault detection results of MFD-Mamba for TE.
Processes 14 02973 g003
Figure 4. Overlapping fault detection results of MambaAD for TE.
Figure 4. Overlapping fault detection results of MambaAD for TE.
Processes 14 02973 g004
Figure 5. Overlapping fault detection results of IFD for TE.
Figure 5. Overlapping fault detection results of IFD for TE.
Processes 14 02973 g005
Figure 6. Overlapping fault detection results of MOLA for TE.
Figure 6. Overlapping fault detection results of MOLA for TE.
Processes 14 02973 g006
Figure 7. Sequential fault detection results of MFD-Mamba for TE.
Figure 7. Sequential fault detection results of MFD-Mamba for TE.
Processes 14 02973 g007
Figure 8. Sequential fault detection results of MambaAD for TE.
Figure 8. Sequential fault detection results of MambaAD for TE.
Processes 14 02973 g008
Figure 9. Sequential fault detection results of IFD for TE.
Figure 9. Sequential fault detection results of IFD for TE.
Processes 14 02973 g009
Figure 10. Sequential fault detection results of MOLA for TE.
Figure 10. Sequential fault detection results of MOLA for TE.
Processes 14 02973 g010
Figure 11. Schematic diagram of the ASP.
Figure 11. Schematic diagram of the ASP.
Processes 14 02973 g011
Figure 12. Overlapping fault detection result of MFD-Mamba for the ASP.
Figure 12. Overlapping fault detection result of MFD-Mamba for the ASP.
Processes 14 02973 g012
Figure 13. Overlapping fault detection result of MambaAD for the ASP.
Figure 13. Overlapping fault detection result of MambaAD for the ASP.
Processes 14 02973 g013
Figure 14. Overlapping fault detection result of IFD for the ASP.
Figure 14. Overlapping fault detection result of IFD for the ASP.
Processes 14 02973 g014
Figure 15. Overlapping fault detection result of MOLA for the ASP.
Figure 15. Overlapping fault detection result of MOLA for the ASP.
Processes 14 02973 g015
Figure 16. Sequential fault detection result of MFD-Mamba for the ASP.
Figure 16. Sequential fault detection result of MFD-Mamba for the ASP.
Processes 14 02973 g016
Figure 17. Sequential fault detection result of MambaAD for the ASP.
Figure 17. Sequential fault detection result of MambaAD for the ASP.
Processes 14 02973 g017
Figure 18. Sequential fault detection result of IFD for the ASP.
Figure 18. Sequential fault detection result of IFD for the ASP.
Processes 14 02973 g018
Figure 19. Sequential fault detection result of MOLA for the ASP.
Figure 19. Sequential fault detection result of MOLA for the ASP.
Processes 14 02973 g019
Table 1. Hyperparameter settings of MFD-Mamba for the TE process and ASP.
Table 1. Hyperparameter settings of MFD-Mamba for the TE process and ASP.
HyperparameterSymbolTEASP
Process-prior weight λ p 0.450.45
Sparsity coefficient λ s 0.0150.015
Non-prior penalty λ n 0.080.08
Residual-tail weight γ 0.300.30
Consistency weight λ c 0.050.05
Relation-prior strength λ r 0.750.75
EWMA smoothing coefficient λ e 0.400.40
Fault prior π F 0.010.01
Number of process blocksB44
Membership threshold δ 0.300.30
Block control confidence level α b 0.990.99
Global calibration confidence level α g 0.990.995
EWMA confidence level α e 0.9990.999
Global Bayesian persistence K / N 2 / 3 2 / 3
Moderate Mahalanobis persistence K / N 2 / 3 2 / 3
Cross-evidence persistence K / N 2 / 2 2 / 2
Table 2. Fault detection performance of different methods for the TE process, reported as mean ± standard deviation over six independent runs.
Table 2. Fault detection performance of different methods for the TE process, reported as mean ± standard deviation over six independent runs.
MethodFault ScenarioFDR (%)FAR (%)Delay
MFD-MambaOverlapping fault99.88 ± 0.030.13 ± 0.010 ± 0.00
Single-block MambaOverlapping fault97.69 ± 0.828.98 ± 4.199.00 ± 0.00
MambaADOverlapping fault98.75 ± 0.700.30 ± 0.0839 ± 5.64
IFDOverlapping fault95.13 ± 0.117.28 ± 3.449 ± 2.16
MOLAOverlapping fault85.69 ± 6.972.21 ± 1.3037.17 ± 31.51
MFD-MambaSequential fault99.81 ± 0.020.20 ± 0.062 ± 0.00/0 ± 0.10
Single-block MambaSequential fault98.88 ± 0.667.66 ± 3.358.00 ± 0.00/1.00 ± 0.00
MambaADSequential fault99.25 ± 2.091.35 ± 1.027 ± 0.98/2 ± 0.52
IFDSequential fault98.56 ± 0.033.81 ± 2.102 ± 0.41/1 ± 0.50
MOLASequential fault94.70 ± 4.6456.82 ± 0.5815.67 ± 13.78/3 ± 0.80
Note: Bold values indicate the best-performing results.
Table 3. Statistical significance analysis of the FDR on the TE process.
Table 3. Statistical significance analysis of the FDR on the TE process.
Fault ScenarioComparisonp-ValueSignificance
OverlappingMFD-Mamba vs. Single-block Mamba0.0012**
MFD-Mamba vs. MambaAD0.0108*
MFD-Mamba vs. IFD<0.0001***
MFD-Mamba vs. MOLA0.0042**
SequentialMFD-Mamba vs. Single-block Mamba0.0182*
MFD-Mamba vs. MambaAD0.5406n.s.
MFD-Mamba vs. IFD<0.0001***
MFD-Mamba vs. MOLA0.0429*
* p < 0.05, ** p < 0.01, *** p < 0.001; n.s. denotes no statistically significant difference.
Table 4. Fault detection performance of different methods for the ASP.
Table 4. Fault detection performance of different methods for the ASP.
MethodFDR (%)FAR (%)Delay
Overlapping fault
MFD-Mamba99.79 ± 0.010.78 ± 0.011 ± 0.00
Single-block Mamba95.58 ± 2.509.38 ± 3.388.00 ± 0.00
MambaAD99.37 ± 0.762.34 ± 1.483 ± 0.00
IFD98.95 ± 0.1115.30 ± 5.675 ± 0.52
MOLA82.23 ± 6.379.21 ± 0.5110.80 ± 3.16
Sequential fault
MFD-Mamba99.58 ± 0.031.56 ± 0.070 ± 0.00/2 ± 0.00
Single-block Mamba96.79 ± 2.3810.62 ± 6.4110.00 ± 0.00/8.00 ± 0.00
MambaAD99.17 ± 2.522.34 ± 1.483 ± 2.66/0 ± 0.00
IFD98.96 ± 0.3215.30 ± 5.673 ± 0.00/2 ± 0.00
MOLA82.23 ± 8.9811.21 ± 1.513.75 ± 0.90/7.75 ± 1.01
Note: Bold values indicate the best-performing results.
Table 5. Statistical significance analysis of the FDR on the ASP.
Table 5. Statistical significance analysis of the FDR on the ASP.
Fault ScenarioComparisonp-ValueSignificance
OverlappingMFD-Mamba vs. Single-block Mamba0.0091**
MFD-Mamba vs. MambaAD0.0313*
MFD-Mamba vs. IFD0.1250n.s.
MFD-Mamba vs. MOLA0.0313*
SequentialMFD-Mamba vs. Single-block Mamba0.0349*
MFD-Mamba vs. MambaAD0.0313*
MFD-Mamba vs. IFD0.0313*
MFD-Mamba vs. MOLA0.0313*
* p < 0.05 , ** p < 0.01 ; n.s. denotes no statistically significant difference.
Table 6. Computational efficiency comparison of different methods on the TE process and ASP.
Table 6. Computational efficiency comparison of different methods on the TE process and ASP.
DatasetMethodParameters
(K)
Model Size
(MB)
Training Time
(s) 
Inference Time
(ms/Sample) 
TEMFD-Mamba36.770.15815.564.378
Transformer27.650.1159.950.286
MambaAD233.440.919142.316.855
IFD72.170.2807.800.051
MOLA12.190.05911.011.034
ASPMFD-Mamba36.250.156320.805.224
Transformer27.130.11335.980.287
MambaAD221.370.873500.355.792
IFD56.730.22119.240.056
MOLA11.580.05725.251.201
Table 7. Sequence -length dependence of representative sequence models.
Table 7. Sequence -length dependence of representative sequence models.
ModelDominant Sequence Modeling CostDependence on L
Transformer O ( L 2 d ) Quadratic
LSTM/GRU O ( L d 2 ) Linear
Mamba O ( L d N ) Linear
MFD-Mamba O ( B L d N ) Linear
Table 8. Mean block-wise Bayesian fault evidence under representative multiple-fault scenarios.
Table 8. Mean block-wise Bayesian fault evidence under representative multiple-fault scenarios.
DatasetOperating StageBlock 1Block 2Block 3Block 4Dominant Block
TEOverlapping fault0.8900.9400.9440.937Block 3
TEFirst fault0.9840.2620.3240.411Block 1
TERecovery0.0020.0020.0010.002
TESecond fault0.1420.6500.0010.997Block 4
ASPOverlapping fault0.8990.9330.6320.711Block 2
ASPFirst fault0.8010.7660.6740.481Block 1
ASPSecond fault0.7100.7530.8300.889Block 4
Note: Bold values indicate the best-performing results.
Table 9. Sensitivity of MFD-Mamba to the number of process blocks B on the TE and ASP processes.
Table 9. Sensitivity of MFD-Mamba to the number of process blocks B on the TE and ASP processes.
Overlapping FaultSequential Fault
Dataset B FDR (%)FAR (%)DelayFDR (%)FAR (%)Delay
TE298.123.11699.251.074/2
TE399.261.22299.250.844/2
TE499.880.12099.810.202/0
TE599.880.12099.810.202/0
TE697.752.661799.253.204/2
ASP295.839.01392.507.5014/0
ASP398.543.99297.212.877/0
ASP499.790.78199.581.562/0
ASP599.790.78299.581.562/0
ASP699.790.78299.581.562/0
Note: Bold values indicate the best-performing results.
Table 10. Ablation results of MFD-Mamba for the TE process.
Table 10. Ablation results of MFD-Mamba for the TE process.
MethodFault ScenarioFDR (%)FAR (%)Delay
MFD-MambaOverlapping fault99.880.130
W/o MambaOverlapping fault96.253.7530
W/o Bayesian fusionOverlapping fault98.131.889
W/o final detectorOverlapping fault97.509.3715
MFD-MambaSequential fault99.810.202/0
W/o MambaSequential fault99.255.637/2
W/o Bayesian fusionSequential fault97.634.375/2
W/o final detectorSequential fault98.563.817/10
Note: Bold values indicate the best-performing results.
Table 11. Ablation results of MFD-Mamba for the ASP.
Table 11. Ablation results of MFD-Mamba for the ASP.
MethodScenarioFDR (%)FAR (%)Delay
MFD-MambaOverlapping99.790.781
W/o Bayesian fusionOverlapping98.252.605
W/o equal block fusionOverlapping91.674.172
W/o MambaOverlapping89.381.049
MFD-MambaSequential99.581.560/2
W/o Bayesian fusionSequential98.833.138/2
W/o equal block fusionSequential96.255.452/5
W/o MambaSequential90.427.1619/0
Note: Bold values indicate the best-performing results.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Sun, C.; Wu, S.; Chen, G.; Li, K. MFD-Mamba: A Mamba-Based Framework for Multiple-Fault Process Monitoring in Industrial Systems. Processes 2026, 14, 2973. https://doi.org/10.3390/pr14182973

AMA Style

Sun C, Wu S, Chen G, Li K. MFD-Mamba: A Mamba-Based Framework for Multiple-Fault Process Monitoring in Industrial Systems. Processes. 2026; 14(18):2973. https://doi.org/10.3390/pr14182973

Chicago/Turabian Style

Sun, Chengyuan, Shijun Wu, Guangbo Chen, and Keqin Li. 2026. "MFD-Mamba: A Mamba-Based Framework for Multiple-Fault Process Monitoring in Industrial Systems" Processes 14, no. 18: 2973. https://doi.org/10.3390/pr14182973

APA Style

Sun, C., Wu, S., Chen, G., & Li, K. (2026). MFD-Mamba: A Mamba-Based Framework for Multiple-Fault Process Monitoring in Industrial Systems. Processes, 14(18), 2973. https://doi.org/10.3390/pr14182973

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop