Next Article in Journal
An Activity-Centric Intelligent Information System for Context-Aware Recommendation in Exploratory Journalistic Research
Next Article in Special Issue
GaussianCopula-Based Synthetic Data Generation for Turbocharger Fault Scenario Simulation and SFOC Degradation Modelling in Two-Stroke Marine Diesel Engines
Previous Article in Journal
Evaluating Creative Methods in AIGC-Assisted Fashion Design Sketch Generation: Mind Mapping, Brainstorming, and SCAMPER
Previous Article in Special Issue
Design of a Lightweight Edge-AI System for Predictive Maintenance on ESP32-S3
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Data-Driven Technique for Fault Detection and Localization of Air Quality Process

1
Faculty of Computers and Information Technology, University of Tabuk, Tabuk 71491, Saudi Arabia
2
National Engineering School of Gabes, University of Gabes, Gabes 6029, Tunisia
3
Department of Computer Science, University of Tabuk, Tabuk 71491, Saudi Arabia
4
Department of Management Information Systems, University of Tabuk, Tabuk 71491, Saudi Arabia
*
Author to whom correspondence should be addressed.
Appl. Sci. 2026, 16(11), 5674; https://doi.org/10.3390/app16115674
Submission received: 5 April 2026 / Revised: 23 May 2026 / Accepted: 25 May 2026 / Published: 5 June 2026

Abstract

Air pollution is primarily caused by human activities such as industrial emissions, road traffic, waste incineration, and fossil fuel power plants. Pollution refers to the presence of harmful substances in the air, such as nitrogen dioxide (NO2), sulfur dioxide (SO2), ozone (O3), carbon monoxide (CO), and other environmental pollutants. Some pollutants pose health risks even at low doses. Given the critical importance of air quality, monitoring air pollution has become an urgent and essential subject. Air quality monitoring relies on accurate data, so changeable environments and sensor issues make using interval diagnostic techniques for addressing uncertainty in systems interesting. In this article, we focus on three key aspects to achieve precise and efficient results: (1) the use of an accurate fault detection method that accounts for data uncertainty while maintaining model symmetry, (2) the implementation of a reliable detection index invariant to symmetric sensor behaviors, and (3) the combination of both to improve fault localization accuracy. This paper presented a fault detection and localization framework designed for uncertain and nonlinear monitoring environments. A novel fault-sensitive detection index was developed and integrated into an elimination-based localization strategy within a reduced-rank interval kernel PCA (RR-IKPCA) model. By exploiting information contained in modified residual subspaces and explicitly accounting for measurement uncertainty, the proposed approach enhances fault sensitivity while preserving robust localization capability, as validated on the AIRLOR air quality monitoring network.

1. Introduction

Continuous monitoring systems have become indispensable in modern environmental and industrial infrastructures, particularly with the rapid deployment of large-scale air quality monitoring. Over the past decades, diagnostic monitoring methodologies have evolved significantly, transitioning from model-based and analytical redundancy approaches toward advanced data-driven techniques [1,2,3]. Early methods relied heavily on first-principles models [4], which, although effective in well-defined systems, were often difficult to implement in complex and dynamic environments such as urban air quality networks. These networks commonly use low-cost IoT sensors, including electrochemical sensors for gases (e.g., CO, NO2) and optical sensors for particulate matter (PM2.5, PM10).
With the increasing availability of real monitoring data, multivariate statistical approaches emerged as practical solutions for capturing correlations among variables and detecting abnormal system behavior [5,6]. These techniques have been successfully applied to real air quality monitoring datasets for identifying sensor faults, calibration issues, and data inconsistencies across monitoring stations [7]. Accurate detection and validation of measurement errors is a crucial first step in air quality network monitoring, as it allows deviations from normal operating conditions to be identified based on observed data patterns [8]. The second is fault isolation, where the goal is to pinpoint the specific sensor or subsystem responsible for the detected anomaly.
Various successful statistical fault detection methods are developed in the literature to detect faults or anomalous changes in key measured air quality variables [9,10]; this task of fault detection is very important but remains limited to testing the presence of faults [11,12]. To identify defect variables, several Principal Component Analysis (PCA)-based techniques have been developed in the literature, like the defect localization procedure using residue structuring [10,13]. The principle is to construct a set of residues, so that each residue is responsive to certain defects and not to others.
However, the nonlinear and non-Gaussian nature of air quality data has motivated the development of more sophisticated approaches, including kernel PCA (KPCA) and its variants, which provide improved representation of complex relationships. In [14,15,16,17] the authors proposed an extension for fault isolation using kernel PCA techniques (KPCA) to deal with nonlinearity.
Recent research has increasingly focused on applying these techniques to real-world monitoring networks, integrating methods such as reduced-rank models [18,19] to solve the problem of large datasets by downsizing the kernel matrix. These monitoring techniques have produced good detection and localization results but only with certain systems. However, in industrial environments, sensor measurements are often affected by imprecision [20,21]. To address this uncertainty in fault detection, numerous studies have been conducted [22,23], leading to the development of several linear Interval Principal Component Analysis (IPCA) approaches documented in the literature. In [24], an MRPCA-based EWMA scheme is proposed for sensor fault detection in air quality monitoring networks. The approach is validated through simulation studies, showing that MRPCA, as a robust interval multivariate statistical method, effectively addresses model uncertainties and improves fault detection capability.
For nonlinear scenarios, the interval kernel PCA (IKPCA) method was introduced by the authors in [25,26]. Furthermore, when dealing with interval-valued data, only a limited number of recent dimensionality reduction techniques based on IKPCA have been proposed, including reduced IKPCA (RIKPCA) [27], reduced-rank IKPCA (RRIKPCA) [28], and interval reduced-rank KPCA using the kernel generalized likelihood ratio test (IRR-KGLRT) [29]. In addition, some of these techniques have been applied to interval-based fault isolation; however, only limited research has addressed fault location in uncertain conditions. For example, the authors in [26,30] proposed a linear fault localization approach based on interval PCA using the ISPE index. Also, a localization method based on adaptive CIPCA was proposed in [13].
In addition to the importance of properly selecting the method used for fault detection, the choice of the fault detection index is also crucial. In most studies, Hotelling’s T2 and the SPE are the main indices used [11]; the former is defined in the principal subspace where eigenvalues dominant, making it difficult to detect defects [31]. The latter, Squared Prediction Error (SPE), is calculated in the residual subspace where it is more sensitive to errors, but existing work shows that this is not always the case; this index fails in some cases to give better detection, hence the idea in [32] to improve this statistic by exploiting modified residual subspaces. The proposed index captures fault-related variations that may remain undetected by classical Hotelling’s T2 and SPE statistics and demonstrates good detection performance. The aim of this study is to apply this index to the diagnosis of air quality monitoring network data.
To locate or identify a defect is defined as the capability of the monitoring process to determine the location of the defect [3,33,34]. The proposed RRIKPCA-based fault detection and localization framework demonstrates strong performance in uncertain and nonlinear monitoring environments, as validated on the AIRLOR air quality monitoring network. AIRLOR, located in Lorraine, France, comprises twenty monitoring stations distributed across rural, peri-urban, and urban sites. Each station measures key air pollutants, including NO, NO2, O3, CO, and SO2, while six stations also record additional meteorological parameters. The primary focus of this study is the detection of sensor faults measuring ozone (O3) and nitrogen oxides (NO and NO2). The observation vector contains 18 controlled variables, corresponding to the concentrations of O3, NO, and NO2 at each station. A total of 400 observations were used for training the reference model, and 1000 observations were used as testing data to evaluate the framework’s performance.
The proposed framework considers several sources of uncertainty commonly encountered in air quality monitoring systems, including meteorological variability and external environmental disturbances, sensor-related effects such as measurement noise and sensor drift due to aging, and calibration inaccuracies. By incorporating these uncertainty sources through interval modeling, the framework provides a more realistic representation of operating conditions and enhances the robustness of the monitoring process. In addition, the integration of interval modeling with the fault-sensitive detection index [D] improves the capability of identifying incipient sensor faults under highly non-stationary air quality data characterized by temporal variability and abrupt changes caused by meteorological conditions or local emission events. Unlike conventional residual-based indicators, which may fail to detect subtle deviations in such environments, the proposed index [D] effectively captures anomalies in the modified residual subspace while remaining robust to normal data fluctuations, thereby improving the reliability and timeliness of fault detection.
The elimination-based localization strategy further strengthens the approach by directly evaluating the impact of removing each variable on the detection index. Compared with reconstruction-based contribution (RBC) methods and contribution plots, which may suffer from smearing effects and ambiguous interpretations, the elimination principle provides clearer fault attribution.
Nevertheless, the proposed method has some limitations. It assumes the presence of a single dominant fault, which may reduce localization accuracy in multi-fault scenarios. In addition, the performance of the RRIKPCA model depends on the availability of representative fault-free training data.
This paper is planned as follows. In Section 2, a summary of interval kernel methods is presented in the first part, while the second part presents the mathematical formulation of the novel index [D]. In Section 3, we present the suggested RR-IKPCA localization method. Section 4 demonstrates the efficiency of the proposed fault detection and isolation method using a real air quality monitoring network data.

2. Interval Fault Detection Methods

For large majority of cases we cannot measure the true value. This impreciseness is caused by the uncertainties Δ Υ . We must take into consideration this uncertainty of the measured data Υ j ( k ) that is composed of Υ ¯ j ( k ) and Υ ¯ j ( k ) is an effect of the incertitude of measurement [26,30]:
Υ ι ( κ ) = Υ ¯ ι ( κ ) , Υ ¯ ι ( κ ) ,
and we define Υ ¯ j ( k ) = Υ j ( k ) Δ Υ as the lower bound and Υ ¯ j ( k ) = Υ j ( k ) + Δ Υ as the upper bound.
X is the form interval of the data matrix X R N × M that contains the N sample of the M process sensors. X is organized as follows:
X = Υ ¯ 1 1 , Υ ¯ 1 1 · Υ ¯ ι 1 , Υ ¯ ι 1 · Υ ¯ M 1 , Υ ¯ M 1                                     Υ ¯ 1 κ , Υ ¯ 1 κ Υ ¯ ι κ , Υ ¯ ι κ Υ ¯ M κ , Υ ¯ M κ                                     Υ ¯ 1 N , Υ ¯ 1 N · Υ ¯ ι N , Υ ¯ ι N · Υ ¯ M N , Υ ¯ M N κ = 1 , , N ;   ι = 1 , , M ,
where N is the number of process variables and M is the number of observations of each process variable.
Υ ι c ( κ ) and Υ ι r ( κ ) are assumed by
Υ ι c κ = 1 2 Υ ¯ ι κ + Υ ¯ ι κ ,   Υ ι r κ = 1 2 Υ ¯ ι κ Υ ¯ ι κ ,
A new numerical data matrix can be improved as follows:
Υ ι ( κ ) = Υ ι c ( κ ) Υ ι r ( κ ) , Υ ι c ( κ ) + Υ ι r ( κ ) ,
Y C R = Y C Y R ,
With Y C = Υ 1 c , Υ 2 c , , Υ r c R r × M and Y R = Υ 1 r , Υ 2 r , , Υ N r R r × M , the new data is Υ C R = Υ c Υ r R 2 M .
The average and the variance of Y ι are calculated by
M ( Y ι = 1 n κ = 1 n Υ ¯ ι ( κ ) + Υ ¯ ι ( κ ) 2 ,
V A R ( Y ι ) = Y ι , Y ι = κ = 1 N 1 3 ( Υ ¯ 2 ι ( κ ) + Υ ¯ ι ( κ ) Υ ¯ ι ( κ ) + Υ ¯ 2 ι ( κ ) ) ,
With the vector Y ι as
Y ι = Υ ι 1 Υ ι κ Υ ι N T = Υ ¯ ι ( 1 ) , Υ ¯ ι ( 1 ) , Υ ¯ ι ( 2 ) , Υ ¯ ι ( 2 ) , , Υ ¯ ι ( N ) , Υ ¯ ι ( N ) T ,
The standardization of Y ι ( κ ) is:
Υ ι ( κ ) M ( Y ι ( κ ) ) V A R ( Y ι ( κ ) ) = Υ ¯ ι ( κ ) M ( Y ι ( κ ) ) V A R ( Y ι ( κ ) ) , Υ ¯ ι ( κ ) M ( Y ι ( κ ) ) V A R ( Y ι ( κ ) ) ,

2.1. Review of IKPCA Methods

  • IKPCA based on interval midpoint–radii:
The authors in [21,25] suggested involving uncertainties in the KPCA approach by two models: a model based on LBs and UBs and a model centered on midpoints and radii. Recently the models have been improved in [27,29] exploiting two reduction methods.
Remember that the kernel matrix K is
K = Γ 1 T Γ 1 Γ 1 T Γ N . . . . Γ N T Γ 1 Γ N T Γ N ,
Γ is a nonlinear projection function in the characteristic space H (feature space):
H , Γ : Υ i R m Γ i = Γ ( Υ i ) R h
Γ is specified by this propriety:
Γ ( Υ ι ) T Γ ( Υ ι ) = k ( Υ ι , Υ ι ) ,
With k(.,.) as the kernel function, there are several kernel functions [35,36]. Among these functions we mainly cite the Gaussian kernel (radial basis function, RBF) that is used in this paper and defined as follows:
k ( Υ ι , Υ ι ) = exp Υ ι Υ ι 2 2 σ 2 ,
where σ is the width of a Gaussian function which controls the flexibility of the kernel. A typical choice for σ is the average minimum distance (d) between two points in the training data set, i.e., σ 2 = c 1 N i = 1 N min j i d 2 ( x i , x j ) , where c is a user-defined scaling parameter [19]. For the AIRLOR air quality monitoring dataset, ccc was selected in the range 0.5–1.0 based on preliminary experiments, balancing sensitivity to sensor faults with robustness to normal measurement variability.
The new form of kernel matrix K C R is given by
K C R = = k ( Υ C R , 1 , Υ C R , 1 ) k ( Υ C R , 1 , Υ C R , r ) . . . . k ( Υ C R , r , Υ C R , 1 ) k ( Υ C R , r , Υ C R , r ) ,
Both the ‘midpoint–radii’ and ‘upper–lower’ representations are introduced as equivalent ways to describe interval uncertainty in sensor measurements. While the midpoint–radii form provides an intuitive interpretation of the central value and deviation, the upper–lower representation is used as the primary formulation in the subsequent development of the RR-IKPCA-based framework. This choice ensures consistency in the mathematical derivation and simplifies the implementation of uncertainty modeling throughout the fault detection and localization process.
  • IKPCA upper–lower:
The K _ data is given by
K _ = k ( Υ ¯ 1 , Υ ¯ 1 ) k ( Υ ¯ 1 , Υ ¯ N ) . . . . k ( Υ ¯ N , Υ ¯ 1 ) k ( Υ ¯ N , Υ ¯ N ) ,
The kernel matrix for upper-bound data is represented by
K ¯ = k ( Υ ¯ 1 , Υ ¯ 1 ) k ( Υ ¯ 1 , Υ ¯ N ) . . . . k ( Υ ¯ N , Υ ¯ 1 ) k ( Υ ¯ N , Υ ¯ N ) ,

2.2. Review of RR-IKPCA Methods

Previous monitoring techniques have been illustrated for fault detection and isolation in air quality network data monitoring and have provided good detection and localization results [37]. However, the main disadvantage of these techniques is the computation time [11] to generate models, which increases with the number of system variables.
The core concept of the proposed RR-KPCA method is to remove dependencies among variables in the feature space while retaining a reduced but representative subset of the original dataset. Let
Y r = Υ 1 , , Υ i , , Υ r T R r × m
denote the reduced dataset, where r is the number of retained observations. The RR-KPCA-based fault detection approach operates in two stages: offline model training and online fault detection. During the offline stage, a compact reference model is constructed to capture the system’s normal behavior. This involves selecting observations that provide linearly independent representations in the transformed feature space, ensuring that only informative and system-relevant data are retained [38].
In the online stage, the derived model is applied to actively monitor the system, enabling the detection of anomalies as they occur. The parameter r is determined dynamically based on the rank of the kernel matrix, such that r a n g ( K r t ) = r .
The rank of the updated kernel matrix is calculated, its value leading to two cases:
If r a n g ( K r t ) = r , the new observation is added to the reduced matrix. Otherwise, if r a n g ( K r t ) < r no change is made and we bring back the kernel matrix to its previous state.
With
K r t = K r t 1 k x t k x t T k ( x 1 , x t ) R r × r
We obtain the reduced data matrix Y r R r × m and construct the reduced core matrix, by keeping only r observations.
For example, the reduced K C R , denoted as K C R r , becomes
K C R r = k r ( x C R , 1 , x C R , 1 ) k r ( x C R , 1 , x C R , r ) . . . . k r ( x C R , r , x C R , 1 ) k r ( x C R , r , x C R , r ) ,
With
k r ( x C R , k ) = ( k r ( x C R , 1 , x C R , k ) , , k r ( x C R , N , x C R , k ) ) T
The reduction phase of RRIKPCA based on interval midpoint–radii is summarized in Figure 1 and the model based on upper and lower bounds in Figure 2.

2.3. Fault Detection Index

  • The Center-Radius Approach Case
In classical PCA-based fault detection, the Squared Prediction Error (SPE) statistic is computed using the full residual subspace obtained after removing the dominant principal components. However, previous studies [39,40] have shown that important fault information may be contained in the least significant principal components, which are often neglected in conventional monitoring schemes.
To improve fault sensitivity for linear systems, ref. [32] proposed a detection index based on modified residual subspaces constructed by excluding a subset of dominant principal components and retaining only the last components associated with smaller eigenvalues. These components are more sensitive to incipient or low-energy faults, as they capture deviations that are not aligned with the main directions of normal process variability.
In this work, we extend this detection index to nonlinear systems by embedding it within an interval kernel PCA (IKPCA) framework. This extension preserves the sensitivity benefits of the least significant principal components while enabling the modeling of nonlinear process behavior and measurement uncertainty. The resulting index is subsequently used as the basis for fault detection and localization in the proposed methodology.
This index named Di corresponds to an SPE calculated with a PCA model with (m-i) principal components.
The parameter i is selected based on the variance contribution of the principal components and the residual analysis criterion. The proposed Di index detects abnormal deviations from the learned data correlation structure rather than high concentration values alone. Therefore, permanently high pollution levels measured near roads are treated as normal operating conditions when they are represented in the training data.
The expression of the new kernel center–range fault detection index is given by
D i C R ( k ) = k C R ( Υ C R , k , Υ C R , k ) k T ( Υ C R , k ) V C R k ( Υ C R , k ) ,
where V C R = C R C R T .
C R = B C R , 1 1 ρ C R , 1 T A C R , 1 , , B C R , l 1 ρ C R T A C R , m i ,
where A C R and B C R are the eigenvector and the eigenvalue of the reduced matrix K C R and ρ C R T is the mapped data.
If D i C R ( k ) > τ i , α 2 , a fault is detected, where τ i , α 2 represent control limits.
The detection thresholds τ i , α 2 [20] can be calculated by
τ i , α 2 = g ( i ) χ h ( i ) , α 2 ,
  • The Upper–Lower Approach Case
The interval form of D i index is represented as follows:
D ¯ i ( k ) = k ( x ¯ r , k , x ¯ r , k ) k T ( x ¯ r , k ) V ¯ r k ( x ¯ r , k ) ,
D ¯ i ( k ) = k ( x ¯ r , k , x ¯ r , k ) k T ( x ¯ r , k ) V ¯ r k ( x ¯ r , k ) ,
where
V ¯ r = ¯ r ¯ r T , ¯ r = B ¯ r , 1 1 ρ ¯ r , 1 T A ¯ r , 1 , , B ¯ r , l 1 ρ ¯ r T A ¯ r , m l , ¯ r = B ¯ r , 1 1 ρ ¯ r , 1 T A ¯ r , 1 , , B ¯ r , l 1 ρ ¯ r T A ¯ r , m l
where A ¯ r , A ¯ r and B ¯ r , B ¯ r are respectively the vectors and the eigenvalues of the reduced K r ¯ and K ¯ r .

3. Suggested Localization Method

3.1. Principle

The proposed localization method is based on the elimination principle, which is conceptually similar to the reconstruction-based contribution (RBC) approach [31,41]. Once a fault has been detected, the influence of each monitored variable on the detection index is assessed by successively removing one variable at a time from the monitoring model. This is achieved by modifying the PCA model through the elimination of the column corresponding to the excluded variable and computing the associated detection index. The relationship between the obtained index and its corresponding threshold is then evaluated through a diagnostic ratio. A significant reduction in this ratio indicates that the eliminated variable has a strong contribution to the detected fault and is therefore identified as the most likely faulty variable. Unlike classical RBC methods developed for deterministic linear PCA models, the proposed approach integrates interval-based fault detection, explicitly accounting for measurement uncertainty and nonlinear behavior, which makes it particularly suitable for fault localization in uncertain and complex monitoring environments.
Localization Method Steps:
Step 1:
Initial Fault Detection
Step 2:
Generation of Sub-Models for Diagnosis:
From the global model (which includes all process variables), an ensemble of m sub-models is generated; each sub-model is created by removing one variable j from the set of monitored variables.
Step 3:
Recalculation of the Detection Index for Each Sub-Model
For each sub-model (after removing variable j), a new detection index is computed.
Step 4:
Calculation of the Diagnostic Ratio
For each sub-model, a ratio is calculated between the new detection index and its corresponding threshold; this ratio reflects how much the detection index changes when variable j is removed.
Step 5:
Faulty Variable Identification
Among all the computed ratios, the smallest one is identified; the variable associated with the minimum ratio is considered the faulty variable.

3.2. Mathematical Formulation

Stork et al. [42] proposed an approach for fault localization. The philosophy behind this method is similar to that of Dunia et al. [43].
The elimination approach is similar to the reconstruction approach [43]. A set of detection indices is generated by extracting at each instant a variable from the set of the variables to be monitored and by modifying the matrix representing the PCA model by elimination of the column corresponding to the eliminated variable. The relationship between the value indices at the instant of detection and their thresholds are calculated. The ratio calculated after eliminating a variable with a value less than 1 indicates that this variable is the wrong variable.
In the event that a fault is detected with the statistical D i C R j , a set of PCA models is generated from the overall model by eliminating each time a variable from the set of variables to be monitored. The quantity is defined by Q r after the elimination of the jth variable. The defective variable is the jth variable, which corresponds to the smallest Q r .
Q r C R = D i C R j S e u i l D i C R j ,
With
D i C R j ( t ) = k r C R ( Υ C R j , Υ C R j ) k r C R T ( Υ C R j ) C R C R T k r C R T ( Υ C R j )
And the denominator is the corresponding detection threshold associated with the reduced model.
QrCR is calculated for each of the 18 reduced models, where each model is obtained by eliminating one variable at a time from the original monitoring model. The diagnostic ratio represents the normalized value of the detection index after variable elimination.
When a non-faulty variable is eliminated, the fault-related information remains in the reduced model, and the detection index stays close to or above its threshold, resulting in a ratio close to or greater than unity. In contrast, eliminating the faulty variable removes the primary source of abnormal variation from the model. This leads to a significant decrease in the detection index and, consequently, to a smaller diagnostic ratio.
Therefore, the variable associated with the minimum ratio corresponds to the variable whose removal most effectively suppresses the fault signature and is identified as the source of the fault.
In the case of the upper- and lower-limit approach, the ratio is in its interval form with Q r = Q ¯ r , Q ¯ r
Q ¯ r = D ¯ i j S e u i l D ¯ i j ,   Q ¯ r = D ¯ i j S e u i l D ¯ i j ,
D ¯ i j ( k ) = k ( x ¯ r , k j , x ¯ r , k j ) k T ( x ¯ r , k j ) C ¯ r k ( x ¯ r , k j ) ,
D ¯ i j ( k ) = k ( x ¯ r , k j , x ¯ r , k j ) k T ( x ¯ r , k j ) C ¯ r k ( x ¯ r , k j ) ,
In the proposed RR-KPCA-based fault detection framework, a diagnostic ratio value below 1 is used to indicate a potentially faulty variable. This threshold arises from the normalized formulation of the indicator, where values near 1 represent the boundary between nominal system behavior and statistically significant deviations from the learned reference model. In practical air quality monitoring systems, measurement noise and natural variability can cause short-term fluctuations around this threshold. Therefore, fault decisions are based not on a single instantaneous reading but on the temporal behavior of the indicator, such as repeated exceedances or persistent deviation below the threshold. This approach ensures that the diagnostic rule remains robust against transient noise while still effectively identifying genuinely faulty sensor measurements.

4. Application to the AIRLOR Air Quality Monitoring Network

AIRLOR Description

The air quality monitoring network AIRLOR, operating in Lorraine, France, is a comprehensive system that accumulates, examines, and validates air quality data across a wide geographical area. It plays a decisive role in ecological surveillance, sensor reliability, and supporting both research and policymaking in air quality management [8]. The details of this application are resumed in Table 1.
Although the AIRLOR monitoring network consists of 20 stations measuring several atmospheric pollutants, the observation vector used in this study was limited to 18 variables corresponding to synchronized measurements of O3, NO, and NO2 collected from a selected subset of monitoring stations. The selection was based on data completeness and temporal consistency over the analyzed period. Stations containing missing measurements, inconsistent sampling intervals, or incomplete historical records were excluded to ensure the reliability and consistency of the RR-IKPCA model training process. Consequently, the retained variables represent the subset of measurements providing stable multivariate correlations suitable for fault detection analysis.
Only six AIRLOR stations provide additional meteorological measurements. In the proposed framework, meteorological effects are not explicitly introduced as independent model inputs for all stations; instead, their influence is indirectly reflected through the correlation structure learned from the multivariate pollutant measurements. Furthermore, environmental deployment characteristics such as sensor height above ground level and proximity to emission sources may influence pollutant variability and inter-variable correlations. These spatial metadata were not explicitly incorporated into the present model and therefore constitute a limitation of the current study. Accordingly, the introduced 2% uncertainty should be interpreted as an assumed measurement uncertainty level used for robustness evaluation rather than a complete representation of all environmental and deployment-related variability.
Table 2. Different faults.
Table 2. Different faults.
FaultsAdditive DefectStationObservations
Fault 120% of the ordinary variation of (NO)Station 1Between 400 and 600
Fault 230% of the ordinary variation of (NO)Station 1Between 400 and 700
Fault 320% of the ordinary variation of (NO2)Station 4Between 300 and 600
Fault 440% of the ordinary variation of (O3)Station 6Between 400 and 600

5. Discussion

5.1. Results and Discussion

This section evaluates the fault detection performance of four RR-IKPCA-based monitoring strategies—RRIKPCA_UL using SPE and [D4], and RRIKPCA_CR using SPE and the proposed sensitive index D 4 C R —under four different fault scenarios. The evaluation is conducted using three quantitative indicators: false alarm rate (FAR), good detection rate (GDR), and average computation time (CT).
The first part of the analysis focuses on the fault detection capability of the different indices. As shown in Figure 3, Figure 4, Figure 5 and Figure 6, methods based on the proposed index consistently outperform those relying on the classical SPE statistic. This behavior can be explained by the fact that SPE is computed in the residual subspace and may dilute fault information when uncertainty and nonlinear effects are present, especially under interval modeling. In contrast, the proposed index exploits the most sensitive residual components, making it more responsive to fault-induced variations.
Among all tested approaches, the RRIKPCA_CR method based on D 4 C R achieves the best overall performance. It maintains a 100% GDR for all fault scenarios, while simultaneously achieving the lowest FAR. In addition, its computation time remains low. This improved efficiency is attributed to the center–range (CR) interval representation, which preserves essential uncertainty information while reducing redundant computations compared to the upper–lower (UL) formulation.
Although the RRIKPCA_UL approach based on [D4] also attains a perfect GDR, its computation time is slightly higher in several scenarios, indicating that the CR formulation offers a better compromise between detection accuracy and computational efficiency. In contrast, SPE-based methods exhibit higher FAR values and reduced detection reliability in certain fault conditions, highlighting their limited sensitivity in uncertain and nonlinear environments.
This difference is particularly evident in the case of fault 4 (Figure 6). Despite the introduction of a severe fault corresponding to a 40% increase in the standard deviation of the O3 signal, the SPE-based methods fail to provide satisfactory detection, as reflected by reduced GDR values in both UL and CR configurations. Such detection uncertainty is unacceptable in air quality monitoring applications, where delayed or missed detection may have direct health implications. Conversely, the proposed D 4 C R index achieves a perfect GDR of 100%, confirming its superior fault sensitivity and robustness.
Table 3 presents a comparative study between the [SPE] and [D] indices in order to determine the most appropriate choice for fault detection, which consequently leads to better fault isolation. It is clear that the proposed index outperforms the others due to the superior results obtained.
Furthermore, this table presents a comparative analysis of different fault detection and localization methods under two operating conditions: ideal measurements without uncertainty (0%) and noisy measurements with 2% uncertainty. The comparison was conducted to evaluate the sensitivity and robustness of each method when measurement uncertainty is introduced. As shown in the table, the addition of uncertainty generally affects the fault attribution and detection performance of the conventional approaches, leading to higher false alarm rates and reduced localization reliability. In contrast, the proposed RR-IKPCA-based framework maintains more stable performance with comparatively lower FAR values and clearer fault isolation capability under uncertain conditions. These results demonstrate that the proposed method is less sensitive to measurement variations and remains effective in realistic air quality monitoring environments where sensor noise and uncertainty are unavoidable.
All experiments were conducted on a MSI GF63 Thin 10SC laptop equipped with an Intel® Core™ i7-10750H processor operating at 2.60 GHz, 16 GB RAM, and Intel® UHD Graphics with 8 GB shared memory. The implementation and simulations were performed in a Windows 10 Professional 64-bit environment with DirectX 12 support. The reported computation time (CT) values were obtained under this hardware and software configuration to ensure reproducibility and provide a representative evaluation of the computational performance of the proposed method.
To summarize, the simulation results clearly demonstrate that combining the elimination-based localization principle with the proposed sensitive detection index significantly enhances fault detection and localization performance. Compared with the classical SPE-based approach, the proposed method provides clearer fault attribution, lower false alarm rates, and higher detection reliability under interval uncertainty and nonlinear system behavior. The reduction in diagnostic ambiguity is particularly evident when the faulty variable is eliminated, which leads to a pronounced decrease in the detection index and enables unambiguous localization. These results confirm that the proposed index is better suited than SPE for elimination-based localization and highlight the effectiveness of the RR-IKPCA_CR framework in uncertain air quality monitoring applications.
Figure 7 and Figure 8 show the fault localization results by elimination for fault (d7), using the two models: RR-IKPCA-UL with the index [D4] and RR-IKPCA- CR with the index D4CR, respectively. The evolution of different D 4 j and D 4 C R j values in the case of fault (d7) shows that eliminating variable v7 leads to the indices D 4 7 and D 4 7 , which no longer indicate the presence of the fault. Figure 9 and Figure 10 present the Qr ratio for each of the 18 models. These figures show that the smallest ratio corresponds to variable v7, indicating that this variable is the source of the fault.
To summarize, this approach aims to locate the faulty variable by analyzing how the detection index behaves when each variable is individually removed. The use of RR-IKPCA and interval analysis makes the fault diagnosis more robust and reliable, especially under uncertainty.

5.2. Comparative Analysis

To position the proposed method with respect to existing approaches, Table 4 shows a comparison between commonly used fault localization techniques in terms of model type, localization principle, uncertainty handling, and limitations.
In practical industrial systems, the assumption of a single-fault occurrence may not always hold, as multiple faults can arise simultaneously under complex operating conditions. In such scenarios, the separation and identification of individual fault signatures become more challenging, potentially affecting the localization accuracy of the proposed RR-IKPCA framework. Although the current study focuses on single-fault scenarios to establish the effectiveness of the method, we acknowledge that multi-fault conditions represent a more realistic and demanding case. The behavior of the proposed approach under simultaneous faults is therefore an important consideration, and its detailed investigation, including performance evaluation and possible methodological extensions to handle fault overlap, is identified as a relevant direction for future research.

6. Conclusions

This paper proposes a fault detection and localization method for diagnosing sensor failures in an air monitoring network by combining interval-based fault detection techniques with the localization-by-elimination principle. The effectiveness of the proposed approach is demonstrated through its application to fault diagnosis in an air quality monitoring network. Quantitative results obtained from the AIRLOR air quality monitoring network demonstrate the effectiveness of the proposed method. The RRIKPCA_CR-based localization scheme achieved a good detection rate (GDR) across all tested fault scenarios, while maintaining a low false alarm rate (FAR).
Despite the promising results, the proposed framework has several limitations. It assumes that faults occur independently and relies on the availability of sufficient historical data for effective model training. Furthermore, the current implementation has not yet been validated under simultaneous multiple-fault conditions or highly dynamic operating environments. Future work will therefore focus on extending the framework to handle multiple concurrent faults, time-varying process dynamics, and adaptive or online learning schemes. In addition, applying the proposed approach to other domains such as water quality monitoring and industrial process control will further evaluate its generality and scalability.

Author Contributions

Conceptualization, I.H.; formal analysis, I.H. and H.L.; methodology, O.T., H.L. and A.A.; software, I.H.; validation, O.T., A.A. and H.L.; writing—original draft preparation, I.H.; writing—review and editing, H.L., E.A. and O.T.; visualization, I.H. and E.A.; supervision, O.T., A.A. and E.A.; project administration, O.T., A.A. and E.A. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding author.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
[X]Interval data matrix
LBLower bound
UBUpper bound
KKernel matrix
SPESquared Prediction Error
AMean of the fault detection index
BVariance of the fault detection index
HFeature space
τ i , α 2 Detection threshold
Γ Transformation function

References

  1. Isermann, R. Fault-Diagnosis Systems: An Introduction from Fault Detection to Fault Tolerance; Springer Science & Business Media: London, UK, 2005. [Google Scholar]
  2. Chiang, L.H.; Russell, E.L.; Braatz, R.D. Fault Detection and Diagnosis in Industrial Systems; Springer Science and Business Media LLC: London, UK, 2001. [Google Scholar]
  3. Blanke, M.; Kinnaert, M.; Lunze, J.; Staroswiec, M. Diagnosis et Fault-Tolerant Control; Springer: Berlin/Heidelberg, Germany, 2003. [Google Scholar]
  4. Zhu, Y.; Zhao, S.; Zhang, Y.; Zhang, C.; Wu, J. A Review of Statistical-Based Fault Detection and Diagnosis with Probabilistic Models. Symmetry 2024, 16, 455. [Google Scholar] [CrossRef] [Scilit]
  5. Jolliffe, I.T. Principal Component Analysis; Springer: Berlin/Heidelberg, Germany, 2002. [Google Scholar]
  6. Wang, J.; Zhou, Z.; Li, Z.; Du, S. A novel fault detection scheme based on mutual k-nearest neighbor method: Application on the industrial processes with outliers. Processes 2022, 10, 497. [Google Scholar] [CrossRef] [Scilit]
  7. Janssen, S.; Dumont, G.; Fierens, F.; Deutsch, F.; Maiheu, B.; Celis, D.; Trimpeneers, E.; Mensink, C. Land use to characterize spatial representativeness of air quality monitoring stations and its relevance for model validation. Atmos. Environ. 2012, 59, 492–500. [Google Scholar] [CrossRef] [Scilit]
  8. Harkat, M.-F.; Mansouri, M.; Nounou, M.; Nounou, H. Enhanced data validation strategy of air quality monitoring network. Environ. Res. 2018, 160, 183–194. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. Fazai, R.; Abdellafou, K.B.; Said, M. Online fault detection and isolation of an AIR quality monitoring network based on machine learning and metaheuristic methods. Int. J. Adv. Manuf. Technol. 2018, 99, 2789–2802. [Google Scholar] [CrossRef] [Scilit]
  10. Harkat, M.-F.; Mourot, G.; Ragot, J. An improved PCA scheme for sensor FDI: Application to an air quality monitoring network. J. Process Control 2006, 16, 625–634. [Google Scholar] [CrossRef] [Scilit]
  11. Ge, Z.; Song, Z.; Gao, F. Review of recent research on data-based process monitoring. Ind. Eng. Chem. Res. 2013, 52, 3543–3562. [Google Scholar] [CrossRef] [Scilit]
  12. Xie, Z. Identification of sensor fault characteristics based on adaptive kernel principal component analysis. Eur. J. Eng. Technol. 2025, 1, 7–15. [Google Scholar]
  13. Hamrouni, I.; Lahdhiri, H.; Ben Abdellafou, K.; Aljuhani, A.; Taouali, O.; Bouzrara, K. Anomaly Detection and Localization for Process Security Based on the Multivariate Statistical Method. Math. Probl. Eng. 2022, 2022, 5580774. [Google Scholar] [CrossRef] [Scilit]
  14. Samuel, R.T.; Cao, Y. Nonlinear process fault detection and identification using kernel PCA and kernel density estimation. Syst. Sci. Control Eng. 2016, 4, 165–174. [Google Scholar] [CrossRef] [Scilit]
  15. Gharahbagheri, H.; Imtiaz, S.A.; Khan, F. Root cause diagnosis of process fault using KPCA and Bayesian network. Ind. Eng. Chem. Res. 2017, 56, 2054–2070. [Google Scholar] [CrossRef] [Scilit]
  16. Wang, Q.; Liu, Y.B.; He, X.; Liu, S.Y.; Liu, J.H. Fault diagnosis of bearing based on KPCA and KNN method. Adv. Mater. Res. 2014, 986–987, 1491–1496. [Google Scholar] [CrossRef] [Scilit]
  17. Zheng, D.; Zhou, L.; Song, Z. Kernel generalization of multi-rate probabilistic principal component analysis for fault detection in nonlinear process. IEEE/CAA J. Autom. Sin. 2021, 8, 1465–1476. [Google Scholar] [CrossRef] [Scilit]
  18. Zhang, Q.; Li, P.; Lang, X.; Miao, A. Improved dynamic kernel principal component analysis for fault detection. Measurement 2020, 158, 107738. [Google Scholar] [CrossRef] [Scilit]
  19. Benkouider, M.S.B. Sensor fault detection using a new online reduced rank kernel PCA for monitoring an air quality monitoring network. Control Eng. Pract. 2020, 103, 104543. [Google Scholar]
  20. Box, G.E.P. Some theorems on quadratic forms applied in the study of analysis of variance problems, I. effect of inequality of variance in the one-way classification. Ann. Math. Stat. 1954, 25, 290–302. [Google Scholar] [CrossRef] [Scilit]
  21. Ait-Izem, T.; Harkat, M.-F.; Djeghaba, M.; Kratz, F. On the application of interval PCA to process monitoring: A robust strategy for sensor FDI with new efficient control statistics. J. Process Control 2018, 63, 29–46. [Google Scholar] [CrossRef] [Scilit]
  22. Cazes, P.; Chouakria, A.; Diday, E.; Schektman, Y. Extension de l’Analyse en Composantes principales à des données de Type Intervalle. Rev. Stat. Appliquée 1997, 45, 5–24. [Google Scholar]
  23. Chouakria, A. Extension des Méthodes d’analyse Factorielle à des Données de Type Intervalle. Ph.D. Thesis, Université des Sciences et Technologies de Lille, Lille, France, 1998. [Google Scholar]
  24. Mansouri, M.; Harkat, M.-F.; Nounou, M.; Nounou, H. Midpoint-radii principal component analysis-based EWMA and application to air quality monitoring network. Chemom. Intell. Lab. Syst. 2018, 175, 55–64. [Google Scholar] [CrossRef] [Scilit]
  25. Le-Rademacher, J.; Billard, L. Symbolic Covariance Principal Component Analysis and Visualization for Interval Valued Data. J. Comput. Graph. Stat. 2012, 21, 413–432. [Google Scholar] [CrossRef] [Scilit]
  26. Chakour, C.; Benyounes, A.; Boudiaf, M. Diagnosis of uncertain nonlinear systems using interval kernel principal components analysis: Application to a weather station. ISA Trans. 2018, 83, 126–141. [Google Scholar] [CrossRef] [Scilit]
  27. Hamrouni, I.; Lahdhiri, H.; Abdellafou, K.; Taouali, O. Fault detection of uncertain nonlinear process using Reduced Interval Kernel Principal Component Analysis (RIKPCA). Int. J. Adv. Manuf. Technol. 2020, 106, 4567–4576. [Google Scholar] [CrossRef] [Scilit]
  28. Hamrouni, I.; Abdellafou, K.B.; Aborokbah, M.; Taouali, O. Dynamic sensor fault detection approach using data-driven techniques. Neural Comput. Appl. 2024, 36, 14291–14307. [Google Scholar] [CrossRef] [Scilit]
  29. Lahdhiri, H.; Taouali, O. Interval valued data driven approach for sensor fault detection of nonlinear uncertain process. Measurement 2021, 171, 108776. [Google Scholar] [CrossRef] [Scilit]
  30. Harkat, M.; Mansouri, M.; Nounou, M.; Nounou, H. Fault detection of uncertain chemical processes using interval partial least squares-based generalized likelihood ratio test. Inf. Sci. 2019, 490, 265–284. [Google Scholar] [CrossRef] [Scilit]
  31. Alcala, C.F. Fault Diagnosis with Reconstruction Based Contributions for Statistical Process Monitoring. Ph.D. Thesis, University of Southern California (USC), Los Angeles, CA, USA, 2011. [Google Scholar]
  32. Harakat, M.F. Détection et Localisation de Défauts par Analyse en Composantes Principales. Ph.D. Thesis, National Polytechnic Institute of Lorraine, Nancy, France, 2003. [Google Scholar]
  33. Brunet, J.; Labarrère, M.; Jaume, D.; Rault, A.; Vergé, M. Détection et Diagnostic de Pannes: Approche par Modélisation; Hermès: Paris, France, 1990. [Google Scholar]
  34. Jafari, N.; Lopes, A.M. Fault Detection and Identification with Kernel Principal Component Analysis and Long Short-Term Memory Artificial Neural Network Combined Method. Axioms 2023, 12, 583. [Google Scholar] [CrossRef] [Scilit]
  35. Alcala, C.F.; Qin, S.J. Reconstruction-based contribution for process monitoring with kernel principal component analysis. Ind. Eng. Chem. Res. 2010, 49, 7849–7857. [Google Scholar] [CrossRef] [Scilit]
  36. Shahzad, F.; Huang, Z.; Memon, W.H. Process Monitoring Using Kernel PCA and Kernel Density Estimation-Based SSGLR Method for Nonlinear Fault Detection. Appl. Sci. 2022, 12, 2981. [Google Scholar] [CrossRef] [Scilit]
  37. Wang, H.; Peng, M.-J.; Yu, Y.; Saeed, H.; Hao, C.-M.; Liu, Y.-K. Fault identification and diagnosis based on KPCA and similarity clustering for nuclear power plants. Ann. Nucl. Energy 2021, 150, 107786. [Google Scholar] [CrossRef] [Scilit]
  38. Fezai, R.; Mansouri, M.; Trabelsi, M.; Hajji, M.; Nounou, H.; Nounou, M. Online reduced kernel GLRT technique for improved fault detection in photovoltaic systems. Energy 2019, 179, 1133–1154. [Google Scholar] [CrossRef] [Scilit]
  39. Lahdhiri, H.; Aljuhani, A.; Ben Abdellafou, K.; Taouali, O. An Improved Fault Diagnosis Strategy for Process Monitoring Using Reconstruction Based Contributions. IEEE Access 2021, 9, 79520–79533. [Google Scholar] [CrossRef] [Scilit]
  40. Zeng, L.; Long, W.; Li, Y. A Novel Method for Gas Turbine Condition Monitoring Based on KPCA and Analysis of Statistics T2 and SPE. Processes 2019, 7, 124. [Google Scholar] [CrossRef] [Scilit]
  41. Mansouri, M.; Nounou, M.; Nounou, H.; Karim, N. Kernel PCA-based GLRT for nonlinear fault detection of chemical processes. J. Loss Prev. Process Ind. 2016, 40, 334–347. [Google Scholar] [CrossRef] [Scilit]
  42. Stork, C.L.; Veltkamp, D.J.; Kowalski, B.R. Identification of multiple sensor disturbances during process monitoring. Anal. Chem. 1997, 69, 5031–5036. [Google Scholar] [CrossRef] [Scilit]
  43. Dunia, R.; Qin, S.J.; Edgar, T.F.; McAvoy, T.J. Identification of faulty sensors using principal component analysis. AIChE J. 1996, 42, 2797–2812. [Google Scholar] [CrossRef] [Scilit]
  44. Wang, Z.; Cheng, J.; Liu, W.; Zou, X.; Tao, F. A fault localization approach based on multi-system PCA and dynamic SDG: Application in train lifting equipment. Robot. Comput.-Integr. Manuf. 2024, 2024, 102694. [Google Scholar] [CrossRef] [Scilit]
  45. Zhou, W.; Yang, W.; Wang, Y.; Zhang, H. Generalized reconstruction-based contribution for multiple faults diagnosis with bayesian decision. In Proceedings of the 2018 IEEE 7th Data Driven Control and Learning Systems Conference (DDCLS), Enshi, China, 25–27 May 2018; pp. 813–818. [Google Scholar]
  46. Ziane, A.; Mekhilef, S.; Dabou, R.; Sahouane, N.; Necaibia, A.; Rouabhia, A.; Khelifi, S.; Lachter, S.; Mostefaoui, M.; Larbi, A.A.; et al. Online fault detection in grid-connected PV system using nonlinear multivariate statistical analysis. Electr. Eng. 2026, 108, 107. [Google Scholar] [CrossRef] [Scilit]
Figure 1. RR-IKPCA based on interval midpoint–radii.
Figure 1. RR-IKPCA based on interval midpoint–radii.
Applsci 16 05674 g001
Figure 2. RR-IKPCA upper–lower bound.
Figure 2. RR-IKPCA upper–lower bound.
Applsci 16 05674 g002
Figure 3. Fault detection performance for RR-IKPCA based on [SPE] and [D]; fault 1.
Figure 3. Fault detection performance for RR-IKPCA based on [SPE] and [D]; fault 1.
Applsci 16 05674 g003
Figure 4. Fault detection performance for RR-IKPCA based on [SPE] and [D]; fault 2.
Figure 4. Fault detection performance for RR-IKPCA based on [SPE] and [D]; fault 2.
Applsci 16 05674 g004
Figure 5. Fault detection performance for RR-IKPCA based on [SPE] and [D]; fault 3.
Figure 5. Fault detection performance for RR-IKPCA based on [SPE] and [D]; fault 3.
Applsci 16 05674 g005
Figure 6. Fault detection performance for RR-IKPCA based on [SPE] and [D4]; fault 4.
Figure 6. Fault detection performance for RR-IKPCA based on [SPE] and [D4]; fault 4.
Applsci 16 05674 g006
Figure 7. Time evolution of D 4 obtained using the elimination method based on RR−IKPCA−UL model with fault 4.
Figure 7. Time evolution of D 4 obtained using the elimination method based on RR−IKPCA−UL model with fault 4.
Applsci 16 05674 g007
Figure 8. Time evolution of D 4 C R obtained using the elimination method based on RR-IKPCA-CR model with fault 4.
Figure 8. Time evolution of D 4 C R obtained using the elimination method based on RR-IKPCA-CR model with fault 4.
Applsci 16 05674 g008
Figure 9. Fault location using report Q r using the elimination method based on RR-IKPCA-UL model.
Figure 9. Fault location using report Q r using the elimination method based on RR-IKPCA-UL model.
Applsci 16 05674 g009
Figure 10. Fault location using interval report Q r using the elimination method based on RR-IKPCA-UL model.
Figure 10. Fault location using interval report Q r using the elimination method based on RR-IKPCA-UL model.
Applsci 16 05674 g010
Table 1. Description of air quality monitoring network AIRLOR.
Table 1. Description of air quality monitoring network AIRLOR.
CategoryDetails
Number of Stations20
Station LocationsRural, peri-rural, urban
Pollutants MonitoredNO, NO2, O3, CO, SO2
Stations with Meteorological Data6 stations
Primary ObjectiveDetect sensor faults measuring O3, NO, and NO2 concentrations
Observation Vector18 controlled variables (O3, NO, NO2 concentrations at each station)
x ( k ) = v 1 ( k )   v 2 ( k )   v 3 ( k ) O 3 ( k )   NO ( k )   NO 2 ( k ) : station 1 v 10 ( k )   v 11 ( k )   v 12 ( k ) O 3 ( k )   NO ( k )   NO 2 ( k ) : station 4 v 16 ( k )   v 17 ( k )   v 18 ( k ) O 3 ( k )   NO ( k )   NO 2 ( k ) : station 6 T
Training Data400 observations
Testing Data1000 observations
Uncertainty Introduced2% of the range of each variable (added to data for robustness testing)
Fault Simulation4 types of faults (described in Table 2; used to test the effectiveness of the proposed RR-IKPCA method)
Model UsedRR-IKPCA (reference model constructed using training data)
Table 3. Fault detection performances of RR-IKPCA based on [SPE] and [D4].
Table 3. Fault detection performances of RR-IKPCA based on [SPE] and [D4].
MethodsFault 1Fault 2Fault 3Fault 4
FAR
95%
GDR
95%
CT(s)FAR
95%
GDR
95%
CT(s)FAR
95%
GDR
95%
CT(s)FAR
95%
GDR
95%
CT(s)
RRKPCA based on SPE5744.53.88894.18905.234744.89
RRKPCA based on D43.8904.142.299122.59652.5804.66
RRIKPCA_UL based on [SPE]6.894.54.896.5814.892.51007.952.5764.99
RRIKPCA_UL based on [D4]21006.8521006.854.251005.0321004.66
RRIKPCA_CR based on [SPE]8.01954.145837.238.981008.984.2583.64.99
RRIKPCA_CR based on [D4]2.41007.232.41004.147.891007.432.141007.96
Table 4. Comparative analysis.
Table 4. Comparative analysis.
MethodModel TypeLocalization PrincipleUncertainty HandlingStrengthsMain LimitationsDetection RateReferences
Partial contribution Linear PCAContribution of each variable to (T2) or SPENoSimple to implement; intuitive interpretationSmearing effect; poor localization for correlated variables84.5[44]
Reconstruction-based contribution (RBC)Linear PCA + Nonlinear (KPCA)Variable-wise reconstruction errorNoImproved localization compared with contribution plotsSensitive to noise; linear assumption; no uncertainty modeling89.7[35,39,45,46]
Partial reconstruction (PR)Linear PCAReconstruction of selected variablesNoBetter discrimination than classical RBCIncreased computational burden; linear assumption91.2[32]
Proposed RRIKPCA + eliminationNonlinear (IKPCA)Elimination-based diagnostic ratio using sensitive residual indexYes (interval-based)High fault sensitivity; robust to uncertainty; clear localizationAssumes single dominant fault; higher complexity for large-scale systems97.8This work
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Hamrouni, I.; Lahdhiri, H.; Taouali, O.; Alshehri, A.; Aloufi, E. Data-Driven Technique for Fault Detection and Localization of Air Quality Process. Appl. Sci. 2026, 16, 5674. https://doi.org/10.3390/app16115674

AMA Style

Hamrouni I, Lahdhiri H, Taouali O, Alshehri A, Aloufi E. Data-Driven Technique for Fault Detection and Localization of Air Quality Process. Applied Sciences. 2026; 16(11):5674. https://doi.org/10.3390/app16115674

Chicago/Turabian Style

Hamrouni, Imen, Hajer Lahdhiri, Okba Taouali, Ali Alshehri, and Esam Aloufi. 2026. "Data-Driven Technique for Fault Detection and Localization of Air Quality Process" Applied Sciences 16, no. 11: 5674. https://doi.org/10.3390/app16115674

APA Style

Hamrouni, I., Lahdhiri, H., Taouali, O., Alshehri, A., & Aloufi, E. (2026). Data-Driven Technique for Fault Detection and Localization of Air Quality Process. Applied Sciences, 16(11), 5674. https://doi.org/10.3390/app16115674

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop