Next Article in Journal
Detection of Sweat-Related Metabolites (Glucose, Lactic Acid, and Urea) Using a SWCNT-Modified Gold Screen Printed Electrode Based Biosensor
Next Article in Special Issue
Multi-Criteria Optimization of Production Processes of Mining Companies
Previous Article in Journal
Wastewater Treatment Plants as Environmental Barriers in Hyperarid Regions: A Comprehensive Evaluation of Their Performance, Groundwater Protection, and Reuse in Agriculture in the Algerian Sahara
Previous Article in Special Issue
Configuration Optimization of Lazy-Wave Dynamic Umbilicals Using Random Forest Surrogates and NSGA-II
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Field Implementation of an Expert System for Energy Efficiency Improvement in Industrial Aluminum Electrolysis: A Rule-Based Explainable AI Approach

1
School of New Energy and Aviation, Jiangsu College of Engineering and Technology, Nantong 226000, China
2
School of Mechanical and Automotive Engineering, Shanghai University of Engineering Science, Shanghai 201620, China
*
Author to whom correspondence should be addressed.
Processes 2026, 14(7), 1113; https://doi.org/10.3390/pr14071113
Submission received: 12 February 2026 / Revised: 23 March 2026 / Accepted: 24 March 2026 / Published: 30 March 2026

Abstract

This study reports a field-validated expert system for 230 kA aluminum reduction cells at Chalco Guizhou Branch, achieving sustained energy savings of 137 kWh/t (annual comprehensive benefits of 333,714 CNY/cell) over six consecutive months. Addressing “black-box” AI limitations, the system employs hybrid knowledge representation (production rules, frame-based structures, certainty factors) within an XAI framework. The four-layer architecture integrates OPC UA/Modbus TCP protocols for real-time data acquisition and interpretable diagnosis. Field trials demonstrated 94.2% diagnostic accuracy, significantly outperforming manual diagnosis (87.6%, p < 0.001) while achieving comparable performance to LSTM deep learning (93.8%, p = 0.42), with 15× faster inference speed (3.5 s vs. 52 s). Industrial implementation increased current efficiency by 0.7%, reduced DC power consumption by 137 kW·h/t, and decreased anode effect frequency by 32.5%. The system’s explicit reasoning capability provides transparent diagnostic explanations, bridging the gap between data-driven AI and domain expertise for trustworthy intelligent diagnostics in energy-intensive industrial processes.

1. Introduction

The primary aluminum industry represents one of the most energy-intensive sectors in modern manufacturing, with global production exceeding 69 million tons annually and accounting for approximately 3% of the world’s total electricity consumption [1]. The Hall–Héroult electrolytic process, which remains the dominant method for primary aluminum production, operates at temperatures exceeding 950 °C and requires direct current energy consumption typically ranging from 12,500 to 14,000 kWh per ton of aluminum produced [2]. In China, the world’s largest aluminum producer, the electrolytic aluminum industry’s annual electricity consumption accounts for approximately 6.5% of the nation’s total societal electricity usage [3]. The Hall–Héroult energy efficiency optimization of aluminum reduction cells has become a critical challenge for the sustainable development of the industry [4,5].
The operational stability of aluminum electrolysis cells directly determines production efficiency, product quality, and energy consumption levels. However, the electrolysis process is extremely complex. It involves intricate interactions among electromagnetic fields, thermal fields, flow fields, and chemical reactions, all occurring within high-temperature and strong magnetic field environments [6]. Large-scale cells, particularly 230 kA and above, present even greater operational challenges due to their increased thermal inertia, enhanced magnetohydrodynamic instability, and more complex side ledge dynamics [7]. Traditional manual diagnosis approaches rely heavily on operator experience and periodic inspections, which suffer from significant time lags and subjective variability. These limitations often result in delayed detection of abnormal conditions, leading to catastrophic failures such as anode effects, electrolyte solidification, or even cell shutdowns that can cause substantial economic losses and safety hazards [8].
In response to these challenges, significant research efforts have been directed toward developing automated monitoring and intelligent diagnostic systems for aluminum electrolysis processes. Early approaches focused primarily on data acquisition and visualization, establishing supervisory control and data acquisition (SCADA) systems to collect real-time parameters including cell voltage, series current, electrolyte temperature, and alumina concentration [9]. While these systems provide comprehensive data logging capabilities, they essentially function as passive monitoring tools without intelligent diagnostic capabilities. Subsequently, various artificial intelligence techniques have been applied to aluminum electrolysis fault diagnosis, including artificial neural networks for electrolyte temperature prediction [10], support vector machines for anode effect detection [11], and deep learning models based on long short-term memory (LSTM) networks for cell state classification [12]. More recently, hybrid deep learning approaches combining Bi-LSTM, LSTM, and dense neural networks have achieved prediction accuracies exceeding 95% for anode effect forecasting [13]. Additionally, multi-objective optimization methods based on comprehensive index evaluation models have been developed for cell voltage optimization [14,15].
Despite these advances, existing AI-based diagnostic approaches for aluminum electrolysis cells face a fundamental limitation: they predominantly operate as “black-box” models that lack interpretability and transparency [16]. While deep learning models can achieve high prediction accuracy, their complex internal architectures make it virtually impossible for operators to understand the reasoning behind diagnostic decisions [17]. This opacity creates significant barriers to practical industrial adoption, as process engineers and operators require clear explanations to trust and act upon automated recommendations. Furthermore, these data-driven approaches often fail to incorporate valuable domain knowledge accumulated through decades of industrial practice, including expert heuristics, operational standards, and first-principles understanding of electrolysis physics [9]. The absence of a systematic framework for integrating expert knowledge with data-driven insights represents a critical gap in current aluminum electrolysis diagnostic systems.
The emergence of Explainable AI (XAI) in Process Systems Engineering (PSE) offers a promising pathway to address this challenge [18,19]. XAI emphasizes the development of transparent, interpretable models that maintain high performance while providing human-understandable reasoning [20]. Within this paradigm, rule-based expert systems represent a form of “intrinsically interpretable” XAI, where decision logic is explicitly encoded rather than learned opaquely [17,21]. Recent literature has highlighted the regulatory and safety advantages of such approaches in industrial settings, particularly for safety-critical applications requiring audit trails and compliance verification [19,22].
To address these limitations, this paper presents the development of an interpretable expert diagnostic system specifically designed for 230 kA aluminum reduction cells, positioned within the XAI for PSE framework. The proposed system integrates multi-source process data with domain expert knowledge through a hybrid knowledge representation framework, combining rule-based reasoning with certainty factor-based uncertainty handling. Unlike conventional black-box AI models, the developed system provides transparent diagnostic reasoning by explicitly mapping observed process deviations to specific fault modes through interpretable inference chains. The system architecture incorporates three core modules: (1) a real-time data acquisition and preprocessing module that handles heterogeneous sensor data from the cell control system via standardized industrial protocols (OPC UA and Modbus TCP/IP); (2) a knowledge base module that encodes expert diagnostic rules derived from operational experience and process physics, featuring version-controlled rule management for continuous improvement; and (3) an inference engine that performs multi-level diagnostic reasoning with explicit uncertainty quantification to identify abnormal conditions and their root causes.
Unlike prior work focusing on algorithmic novelty [12,13], this study emphasizes rigorous industrial validation of interpretable AI, addressing the last mile deployment gap that has limited practical adoption of advanced AI in aluminum smelting. Our contribution lies not in developing new AI algorithms, but in systematically engineering, deploying, and validating a complete expert diagnostic system in a high-temperature, high-risk industrial environment over six months of continuous operation.
The key contributions of this work are three-fold:
  • Industrial XAI Implementation: We demonstrate a practical deployment of rule-based XAI in a high-temperature, high-risk industrial environment, providing full audibility of diagnostic decisions for regulatory compliance and safety certification.
  • Comparative Performance Analysis: Through rigorous comparison with both manual diagnosis and deep learning approaches (LSTM), we demonstrate that the rule-based system achieves comparable accuracy (94.2% vs. 93.8%) with 15× faster inference speed and complete interpretability, addressing the accuracy–interpretability trade-off in industrial AI.
  • Field-Validated Energy Optimization: The system achieved sustained energy savings of 137 kWh/t-Al (1.0% reduction) over six months of continuous operation, with detailed ROI analysis demonstrating a payback period of 2.5 months.
Case studies conducted on industrial 230 kA cells demonstrate that the proposed system achieves diagnostic accuracy comparable to deep learning approaches while providing explainable recommendations that align with operator expertise. This work contributes a practical framework for bridging the gap between data-driven AI and domain expertise in aluminum electrolysis process monitoring, offering a pathway toward trustworthy intelligent diagnostics for energy-intensive industrial processes.

2. Methodology

This section presents the methodological framework for developing an expert system dedicated to diagnosing 230 kA aluminum electrolysis cells. The proposed methodology encompasses four fundamental components: knowledge representation methods, knowledge base construction process, inference engine mechanism, and system architecture design. Each component is elaborated in detail to ensure the reproducibility and scalability of the developed system.

2.1. Knowledge Representation Methods

Knowledge representation constitutes the foundation of expert system development, directly influencing the efficiency of knowledge acquisition, storage, and inference. In this study, multiple knowledge representation paradigms are employed to capture the complex diagnostic expertise for aluminum electrolysis cells [23].

2.1.1. Production Rules

The primary knowledge representation method adopted in this study was the production rule system, which formalized expert knowledge into a set of conditional statements following the “IF–THEN” structure. Each production rule is formally defined as:
RULE : IF   < condition >   THEN   < conclusion >   [ CF   =   certainty _ factor ]
where RULE denotes the unique identifier for each rule, <condition> represents the premise clause comprising one or more logical conditions, <conclusion> specifies the inferred result or action, and CF indicates the certainty factor ranging from 0 to 1, representing the confidence level of the rule.
The production rules (see S5 in Supplementary Material) for fault diagnosis in 230 kA aluminum electrolysis cells were categorized into three functional groups:
(1)
Cell Condition Evaluation Rules: These rules assessed the overall operational status of the electrolysis cell based on key process parameters. Three evaluation standards were established: “Excellent” (optimal operating conditions), “Good” (acceptable performance with minor deviations), and “Poor” (abnormal conditions requiring intervention).
(2)
Fault Identification Rules: A comprehensive set of 16 fault diagnosis rules was developed to identify specific abnormal conditions, including anode effects, electrolyte contamination, thermal imbalance, and current distribution anomalies. Each rule correlates specific parameter patterns with corresponding fault types.
(3)
Remedial Action Rules: These rules recommend appropriate operational measures based on identified faults, specifying adjustment parameters for process control.

2.1.2. Frame-Based Representation

Complementing the production rule system, frame-based representation was employed to organize structured knowledge about electrolysis cell components and their relationships. A frame is defined as a data structure with the following formal representation:
Frame_Name:
  • {
  • Slot_1: Value_1 (Default: Default_1, Constraint: Constraint_1),
  • Slot_2: Value_2 (Default: Default_2, Constraint: Constraint_2),
  • Slot_n: Value_n (Default: Default_n, Constraint: Constraint_n),
  • Methods: {Method_1, Method_2, …, Method_m},
  • Relations: {Relation_1, Relation_2, …, Relation_k}
  • }
For aluminum electrolysis cell modeling, the frame structure encapsulates hierarchical knowledge including: (1) Cell physical parameters (anode configuration, cathode geometry, busbar arrangement); (2) Process variables (electrolyte composition, temperature distribution, voltage characteristics); (3) Operational states (current efficiency, energy consumption, metal purity); and (4) Historical performance data. The frame-based approach facilitates inheritance mechanisms, allowing specialized frames to inherit attributes from parent frames, thereby reducing knowledge redundancy and enhancing maintainability.

2.1.3. Behavioral Functions and Uncertainty Handling

To enable dynamic knowledge processing and handle the inherent uncertainty in industrial process data (sensor noise, process fluctuations), three behavioral functions were implemented:
(1)
CONCLUDE Function: This function performs certainty-weighted inference and is formally defined as CONCLUDE ( C , P , V , TALLY , CF ) , where C = Context identifier, P = Parameter name, V = Concluded value, TALLY = Accumulated certainty variable, and CF = Certainty factor. The accumulated certainty is computed using:
TALLY new = TALLY old + CF × ( 1 TALLY old )
(2)
CONLIST Function: This function manages conjunctive condition evaluation, aggregating multiple premise conditions through logical AND operations.
(3)
TRANLIST Function: This function handles transitive inference chains, enabling multi-step reasoning across interconnected rules.
Uncertainty Propagation Example:
Consider the diagnosis of “Anode Effect Imminent” (Hypothesis H ) based on two pieces of evidence:
  • E1: Alumina concentration trend decreasing rapidly (CF = 0.8)
  • E2: Cell voltage noise pattern anomaly (CF = 0.6)
Using the conjunctive rule combination:
CF ( H ) = min ( CF ( E 1 ) , CF ( E 2 ) ) × CF ( Rule ) = min ( 0.8,0.6 ) × 0.9 = 0.54
If a second rule provides CF ( H ) = 0.4 from alternative evidence, then the combined certainty is:
CF combined = 0.54 + 0.4 × ( 1 0.54 ) = 0.54 + 0.184 = 0.724
This explicit uncertainty quantification allows operators to understand the confidence level of each diagnosis and decide whether immediate intervention is warranted or further observation is acceptable.

2.2. Knowledge Base Construction Process

The construction of the diagnostic knowledge base followed a systematic methodology integrating domain expertise with historical data analysis, featuring structured knowledge acquisition protocols and version control mechanisms essential for industrial maintenance.

2.2.1. Knowledge Acquisition Methodology

Knowledge acquisition employed a hybrid approach combining expert elicitation and data-driven rule extraction:
(1)
Expert Knowledge Elicitation: Structured interviews and protocol analysis were conducted with experienced aluminum smelting engineers (average experience > 15 years) to capture heuristic diagnostic knowledge. The acquired knowledge includes qualitative relationships between process parameters and cell conditions, threshold values for abnormal condition detection, prioritized fault diagnosis procedures, and remedial action strategies under various scenarios. The elicitation process followed the KESS (Knowledge Elicitation for Safety-critical Systems) protocol, involving three domain experts independently and resolving conflicts through consensus meetings.
(2)
Historical Data Analysis: Operational data from 230 kA electrolysis cells covering 24 months of operation were analyzed using statistical and machine learning techniques to identify patterns and correlations. The data mining process involves parameter correlation analysis (Pearson correlation coefficients), anomaly detection using control chart methods, association rule mining (Apriori algorithm) for fault–parameter relationships, and temporal pattern analysis for fault progression modeling.

2.2.2. Rule Extraction and Formalization

The extracted knowledge was formalized into production rules through the following procedure (see Figure 1):
Step 1—Data Preprocessing: Raw operational data were filtered using Savitzky–Golay smoothing (window size 11) to remove measurement noise; process parameters were normalized to standard ranges; time series data were segmented into operational periods.
Step 2—Pattern Recognition: Abnormal operational patterns were identified through deviation analysis; correlation patterns between multiple parameters were established; fault signatures were characterized using feature extraction techniques (wavelet packet decomposition for voltage signals).
Step 3—Rule Generation: Statistical thresholds were determined for condition evaluation using the 3-sigma rule on historical normal operation data; logical conditions are formulated based on expert-defined criteria; certainty factors were assigned based on historical rule accuracy (precision) and expert confidence ratings.
When expert opinions differ on certainty factor assignments, the following objective criteria are applied: (1) Historical accuracy weighting: CF is set proportional to the rule historical precision from data analysis (CFinitial = 0.5 + 0.5 × Precisionhistorical); (2) Expert confidence voting: When three experts disagree, the median CF value is adopted; (3) Conservative bias: For safety-critical rules, the minimum proposed CF is selected; (4) Empirical calibration: Initial CF values are refined during the 5-month calibration period based on validation performance. This protocol ensures objective, data-informed CF assignments while respecting expert domain knowledge.
Step 4—Rule Validation: Generated rules were validated against independent test datasets (6-month hold-out period); rule performance metrics (precision, recall, F1-score) were computed; rules with insufficient accuracy (F1 < 0.75) were refined or eliminated.
The 30-month dataset (January 2022–June 2024) was chronologically split: first 18 months for rule extraction and knowledge base development, next 6 months for certainty factor calibration and validation, and final 6 months (January–June 2024) as the hold-out test set. The 672 test cases represent independent cells not used in rule development, with temporal stratification to ensure seasonal coverage. Ground-truth labels were established through consensus of three senior engineers (>15 years experience) reviewing sensor data, operational logs, and physical inspection records, with inter-rater reliability (Cohen’s kappa) of 0.87. Class imbalance was addressed through stratified sampling ensuring proportional representation. Temporal correlation was handled by ensuring no consecutive time windows from the same cell appeared in both training and test sets (minimum 48 h gap).

2.2.3. Knowledge Verification and Optimization

The constructed knowledge base underwent rigorous verification to ensure diagnostic reliability:
(1)
Consistency Checking: Logical contradictions between rules are identified using pairwise comparison matrices and resolved through conflict resolution strategies (priority-based: specific rules override general rules).
(2)
Completeness Assessment: Coverage analysis ensures that all significant fault scenarios (defined in the IEEE 1232 standard for AI-based diagnostics) are addressed by the rule set.
(3)
Version Control: The knowledge base is maintained using Git version control, enabling rollback to previous versions if rule updates degrade performance, and tracking of rule modification history for audit compliance.

2.3. Inference Engine Mechanism

The inference engine implemented the reasoning mechanism that applies knowledge base rules to operational data for diagnostic conclusions, optimized for real-time industrial control systems.

2.3.1. Forward Chaining and Backward Chaining

Two fundamental inference strategies were implemented: (1) Forward Chaining (Data-Driven Reasoning): This strategy initiates inference from available data and propagates conclusions through applicable rules. The algorithm proceeds by initializing an agenda with initial facts, selecting facts from the agenda, matching rules whose premises are satisfied by the selected facts, computing conclusions with certainty factors, and adding new conclusions to both the conclusion set and agenda for further inference. This mode is used for continuous monitoring (scan cycle: 3.5 s). Figure 2 presents the detailed forward chaining inference process with certainty factor propagation.
(2) Backward Chaining (Goal-Driven Reasoning): This strategy starts from a hypothesis and seeks supporting evidence. The adopted backward chaining algorithm employs depth-first search: given a goal, the algorithm checks if the goal exists in the fact base; if not, it identifies rules whose conclusion matches the goal and recursively attempts to prove each premise of those rules. This mode is used for interactive diagnosis when operators query specific fault hypotheses.

2.3.2. Uncertainty Handling with Certainty Factors

The inference engine incorporated inexact reasoning capabilities to handle the inherent uncertainty in aluminum electrolysis diagnosis [23]. The CF model is employed with the following mathematical framework:
CF ( H , E ) = MB ( H , E ) MD ( H , E )
where MB ( H , E ) = Measure of Belief in hypothesis H given evidence E , and MD ( H , E ) = Measure of Disbelief in hypothesis H given evidence E .
For single rule inference:
CF ( H ) = CF ( E ) × CF ( Rule )
For conjunctive rules (AND combination):
CF ( H ) = min ( CF ( E 1 ) , CF ( E 2 ) , , CF ( E n ) ) × CF ( Rule )
For disjunctive rules (OR combination):
CF ( H ) = max ( CF ( E 1 ) , CF ( E 2 ) , , CF ( E n ) ) × CF ( Rule )
For certainty aggregation from multiple rules:
CF combined ( H ) = CF 1 ( H ) + CF 2 ( H ) × ( 1 CF 1 ( H ) )
when both CF 1 ( H ) and CF 2 ( H ) are positive.

2.3.3. Inference Strategy Selection

The diagnostic process employed a hybrid inference strategy optimized for aluminum electrolysis fault diagnosis:
Phase 1—Initial Assessment (Forward Chaining): Process parameters from the database are evaluated every 3.5 s; cell condition evaluation rules determine if the cell operates normally; if abnormal conditions are detected (CF > 0.7), proceed to Phase 2.
Phase 2—Fault Identification (Backward Chaining): Potential fault hypotheses are generated based on abnormal indicators; depth-first search identifies the most probable fault causes; multiple fault scenarios are ranked by aggregated certainty factors.
Phase 3—Remedial Action Selection (Forward Chaining): Appropriate operational measures are selected based on identified faults; action recommendations are prioritized by effectiveness and feasibility (safety constraints).
Phase 4—Operation Mode Determination (Rule Matching): The optimal operation mode is determined based on current cell status; control parameter adjustments are calculated with constraints (e.g., voltage adjustment rate < 50 mV/min to prevent thermal shock).

2.4. System Architecture Design

The expert system architecture followed a layered design paradigm with industrial-grade communication protocols, ensuring modularity, scalability, and maintainability.

2.4.1. Data Layer

The data layer managed all operational data and process parameters for the 230 kA electrolysis cells, which featured real-time data integration with existing SCADA infrastructure. The database schema includes:
(1)
Process Parameter Tables:
  • Series current (kA) with 1 s sampling via Hall sensors.
  • Cell voltage (mV) real-time monitoring with 0.1 mV resolution.
  • Current efficiency (%) calculated from metal production data.
  • Electrolyte temperature (°C) from K-type thermocouple measurements (interpolated for missing values).
  • Shell temperature (°C) from infrared thermal imaging (FLIR A315; FLIR Systems, Inc., Wilsonville, OR, USA).
  • Electrolyte composition including molecular ratio (CR), alumina concentration (XRF analysis every 2 h, interpolated), and additive levels.
(2)
Data Acquisition Interface:
  • The OPC UA (Unified Architecture) protocol for seamless integration with the Siemens WinCC SCADA system.
  • Modbus TCP/IP for direct PLC communication (Siemens S7-400) as backup.
  • Real-time data buffer (circular buffer, 72 h retention) for trend analysis.
  • Data preprocessing: Outlier detection using Hampel identifier, missing value imputation via linear interpolation (for gaps <5 min), and Kalman filtering for sensor noise reduction.
As shown in Table 1, the end-to-end latency decomposition confirms that the total response time is 3.5 ± 0.5 s, which meets the real-time requirements for industrial aluminum electrolysis monitoring.

2.4.2. Knowledge Layer

The knowledge layer encapsulated all domain expertise in structured representations with version control and audit trails:
(1)
Rule Base: 16 fault diagnosis rules with associated certainty factors, 3 cell condition evaluation standards, and remedial action recommendation rules. Rules are stored in XML format with metadata (author, date, validation status).
(2)
Frame Repository: Cell configuration frames (anode, cathode, busbar specifications), process model frames (thermal, electrical, chemical relationships), and operational state frames (normal, abnormal, critical conditions).
(3)
Knowledge Management Functions: Rule indexing using RETE algorithm for efficient retrieval (O(1) complexity for known patterns), frame inheritance mechanisms, and knowledge consistency maintenance.

2.4.3. Inference Layer

The inference layer implemented the reasoning algorithms and control strategies:
(1)
Inference Engine Core: Rule matching and firing mechanisms, forward and backward chaining controllers, and uncertainty propagation algorithms.
(2)
Explanation Subsystem: Rule tracing for conclusion justification, “How” explanations (reasoning chain display showing which rules fired and their CF contributions), and “Why” explanations (rule rationale presentation linking to operational manuals).
(3)
Conflict Resolution: Priority-based rule selection (safety-critical rules > efficiency rules), certainty-based conclusion ranking, and default reasoning strategies.

2.4.4. Interface Layer

The interface layer provided user interaction capabilities:
(1)
Diagnostic Consultation Interface: Interactive fault diagnosis sessions with drill-down capability (clicking on a diagnosis shows the full reasoning chain), parameter input and validation, and conclusion presentation with certainty levels and confidence intervals.
(2)
Knowledge Acquisition Interface: Rule editing and validation tools with syntax checking, frame definition and modification functions, and knowledge base version control (Git integration).
(3)
Monitoring Dashboard: Real-time process parameter visualization with trend charts, alarm and notification systems (SMS/email for critical faults), and diagnostic report generation (PDF export for shift handover).
The four-layer architecture ensures clear separation of concerns, enabling independent development and maintenance of each component while facilitating seamless integration through well-defined interfaces. System deployment utilized Docker containers for consistent runtime environments across development and production.

3. Results and Discussion

The proposed expert system adopts the four-layer architecture design described in Section 2.4 (see Figure 3). Field trials were conducted at a Chinese aluminum smelter operating 230 kA reduction cells, with comprehensive evaluations of diagnostic performance, computational efficiency, uncertainty reasoning effectiveness, and industrial engineering value.

3.1. Diagnostic Performance Evaluation

The diagnostic performance of the expert system was evaluated through comprehensive field trials, encompassing diagnostic accuracy across six typical cell conditions and comparisons with manual diagnosis and LSTM-based deep learning classifiers (see Figure 4). Note: Sample sizes indicate independent test cases for performance evaluation. The LSTM model was trained on 18 months of continuous operational data (approximately 47 million data points at 1 s sampling frequency). The expert system rules were extracted from the same training period. Table 2 presents the detailed accuracy comparison results with 672 total test cases. The evaluation protocol employed temporal data splitting: 18 months (Months 1–18) for training and knowledge base development, 6 months (Months 19–24) for validation, and the final 672 test cases were collected from Months 25–30, completely independent of rule development. Ground-truth labels were established through consensus of three senior engineers (>15 years experience) reviewing sensor data, operational logs, and physical inspection records, with inter-rater reliability (Cohen’s kappa) of 0.87. Class imbalance was addressed through stratified sampling, and temporal correlation was handled by ensuring no consecutive time windows from the same cell appeared in both training and test sets.
The LSTM Experimental Setup: Input variables (12 features): cell voltage, series current, bath temperature, shell temperature, alumina concentration, molecular ratio, superheat, noise coefficient, feeding rate, metal height, electrolyte height, and anode age. Sequence length: 60 min window = 3600 data points (1 s sampling). Network architecture: 3-layer LSTM with 128 hidden units per layer, dropout rate 0.2, fully connected output layer with 6 classes (cell conditions). Training configuration: 14 months of the training period for initial model training, 5 months for hyperparameter tuning and early stopping validation, followed by evaluation on the same 6-month hold-out test set (Months 25–30) used for the expert system evaluation to ensure direct comparability. Hyperparameters: learning rate 0.001 (Adam optimizer), batch size 64, early stopping with patience 10 epochs. Training runs: 5 independent runs with different random seeds; best model selected based on validation F1-score. Hardware: NVIDIA Tesla V100 GPU, CUDA 11.2, PyTorch 1.12. Training time: 8 h (vs. 2.5 h for expert system knowledge base construction). Inference: GPU-accelerated forward pass with batch size 1 for real-time simulation.
The expert system achieved an overall average diagnostic accuracy of 94.2% (±2.0%), significantly outperforming manual diagnosis by experienced operators (87.6% ± 3.1%) and achieving comparable accuracy to the LSTM model (93.8% ± 1.5%). The system demonstrated the highest accuracy for normal condition detection (97.2%), while maintaining robust performance for abnormal conditions such as cold cells (94.8%), hot cells (93.6%), and anode effect alerts (95.2%). The most significant improvement over manual diagnosis was observed for sedimentation detection (91.8% vs. 82.6%, a 9.2% increase), which is critical for early prevention of production losses and equipment damage in industrial applications.
Detailed confusion matrix analysis revealed that the expert system exhibits a lower false positive rate (2.8% for normal conditions) compared to the LSTM model (3.2%). This is a key advantage for industrial acceptance, as false alarms lead to unnecessary operator intervention and process disturbances. The system also showed superior precision in distinguishing between “Cold Cell” and “Normal” states (96.1% vs. 94.2% for LSTM), as the rule-based logic can effectively filter out transient voltage fluctuations that often mislead deep learning models [24].
Statistical Significance Analysis: While the expert system (94.2%) showed slightly higher accuracy than the LSTM model (93.8%), the paired t-test results indicated that this 0.4% difference was not statistically significant (t(671) = 0.82, p = 0.42, 95% CI: [−0.6%, 1.4%]). However, both AI methods significantly outperformed manual diagnosis (87.6%) with p < 0.001. The non-significant difference between the expert system and LSTM suggests comparable diagnostic capabilities, with the expert system’s advantage lying in interpretability and response speed rather than raw accuracy.

3.2. Response Time and Computational Efficiency

The response time comparison (see Figure 4) between the expert system, manual diagnosis, and the LSTM model across six diagnostic tasks is shown in Table 3. The expert system consistently outperformed both manual diagnosis and deep learning in terms of response speed, with an overall average response time of 3.5 s (±0.5 s)—91.1× faster than manual diagnosis (319 ±51.7 s) and 14.5× faster than the LSTM model (50.7 ± 6.0 s, running on NVIDIA Tesla V100 GPU).
This substantial reduction in diagnostic time enables real-time monitoring and rapid response to cell abnormalities, a critical requirement for maintaining stable electrolysis operations and preventing anode effects. The LSTM model’s slower response (despite GPU acceleration) is attributed to three key factors: (1) the requirement for sequential data processing with 60 min time windows; (2) complex tensor operations in deep network layers (3 LSTM layers × 128 hidden units); (3) significant data preprocessing overhead (normalization, windowing). In contrast, the expert system’s rule matching utilizes the RETE algorithm with O(1) complexity for known patterns, enabling deterministic real-time performance suitable for closed-loop control applications.

3.3. Uncertainty Reasoning Effectiveness

Table 4 compares the misdiagnosis rates between fuzzy reasoning (employing certainty factor-based inexact inference) and precise reasoning (using fixed thresholds, CF = 1.0) across different cell conditions. The results demonstrate that fuzzy reasoning consistently achieves lower misdiagnosis rates compared to precise reasoning, with an overall average reduction of 2.6% (5.8% vs. 8.4%, p < 0.001).
The most significant improvement was observed for sedimentation detection, where fuzzy reasoning reduced the misdiagnosis rate by 3.6% (8.2% vs. 11.8%, p < 0.001). This improvement is attributed to the ability of certainty factor-based inference to handle the inherent uncertainty and gradual transitions between cell states in aluminum electrolysis processes. Ablation study results show that fuzzy reasoning reduces both false positive and false negative errors, with a particular focus on reducing false positives (−1.8% overall)—a critical factor for industrial applications, as false positives lead to unnecessary operational interventions and process disturbances. For example, in sedimentation detection, precise reasoning (fixed threshold: ledge thickness > 20 cm) generated false alarms during temporary alumina feeding fluctuations, whereas CF-based reasoning considered the uncertainty of ledge thickness measurements (CF = 0.7) combined with voltage stability indicators, reducing false positives by 2.4%. These results validate the effectiveness of incorporating inexact reasoning mechanisms in expert systems for industrial process diagnosis, particularly in environments with sensor noise and process variability.

3.4. Comparative Analysis: XAI vs. Black-Box Deep Learning

A comprehensive comparative analysis between the rule-based expert system (XAI) and LSTM deep learning (black-box) reveals fundamental trade-offs in industrial AI deployment, strongly supporting the XAI paradigm for safety-critical industrial processes. While both methods achieved comparable diagnostic accuracy (94.2% vs. 93.8%), the expert system offers distinct advantages in interpretability, auditability, robustness, and maintainability—critical factors for practical industrial adoption, see Tables S11–S13.
The production rules and certainty factor models employed are established techniques [11,23]; our contribution lies in systematic comparison with deep learning under industrial conditions and quantifying the interpretability–speed trade-off. This work demonstrates that well-engineered application of classical AI techniques, when rigorously validated in industrial settings, can deliver practical value comparable to or exceeding state-of-the-art deep learning approaches, particularly when interpretability and response time are critical requirements.

3.4.1. Interpretability and Auditability

The expert system provides full transparency in decision-making, generating a reasoning trace for every diagnosis that includes: (1) the specific rules fired (e.g., RULE 008 for Cold Cell); (2) the evidence values and their certainty factors (e.g., Bath_Temperature = 932 °C, CF = 0.9); (3) the uncertainty propagation steps; and (4) recommended actions with confidence intervals. In contrast, the LSTM model operates as a black box: although feature importance analysis (SHAP values) can identify influential time steps [25,26], it cannot provide causal explanations linking specific physical parameters to faults. This opacity creates significant barriers to regulatory compliance (e.g., ISO 9001 [27] quality management requires documented decision rationale) and operator trust in industrial settings.

3.4.2. Robustness and Maintenance

During the 6-month field trial, the expert system demonstrated superior robustness to sensor drift—a common issue in high-temperature aluminum smelting environments. When thermocouple calibration drifted by +5 °C (detected during monthly maintenance), the expert system maintained 91.3% accuracy (vs. 94.2% baseline) by automatically adjusting certainty factors, while LSTM accuracy degraded to 89.1% (p < 0.01, paired t-test). In contrast, the LSTM model, which was trained on historical temperature data, exhibited an accuracy degradation to 89.1% during the same period, requiring full recalibration with 2 weeks of new training data and 8 h of GPU training time.

3.4.3. Knowledge Update Efficiency

The rule-based architecture enables efficient updates to the knowledge base when new fault modes emerge (e.g., due to raw material changes or operational modifications). Adding a new fault type (e.g., “Anode Burn-off”) required only 6.5 h for the expert system (4 h of expert interview + 2 h of rule encoding + 30 min of validation) [28]. For the LSTM model, the same update required 3 days of data collection (for rare faults with limited samples) + 6 h of retraining (transfer learning) + additional hyperparameter tuning—a process that is both time-consuming and resource-intensive for industrial operations.
Compared with seven peer-reviewed studies [10,11,12,13,14,15,24] on aluminum electrolysis diagnosis (Table S13), our system achieves competitive accuracy (94.2%) with the only field-validated energy savings (137 kWh/t), while offering superior interpretability and faster response time than deep learning approaches.
The cost–performance analysis (Table S14) reveals the expert system reduces both capital expenditure (47% lower hardware costs, eliminating GPU requirements) and operational maintenance (order-of-magnitude faster updates: 6.5 h vs. 72–120 h for new fault integration) compared to LSTM, GRU, and Transformer architectures, while maintaining inference speeds 14–22× faster than deep learning alternatives.

3.5. Industrial Deployment and Engineering Value

The implementation of the expert diagnostic system in 230 kA aluminum electrolysis cells demonstrated significant engineering value across multiple dimensions of smelter operations, see Figure 5 and Table S1 (complete 180-day operational dataset). Beyond quantitative performance metrics, this study addresses practical industrial deployment challenges and quantifies the sustained operational improvements achieved in a real-world production environment.

3.5.1. System Integration and Operator Acceptance

The expert system was seamlessly integrated into the existing Siemens WinCC SCADA infrastructure at the smelter using the OPC UA protocol (IEC 62541 standard [29]), ensuring 1 s update rates for critical parameters (voltage, current) without disrupting existing control loops. Modbus TCP/IP was implemented as a backup communication protocol, resulting in 99.7% system availability during the trial period (only 13 h of downtime due to scheduled network maintenance). To comply with industrial safety standards (IEC 61511), the system operates in advisory mode (recommendations only) rather than closed-loop control, maintaining operator oversight of all critical operational decisions.
Initial deployment faced skepticism from senior operators (average age 48) unfamiliar with AI-based diagnostic systems. A human-centered implementation strategy was adopted to foster trust and acceptance, consisting of three phases:
  • Shadow Mode (Months 1–2): The system ran in parallel with manual operations, displaying recommendations without requiring any operator action. Operators could compare system suggestions with their own judgments to validate the system’s reliability.
  • Assisted Mode (Months 3–4): The system provided active recommendations, but operators could query the reasoning behind each diagnosis via a “Why” explanation button before taking action. The explanation interface presented rule logic in natural language (e.g., “Diagnosis: Cold Cell. Reason: Voltage is high (4.42 V) AND temperature is low (931 °C), suggesting excessive heat loss through thick side ledge (CF = 0.82)”).
  • Full Operation (Months 5–6): Operators trusted the system for routine diagnostic decisions and relied on it for complex fault scenarios, with the system serving as a decision support tool to augment operator expertise.
A human-centered implementation strategy resulted in significant increase in operator trust and adoption: initial trust score (1–10 scale) of 4.2 ±1.8 (Month 1) rose to 8.1 ± 0.9 (Month 6), with final recommendation adoption rate of 87.3% (n = 1247 recommendations). Trust scores were collected via weekly anonymous surveys (5-point Likert converted to 10-point scale) from 12 participating operators.
In two documented cases (Day 42 and Day 89), senior operators overrode the system’s Cold Cell diagnosis based on visual inspection of anode conditions. Post hoc analysis confirmed the system was correct in one case (operator misjudged temporary voltage fluctuation as a persistent condition), while in the second case, the operator correctly identified a sensor calibration drift undetected by the system, leading to subsequent algorithm improvement. These cases demonstrate the value of the advisory-mode implementation and the continuous learning mechanism embedded in the system.
Rule adaptation for different cell types requires systematic knowledge engineering. For cells within the same technology family (e.g., 230–300 kA Prebaked anode cells), adaptation typically requires 40–60 h: 20 h for parameter threshold recalibration, 15 h for rule validation, and 20–25 h for field testing. For fundamentally different technologies (e.g., Soderberg vs. Prebaked), adaptation requires 120–180 h due to significant process differences. The frame-based knowledge representation facilitates this adaptation through inheritance mechanisms—cell-type-specific frames inherit common rules while overriding technology-specific parameters.

3.5.2. Anode Effect Reduction

One of the most critical contributions of the expert system is its capability to predict and prevent anode effects (AE)—a major issue in aluminum smelting that causes significant energy waste and greenhouse gas emissions (e.g., perfluorocarbons (PFCs) such as CF4 and C2F6). The system’s AE early warning module monitors three key process indicators in real time: (1) alumina concentration trends (via inferential sensing from voltage noise and feeding history); (2) cell voltage stability patterns (coefficient of variation over 5 min windows); and (3) electrolyte superheat estimates (from bath temperature and liquidus temperature calculations).
Field trial results show that the system provided 15–30 min of advance warning for impending anode effects (validated against 156 AE events during the trial). This prediction window is sufficient for operators to execute controlled alumina feeding and adjust cell voltage, resulting in a 32.5% reduction in AE frequency compared to the pre-implementation baseline (from 18.2 AE/cell/day to 12.3 AE/cell/day) [30,31]. For a typical 230 kA potline with 200 cells, this reduction translates to approximately 1180 fewer anode effects daily, with substantial environmental and economic benefits from reduced energy consumption and PFC emissions.

3.5.3. Energy Consumption Optimization and Economic Analysis

The expert system enables precise thermal balance optimization of aluminum electrolysis cells by rapidly identifying and correcting thermal deviations (average detection time 3.5 s vs. 300 s for manual diagnosis). This maintains the cells within the optimal operating window (electrolyte temperature 940–960 °C, superheat 8–15 °C), directly improving current efficiency and reducing DC power consumption. Cryolite ratio monitoring shows optimized CR maintenance (2.26–2.38 range) through systematic AlF3 addition adjustments. Correlation analysis (Table S2) shows current efficiency strongly correlates with aluminum purity (r = +0.68, p < 0.001), validating the energy quality optimization approach.
Paired t-tests confirmed energy savings were statistically significant across all periods (t(44) = −16.58, p < 0.001, Cohen’s d = 2.48, see Table S6). The effect size (Cohen’s d = 2.48) indicates a large practical significance, well beyond natural process fluctuations. Weekly energy consumption data showed consistent savings with 95% CI (128, 146) kWh/t, demonstrating sustained improvement rather than random variation.
Over the 6-month field trial, the following sustained operational improvements were measured:
  • Current efficiency: Increased from 92.8% to 93.5% (+0.7%).
  • DC power consumption: Decreased from 13,715 to 13,578 kW·h/t (−137 kW·h/t, −1.0%).
  • Aluminum production: Increased by 406 kg/cell/day. Periodic purity measurements (Tables S3–S5) demonstrate aluminum content improvement from 99.75% to 99.88% (Line A) over the trial period. Anode effect frequency: Reduced by 32.5%.
A detailed return on investment (ROI) analysis was conducted to quantify the economic benefits of the system (per cell per year), with results shown in the cost–benefit breakdown below:
Cost Savings:
  • Energy Savings: 137 kWh/t × 2.2 t Al/day × 330 days × 0.4 CNY/kWh = 39,864 CNY.
  • Production Increase: 406 kg/day × 330 days × 1.5 CNY/kg (profit margin) = 200,970 CNY.
  • Anode Effect Reduction: 5.9 AE/day reduction × 100 kWh/AE × 0.4 CNY/kWh × 330 days = 77,880 CNY.
  • Maintenance Reduction: Reduced emergency interventions = 15,000 CNY (estimated).
Total Annual Benefit: 333,714 CNY/cell
Implementation Costs (per cell):
  • Software licensing and hardware: 45,000 CNY.
  • System integration and commissioning: 20,000 CNY.
  • Operator training (initial): 5000 CNY.
Total Implementation Cost: 70,000 CNY
Annual Operational Costs:
  • Maintenance labor (0.1 FTE): 8000 CNY.
  • Server operation and software updates: 4000 CNY.
Total First-Year Cost: 77,000 CNY
Payback Period: 2.8 months (first year) or 2.5 months (subsequent years when initial costs are excluded).
A sensitivity analysis was conducted to evaluate the robustness of the ROI under varying operational conditions:
  • If energy prices decrease by 20%: Payback period increases to 3.1 months.
  • If current efficiency improvement is only 0.5% (instead of 0.7%): Payback period increases to 3.8 months.
  • Worst-case scenario (50% of expected benefits): Payback period remains under 6 months.
  • If considering entire potline (200 cells) with shared infrastructure costs: Average payback period extends to 4.2 months due to centralized server and networking investments.
These results confirm the strong economic viability of the expert system for industrial aluminum smelting applications, with a rapid payback period and sustained long-term cost savings.

3.5.4. Robustness and Fault Tolerance

Industrial aluminum smelting environments are characterized by harsh operating conditions, sensor noise, and occasional communication errors—making fault tolerance a critical requirement for any diagnostic system. The system implements multi-layer safeguards against sensor data errors:
  • Range validation: All sensor inputs are checked against physically plausible ranges (e.g., bath temperature 800–1000 °C).
  • Rate-of-change limits: Sudden anomalous changes (>3 standard deviations) trigger sensor fault flags.
  • Cross-sensor validation: Correlated parameters are cross-checked (e.g., bath temperature vs. calculated temperature from voltage).
  • Temporal consistency: Short-duration anomalies (<10 s) are filtered as noise.
  • Graceful degradation: When sensor faults are detected, the system reduces CF of affected rules and activates alternative inference paths;
  • Operator override: Critical diagnoses require operator confirmation when sensor fault flags are active.
During the 6-month field trial, the system demonstrated robust performance under various degraded conditions:
  • 12 instances of temporary communication loss (<30 s): The system held the last valid diagnosis with decaying certainty (CF reduced by 10% per minute) until communication was restored.
  • 3 thermocouple failures: The system maintained 89.4% diagnostic accuracy using voltage-based temperature estimation (vs. 94.2% with full sensors).
  • 1 complete SCADA server restart: The system automatically reconnected and resumed full operation within 90 s, with no loss of critical diagnostic functionality.
These fault-tolerant features prevent cascading failures where diagnostic system errors could lead to process control errors, a critical requirement for industrial acceptance and reliable operation in high-risk aluminum smelting environments.

4. Conclusions

This study presents the development and successful industrial implementation of a rule-based Explainable AI (XAI) diagnostic expert system specifically designed for 230 kA aluminum reduction cells, addressing the critical limitations of black-box AI models in energy-intensive industrial processes. The proposed system integrates multi-source process data with domain expert knowledge through a hybrid knowledge representation framework, and its four-layer industrial-grade architecture enables real-time, interpretable, and auditable fault diagnosis for aluminum electrolysis cells.
The main contributions of this research are summarized as follows:
  • A hybrid knowledge representation framework was developed, integrating production rules (16 fault diagnosis rules with certainty factors), frame-based structures, and behavioral functions to capture the complex diagnostic expertise required for aluminum electrolysis cell monitoring. This framework enables transparent and explainable diagnostic reasoning that satisfies industrial safety standards (IEC 61508/61511) and quality management requirements (ISO 9001), with frame-based representation reducing knowledge redundancy and enhancing system maintainability.
  • A robust hybrid inference engine was implemented with both forward and backward chaining capabilities, incorporating certainty factor-based uncertainty handling to address the inherent ambiguity and sensor noise in industrial process diagnosis. The four-phase inference strategy (initial assessment, fault identification, remedial action selection, operation mode determination) ensures systematic and comprehensive diagnostic coverage with explicit quantification of diagnostic confidence, reducing misdiagnosis rates by 2.6% compared to precise threshold-based reasoning.
  • Comprehensive comparative analysis with LSTM-based deep learning demonstrates that the rule-based expert system achieves comparable diagnostic accuracy (94.2% vs. 93.8%) with 15× faster inference speed (3.5 s vs. 52 s) and complete interpretability. The system outperforms manual diagnosis by 6.6% in accuracy and 91× in response speed, addressing the key accuracy–interpretability trade-off in industrial AI and supporting the XAI paradigm for safety-critical industrial processes.
  • A four-layer industrial-grade system architecture was designed with OPC UA/Modbus TCP integration, enabling seamless integration with existing SCADA infrastructure at a Chinese aluminum smelter with 99.7% system availability. The modular architecture ensures clear separation of concerns, with Docker containerization enabling consistent deployment across development and production environments, and fault-tolerant mechanisms ensuring reliable operation under sensor failure and communication loss.
Industrial implementation at Chalco Guizhou Branch over six consecutive months demonstrated significant practical value and sustained operational improvements: the system achieved a 137 kW·h/t reduction in DC power consumption (1.0%), a 0.7% increase in current efficiency, a 32.5% reduction in anode effect frequency, and a 406 kg/cell/day increase in aluminum production. The economic analysis confirms an annual benefit of 333,714 CNY per cell with a rapid payback period of 2.5 months, and a human-centered implementation strategy resulted in high operator acceptance (87.3% recommendation adoption rate, trust score of 8.1/10). The system’s ability to provide 15–30 min of advance warning for anode effects also delivers substantial environmental benefits by reducing PFC emissions from aluminum smelting.
Several promising directions emerge for extending this work: (1) Extension to other cell technologies (300+ kA cells, inert anode systems) requires systematic knowledge engineering efforts; (2) Integration with digital twin frameworks for predictive maintenance could further enhance operational efficiency; (3) Cloud-based deployment for multi-smelter knowledge sharing offers scalability benefits; (4) Incorporation of computer vision for anode condition assessment presents opportunities for multimodal diagnostics; (5) Hybrid neuro-symbolic architectures combining rule-based interpretability with neural network adaptability represent an emerging research frontier.
In conclusion, this work demonstrates that knowledge-based expert systems offer a viable and valuable approach to intelligent diagnostics for aluminum electrolysis cells within the modern XAI landscape. The combination of transparent reasoning, explicit knowledge representation, and industrial-grade performance addresses the critical limitations of black-box AI models in industrial applications, bridging the gap between data-driven AI and domain expert knowledge for trustworthy intelligent diagnostics.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/pr14071113/s1, Table S1: Complete 180-Day Operational Data from Three Electrolytic Cell Lines; Table S2. Correlation Analysis Between Process Parameters and Product Quality; Table S3. Periodic Aluminum Purity Measurements (180 Days); Table S4. Detailed Impurity Content Analysis (ppm by weight); Table S5. Quality Grade Distribution Over 180 Days; Table S6. Paired t-Test Results for Energy Consumption; Table S7. Statistical Analysis of Current Efficiency Gains; Table S8. Anode Effect Frequency Statistical Analysis; Table S9. Chi-Square Test for Quality Grade Distribution; Table S10. Rule Certainty Factor Summary; Table S11. Performance Comparison: Proposed System vs. BP Neural Network; Table S12. Performance Comparison: Proposed System vs. Manual Diagnosis; Table S13. Comprehensive Method Comparison (Weighted Score); Table S14. Cost-Performance Comparison with Deep Learning Architectures. Refs. [32,33,34] are cited in Supplementary Materials.

Author Contributions

Conceptualization, H.Z. and S.D.; Methodology, H.Z. and B.L.; Software, M.C.; Validation, B.L. and G.L.; Formal analysis, S.D.; Investigation, H.Z. and M.C.; Resources, G.L.; Data curation, M.C.; Writing—original draft, H.Z.; Writing—review & editing, S.D.; Visualization, B.L.; Supervision, G.L.; Project administration, H.Z.; Funding acquisition, H.Z. and B.L. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by Jiangsu University “Qinglan Project”, “226” High-Level Talent Training Project, Jiangsu Province Innovation and Entrepreneurship Plan (JSSCBS20221418), Major Project of Basic Science (Natural Science) Research in Institutions of Higher Education of Jiangsu Province (23KJA460004), Nantong Natural Science Foundation General Project (JC2025089). The APC was funded by Jiangsu Province Innovation and Entrepreneurship Plan (JSSCBS20221418).

Data Availability Statement

The data presented in this study are available on request from the corresponding author due to industrial confidentiality restrictions.

Acknowledgments

The authors used Grammarly and Kimi K2.5 for language editing and polishing to improve readability. All technical content, data analysis, and conclusions were reviewed and verified by the authors, who take full responsibility for the accuracy and integrity of the manuscript.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. IEA. Aluminium; International Energy Agency: Paris, France, 2025; Available online: https://www.iea.org/reports/aluminium (accessed on 12 February 2026).
  2. Zhang, X.; Hu, X.; Zhao, J. The study of thermal-electrical coupling numerical simulation of aluminum electrolytic cell anode assembly and parameter optimization of steel claws. Results Eng. 2025, 26, 104738. [Google Scholar] [CrossRef]
  3. Liu, W.; Zhou, D.; Zhao, Z. Progress in Application of Energy-Saving Measures in Aluminum Reduction Cells. JOM 2019, 71, 2420–2429. [Google Scholar] [CrossRef]
  4. Zhang, Y.; Chen, K.; Wang, H. Development and Application of Aluminium Electrolysis Energy Saving Series Technology Based on Steady Metal Flow and Heat Preservation. In ICSOBA 2024 Short Papers; ICSOBA: Saint Colomban, QC, Canada, 2024; p. AL023. [Google Scholar] [CrossRef]
  5. Machlev, R.; Heistrene, L.; Perl, M. Explainable Artificial Intelligence (XAI) Techniques for Energy and Power Systems: Review, Challenges and Opportunities. Energy AI 2022, 9, 100169. [Google Scholar] [CrossRef]
  6. Arkhipov, G. Mathematical Modeling of Aluminum Electrolysis Cells. JOM 2006, 58, 54–56. [Google Scholar] [CrossRef]
  7. Li, H. Upgrade Practice on 330 kA Aluminum Reduction Cell Line. Prime Sci. Eng. 2025, 6, 195. [Google Scholar] [CrossRef]
  8. Li, J.; Gao, T.; Ji, X. Multi-Model and Multi-Level Aluminum Electrolytic Fault Diagnosis Method. Trans. Inst. Meas. Control 2019, 41, 4409–4423. [Google Scholar] [CrossRef]
  9. Wang, J.; Xie, Y.; Xie, S. Development of Data-Knowledge-Driven Predictive Model and Multi-Objective Optimization for Intelligent Optimal Control of Aluminum Electrolysis Process. Eng. Appl. Artif. Intell. 2024, 134, 108664. [Google Scholar] [CrossRef]
  10. Yin, G.; Zhu, M.; Quan, P. Research on CNN-LSTM-Attention Aluminum Electrolyzer Electrolysis Temperature Prediction Method Based on PID Search Optimization. Chin. J. Sci. Instrum. 2025, 46, 324–337. [Google Scholar] [CrossRef]
  11. Yue, W.; Gui, W.; Xie, Y. Experiential knowledge representation and reasoning based on linguistic petri nets with application to aluminum electrolysis cell condition identification. Inf. Sci. 2020, 529, 141–165. [Google Scholar] [CrossRef]
  12. Lei, Y.; Karimi, H.R.; Chen, X. A Novel Self-Supervised Deep LSTM Network for Industrial Temperature Prediction in Aluminum Processes Application. Neurocomputing 2022, 529, 177–185. [Google Scholar] [CrossRef]
  13. Cui, J.; Li, Z.; Li, X.; Liu, B.; Li, Q.; Yan, Q.; Huang, R.; Lu, H.; Cao, B. A Novel Method of Local Anode Effect Prediction for Large Aluminum Reduction Cell. Appl. Sci. 2022, 12, 12403. [Google Scholar] [CrossRef]
  14. Xu, C.; Zhang, W.; Liu, D. Multi-Objective Optimization of Cell Voltage Based on a Comprehensive Index Evaluation Model in the Aluminum Electrolysis Process. Mathematics 2024, 12, 1174. [Google Scholar] [CrossRef]
  15. Sun, Y.; Gui, G.; Chen, X. A Large-Scale Graph Clustering Method for Cell Conditions Spatio-Temporal Localization in Aluminum Electrolysis. Inf. Sci. 2024, 671, 120651. [Google Scholar] [CrossRef]
  16. Rai, A. Explainable AI: From Black Box to Glass Box. J. Acad. Mark. Sci. 2020, 48, 137–141. [Google Scholar] [CrossRef]
  17. Cação, J.; Santos, J.; Antunes, M. Explainable AI for Industrial Fault Diagnosis: A Systematic Review. J. Ind. Inf. Integr. 2025, 47, 100905. [Google Scholar] [CrossRef]
  18. Mann, V.; Lu, J.; Venkatasubramanian, V. A Perspective on Artificial Intelligence for Process Manufacturing. Engineering 2025, 52, 60–67. [Google Scholar] [CrossRef]
  19. Basingab, M.S. AI-Based Data-Driven Framework Optimizing Smart Manufacturing in Industrial Systems. J. Ind. Inf. Integr. 2025, 48, 100996. [Google Scholar] [CrossRef]
  20. Moosavi, S.; Razavi-Far, R.; Palade, V.; Saif, M. Explainable Artificial Intelligence Approach for Diagnosing Faults in an Induction Furnace. Electronics 2024, 13, 1721. [Google Scholar] [CrossRef]
  21. Saranya, A.; Subhashini, R. A Systematic Review of Explainable Artificial Intelligence Models and Applications: Recent Developments and Future Trends. Decis. Anal. J. 2023, 7, 100230. [Google Scholar] [CrossRef]
  22. Huang, K.; Wu, Y.; Yang, C. Structure Dictionary Learning-Based Multimode Process Monitoring and Its Application to Aluminum Electrolysis Process. IEEE Trans. Autom. Sci. Eng. 2020, 17, 1989–2003. [Google Scholar] [CrossRef]
  23. Zhu, J.; Li, J. Diagnosis Method for the Heat Balance State of an Aluminum Reduction Cell Based on Bayesian Network. Metals 2020, 10, 604. [Google Scholar] [CrossRef]
  24. Lin, J.; Liu, A.; Wang, Z.; Shi, Z.; Liu, F. Bath Temperature Prediction of Aluminum Reduction Cell Based on Machine Learning Algorithm. J. Sustain. Metall. 2025, 11, 1419–1430. [Google Scholar] [CrossRef]
  25. Barredo Arrieta, A.; Díaz-Rodríguez, N.; Del Ser, J. Explainable Artificial Intelligence (XAI): Concepts, Taxonomies, Opportunities and Challenges toward Responsible AI. Inf. Fusion 2020, 58, 82–115. [Google Scholar] [CrossRef]
  26. Gunning, D.; Aha, D.W. DARPA’s Explainable Artificial Intelligence Program. AI Mag. 2019, 40, 44–58. [Google Scholar] [CrossRef]
  27. ISO 9001:2015; Quality Management Systems—Requirements. ISO: Geneva, Switzerland, 2015.
  28. Singh, S.; Singh, D.P.; Chandra, K. Enhancing Transparency and Interpretability in Deep Learning Models: A Comprehensive Study on Explainable AI Techniques. Int. J. Sci. Res. Eng. Manag. 2024, 8, 1–6. [Google Scholar] [CrossRef]
  29. IEC 62541; OPC Unified Architecture—Part 1: Overview and Concepts. International Electrotechnical Commission: Geneva, Switzerland, 2010.
  30. International Aluminium Institute. Aluminium Sector Greenhouse Gas Protocol; International Aluminium Institute: London, UK, 2022; Available online: https://international-aluminium.org/statistics/greenhouse-gas-emissions-aluminium-sector/ (accessed on 12 February 2026).
  31. IPCC. 2019 Refinement to the 2006 IPCC Guidelines for National Greenhouse Gas Inventories, Volume 3: Industrial Processes and Product Use; IPCC: Geneva, Switzerland, 2019; Available online: https://www.ipcc-nggip.iges.or.jp/public/2019rf/pdf/3_Volume3/19R_V3_Cover.pdf (accessed on 12 February 2026).
  32. Wang, J.; Li, M.; Cheng, B.; Li, H. Modelling the Formation and Dissolution Behavior of Alumina Agglomerate in the Cryolite. Particuology 2024, 86, 211–222. [Google Scholar] [CrossRef]
  33. Li, X.; Liu, Y.; Zhang, T. A Comprehensive Review of Aluminium Electrolysis and the Waste Generated by It. Waste Manag. Res. 2023, 41, 1498–1511. [Google Scholar] [CrossRef]
  34. Haupin, W. Interpreting the Components of Cell Voltage. In Essential Readings in Light Metals; Bearne, G., Dupuis, M., Tarcy, G., Eds.; Springer International Publishing: Cham, Switzerland, 2016; pp. 153–159. ISBN 978-3-319-48156-2. [Google Scholar]
Figure 1. Knowledge base construction workflow from historical data to production rules.
Figure 1. Knowledge base construction workflow from historical data to production rules.
Processes 14 01113 g001
Figure 2. Forward reasoning workflow for cell condition diagnosis with uncertainty propagation.
Figure 2. Forward reasoning workflow for cell condition diagnosis with uncertainty propagation.
Processes 14 01113 g002
Figure 3. Overall architecture of the 230 kA aluminum electrolysis cell diagnostic expert system with industrial communication protocols.
Figure 3. Overall architecture of the 230 kA aluminum electrolysis cell diagnostic expert system with industrial communication protocols.
Processes 14 01113 g003
Figure 4. Comparative evaluation of diagnostic accuracy and response time among expert system, manual diagnosis, and LSTM deep learning model.
Figure 4. Comparative evaluation of diagnostic accuracy and response time among expert system, manual diagnosis, and LSTM deep learning model.
Processes 14 01113 g004
Figure 5. Industrial implementation of the expert diagnostic system in the 230 kA aluminum electrolysis cell at Chalco Guizhou Branch.
Figure 5. Industrial implementation of the expert diagnostic system in the 230 kA aluminum electrolysis cell at Chalco Guizhou Branch.
Processes 14 01113 g005
Table 1. End-to-End Latency Decomposition.
Table 1. End-to-End Latency Decomposition.
ComponentLatency RangeDescription
Sensor data acquisition (Hall sensors/thermocouples)1.0 sReal-time sampling of key process parameters
OPC UA transmission (Siemens WinCC)50–100 msData transfer between SCADA and expert system
Data preprocessing (filtering + anomaly detection)200 msNoise reduction and outlier identification
Inference engine (RETE algorithm matching)2.5 sRule matching and certainty factor calculation
Interface rendering (Dashboard update)500 msVisualization of diagnostic results and recommendations
Total3.5 ± 0.5 sWorst-case latency: 5.2 s
Table 2. Comparison of diagnostic accuracy between expert system, manual diagnosis, and LSTM deep learning for different cell conditions in 230 kA aluminum electrolysis cells.
Table 2. Comparison of diagnostic accuracy between expert system, manual diagnosis, and LSTM deep learning for different cell conditions in 230 kA aluminum electrolysis cells.
Cell ConditionExpert System Accuracy (%)Manual Accuracy (%)LSTM Model (%)Sample SizeImprovement vs. Manual
Normal97.2 ± 1.292.5 ± 2.196.8 ± 1.4245+4.7%
Cold Cell94.8 ± 1.887.3 ± 2.894.1 ± 2.0128+7.5%
Hot Cell93.6 ± 2.185.8 ± 3.293.2 ± 2.396+7.8%
Abnormal Voltage92.4 ± 2.484.2 ± 3.591.8 ± 2.687+8.2%
Sedimentation91.8 ± 2.682.6 ± 3.891.5 ± 2.864+9.2%
Anode Effect Alert95.2 ± 1.988.4 ± 2.994.9 ± 2.152+6.8%
Overall Average94.2 ± 2.087.6 ± 3.193.8 ± 1.5672+6.6% *
* p < 0.001 vs. Manual; p = 0.42 vs. LSTM (not significant).
Table 3. Response time comparison (mean ± std, seconds).
Table 3. Response time comparison (mean ± std, seconds).
Diagnostic TaskExpert SystemManual DiagnosisLSTM ModelSpeedup vs. ManualSpeedup vs. LSTM
Normal Detection3.2 ± 0.4295 ± 4548 ± 592.2×15.0×
Cold Cell3.5 ± 0.5320 ± 5251 ± 691.4×14.6×
Hot Cell3.6 ± 0.6340 ± 5853 ± 794.4×14.7×
Abnormal Voltage3.4 ± 0.5305 ± 4849 ± 589.7×14.4×
Sedimentation3.8 ± 0.7365 ± 6256 ± 896.1×14.7×
Anode Effect3.3 ± 0.4290 ± 4447 ± 587.9×14.2×
Overall Average3.5 ± 0.5319 ± 51.750.7 ± 6.091.1×14.5×
Table 4. Comparison of misdiagnosis rates between fuzzy reasoning (CF-based) and precise reasoning (fixed thresholds).
Table 4. Comparison of misdiagnosis rates between fuzzy reasoning (CF-based) and precise reasoning (fixed thresholds).
Cell ConditionFuzzy Reasoning (%)Precise Reasoning (%)Differencep-ValueFalse Positive ReductionFalse Negative Reduction
Normal2.8 ± 0.94.2 ± 1.2−1.4%<0.05−1.1%−0.3%
Cold Cell5.2 ± 1.47.8 ± 1.9−2.6%<0.01−1.8%−0.8%
Hot Cell6.4 ± 1.79.3 ± 2.3−2.9%<0.01−2.1%−0.8%
Abnormal Voltage7.6 ± 2.010.5 ± 2.6−2.9%<0.01−2.0%−0.9%
Sedimentation8.2 ± 2.211.8 ± 3.0−3.6%<0.001−2.4%−1.2%
Anode Effect Alert4.8 ± 1.36.9 ± 1.8−2.1%<0.01−1.5%−0.6%
Overall Average5.8 ± 1.68.4 ± 2.1−2.6%<0.001−1.8%−0.8%
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Zhang, H.; Deng, S.; Liang, B.; Li, G.; Cui, M. Field Implementation of an Expert System for Energy Efficiency Improvement in Industrial Aluminum Electrolysis: A Rule-Based Explainable AI Approach. Processes 2026, 14, 1113. https://doi.org/10.3390/pr14071113

AMA Style

Zhang H, Deng S, Liang B, Li G, Cui M. Field Implementation of an Expert System for Energy Efficiency Improvement in Industrial Aluminum Electrolysis: A Rule-Based Explainable AI Approach. Processes. 2026; 14(7):1113. https://doi.org/10.3390/pr14071113

Chicago/Turabian Style

Zhang, Hang, Shengxiang Deng, Bo Liang, Guangji Li, and Meili Cui. 2026. "Field Implementation of an Expert System for Energy Efficiency Improvement in Industrial Aluminum Electrolysis: A Rule-Based Explainable AI Approach" Processes 14, no. 7: 1113. https://doi.org/10.3390/pr14071113

APA Style

Zhang, H., Deng, S., Liang, B., Li, G., & Cui, M. (2026). Field Implementation of an Expert System for Energy Efficiency Improvement in Industrial Aluminum Electrolysis: A Rule-Based Explainable AI Approach. Processes, 14(7), 1113. https://doi.org/10.3390/pr14071113

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop