CNN–Transformer–KAN: A Hybrid Deep-Learning Framework with an Inspectable KAN Classification Head for Industrial Process Fault Diagnosis
Abstract
1. Introduction
- We propose CTKAN, a hybrid CNN–Transformer–KAN architecture that couples local temporal feature extraction (CNN) with global inter-time-step dependency modeling (Transformer) and a classification head whose per-edge B-spline activation functions can be plotted and inspected individually. The novelty lies in this inspectable-by-construction classification design rather than in maximizing benchmark accuracy.
- Under the reported TE protocol, CTKAN achieves a competitive Macro-F1 of 91.38 ± 0.26% over ten independent training runs among the four main-table architectures, and a factorial ablation quantifies the contribution of the Transformer stage.
- At matched parameter capacity (≈526k), the KAN and MLP classification heads yield comparable Macro-F1 (91.38 ± 0.26% vs. 91.43 ± 0.37%) within seed-to-seed variability; the distinguishing benefit of the KAN head is its directly inspectable B-spline activation curves rather than an accuracy gain.
- Through per-class analyses of hard-to-diagnose TE faults (IDV(3), IDV(9), IDV(15)), we document specific strengths and weaknesses of the model rather than presenting CTKAN as universally superior.
2. Related Work
2.1. Convolutional and Recurrent Approaches
2.2. Transformer-Based and Hybrid Approaches
2.3. Kolmogorov–Arnold Networks in Fault Diagnosis
2.4. Transparency of the Classification Layer
2.5. Comparative Analysis and Research Gap
3. Proposed Method
3.1. Problem Formulation
3.2. Overall Architecture
3.3. CNN Encoder for Local Feature Extraction
3.4. Transformer Encoder for Global Dependency Modeling
3.5. KAN Classification Head
3.6. Training Procedure
3.7. Overall Algorithm
| Algorithm 1. Offline training and online inference of CTKAN |
| Input: training set Dtr, validation set Dval; epochs E; batch size B; learning-rate bounds (ηmax, ηmin). Output: trained parameters θ* of the CTKAN model. Offline training stage: 1: Compute per-variable mean μ and standard deviation σ on Dtr 2: Standardize all windows in Dtr and Dval using Equation (3) 3: Initialize parameters θ = {θcnn, θtrans, θkan} 4: for e = 1 to E do 5: for each mini-batch (X, y) in Dtr do 6: Hc ← CNN-Encoder(X) // Equations (10)–(12) 7: Ht ← Transformer-Encoder(Hc) // Equations (13)–(20) 8: z ← LayerNorm(MeanPool(Ht)) // Equation (8) 9: p ← softmax(KAN-Head(z)) // Equations (9) and (21)–(24) 10: L ← CrossEntropy(p, y) // Equation (25) 11: Update θ with Adam using the gradient of L 12: end for 13: Update the learning rate by cosine annealing // Equation (26) 14: Evaluate Macro-F1 on Dval; keep the best θ* so far 15: end for 16: return θ* Online inference stage: 17: Acquire a new process window Xnew and standardize it with stored μ, σ // Equation (3) 18: p ← fθ*(Xnew) // Equations (6)–(9) 19: ŷ ← argmaxc pc // Equation (5) 20: (optional) Plot the KAN edge activations φ for classification-layer auditing // Equation (21) 21: return predicted fault ŷ and inspection plots |
4. Experiments and Results
4.1. Experimental Details
4.1.1. Datasets
4.1.2. Baselines
4.1.3. Ablation Design
4.2. Main Comparison Results on the TE Dataset
4.3. Ablation Study
4.4. Inspectability Analysis
4.4.1. Feature Space Visualization via t-SNE
4.4.2. Attention Weight Analysis
4.4.3. KAN Activation Curve Visualization
4.5. Analysis of Hard-to-Diagnose Faults
4.6. Hyperparameter-Sensitivity and Convergence Analysis
4.7. Linking Inspectable Latent Dimensions to Process Variables
5. Discussion
5.1. Analysis of Main Results
5.2. Parameter Efficiency
5.3. Limitations
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| CNN | Convolutional Neural Network |
| KAN | Kolmogorov–Arnold Network |
| CTKAN | CNN–Transformer–KAN |
| TE | Tennessee Eastman (process) |
| MLP | Multi-Layer Perceptron |
| LSTM | Long Short-Term Memory |
| RNN | Recurrent Neural Network |
| PCA | Principal Component Analysis |
| PLS | Partial Least Squares |
| t-SNE | t-distributed Stochastic Neighbor Embedding |
References
- Venkatasubramanian, V.; Rengaswamy, R.; Yin, K.; Kavuri, S.N. A review of process fault detection and diagnosis: Part I. Comput. Chem. Eng. 2003, 27, 293–311. [Google Scholar]
- Yin, S.; Ding, S.X.; Xie, X.; Luo, H. A review on basic data-driven approaches for industrial process monitoring. IEEE Trans. Ind. Electron. 2014, 61, 6418–6428. [Google Scholar] [CrossRef]
- Rakholia, R.; Suarez-Cetrulo, A.L.; Singh, M.; Simon Carbajo, R. Integrating AI and IoT for predictive maintenance in Industry 4.0 manufacturing environments: A practical approach. Information 2025, 16, 737. [Google Scholar] [CrossRef]
- Wang, T.; Dong, J.; Xie, T.; Diallo, D.; Benbouzid, M. A self-learning fault diagnosis strategy based on multi-model fusion. Information 2019, 10, 116. [Google Scholar]
- Daga, A.P.; Garibaldi, L. Machine vibration monitoring for diagnostics through hypothesis testing. Information 2019, 10, 204. [Google Scholar] [CrossRef]
- Calabrese, M.; Cimmino, M.; Fiume, F.; Manfrin, M.; Romeo, L.; Ceccacci, S.; Paolanti, M.; Toscano, G.; Ciandrini, G.; Carrotta, A.; et al. SOPHIA: An event-based IoT and machine learning architecture for predictive maintenance in Industry 4.0. Information 2020, 11, 202. [Google Scholar]
- Fernandes, S.; Antunes, M.; Santiago, A.R.; Barraca, J.P.; Gomes, D.; Aguiar, R.L. Forecasting appliances failures: A machine-learning approach to predictive maintenance. Information 2020, 11, 208. [Google Scholar] [CrossRef]
- Solanes, J.E.; Frances-Falip, A.; Gracia, L.; Munoz, A. Bridging virtual and physical realms in industrial metaverses for enhanced process control. Information 2026, 17, 71. [Google Scholar] [CrossRef]
- Downs, J.J.; Vogel, E.F. A plant-wide industrial process control problem. Comput. Chem. Eng. 1993, 17, 245–255. [Google Scholar] [CrossRef]
- Ku, W.; Storer, R.H.; Georgakis, C. Disturbance detection and isolation by dynamic principal component analysis. Chemom. Intell. Lab. Syst. 1995, 30, 179–196. [Google Scholar] [CrossRef]
- Venkatasubramanian, V.; Rengaswamy, R.; Kavuri, S.N.; Yin, K. A review of process fault detection and diagnosis: Part III. Comput. Chem. Eng. 2003, 27, 327–346. [Google Scholar]
- Zhao, R.; Yan, R.; Chen, Z.; Mao, K.; Wang, P.; Gao, R.X. Deep learning and its applications to machine health monitoring. Mech. Syst. Signal Process. 2019, 115, 213–237. [Google Scholar] [CrossRef]
- Wu, H.; Zhao, J. Deep convolutional neural network model based chemical process fault diagnosis. Comput. Chem. Eng. 2018, 115, 185–197. [Google Scholar] [CrossRef]
- Chen, S.; Yu, J.; Wang, S. One-dimensional convolutional auto-encoder-based feature learning for fault diagnosis of multivariate processes. J. Process Control 2020, 87, 54–67. [Google Scholar]
- Hochreiter, S.; Schmidhuber, J. Long short-term memory. Neural Comput. 1997, 9, 1735–1780. [Google Scholar] [CrossRef] [PubMed]
- Mou, M.; Zhao, X.; Liu, K.; Hui, Y. Bidirectional recurrent neural network-based chemical process fault diagnosis. Ind. Eng. Chem. Res. 2020, 59, 824–834. [Google Scholar]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems; NIPS Foundation: La Jolla, CA, USA, 2017; pp. 5999–6009. [Google Scholar]
- Zhang, K.; Huang, W.; Hou, Y.; Chen, G. Generalized transformer in fault diagnosis of Tennessee Eastman process. Neural Comput. Appl. 2022, 34, 8575–8585. [Google Scholar] [CrossRef]
- Labbaf-Khaniki, M.A.; Manthouri, M.; Ajami, M. Twin Transformer using gated dynamic learnable attention mechanism for fault detection and diagnosis in the Tennessee Eastman process. arXiv 2024, arXiv:2403.10842. [Google Scholar]
- Cao, Y.; Tang, X.; Deng, X.; Wang, P. Fault detection of complicated processes based on an enhanced transformer network with graph attention mechanism. Process Saf. Environ. Prot. 2024, 186, 783–797. [Google Scholar] [CrossRef]
- Zhang, Z.; Xu, M.; Wang, S.; Guo, X.; Gao, J.; Hu, A.P. Sequence-aware vision transformer with feature fusion for fault diagnosis in complex industrial processes. Entropy 2025, 27, 181. [Google Scholar] [CrossRef] [PubMed]
- Chen, H.; Cen, G.; Yang, J.; Si, T.; Cheng, C. Fault diagnosis of the dynamic chemical process based on the optimized CNN-LSTM network. ACS Omega 2022, 7, 34389–34400. [Google Scholar] [CrossRef] [PubMed]
- Kai, T.C.Y.; Saptoro, A.; Putra, Z.A.; Lim, K.H.; Yeo, W.S.; Sunarso, J. Supervised deep learning algorithms for process fault detection and diagnosis under different temporal subsequence length of process data. Appl. Intell. 2025, 55, 883. [Google Scholar] [CrossRef]
- Jang, K.; Pilario, K.E.; Lee, N.; Moon, I.; Na, J. Explainable artificial intelligence for fault diagnosis of industrial processes. IEEE Trans. Ind. Inform. 2025, 21, 4–11. [Google Scholar]
- Cacao, J.; Santos, J.; Antunes, M. Explainable AI for industrial fault diagnosis: A systematic review. J. Ind. Inf. Integr. 2025, 47, 100905. [Google Scholar] [CrossRef]
- Liu, Z.; Wang, Y.; Vaidya, S.; Ruehle, F.; Halverson, J.; Soljacic, M.; Hou, T.Y.; Tegmark, M. KAN: Kolmogorov-Arnold Networks. In Proceedings of the ICLR; ICLR: Rio de Janeiro, Brazil, 2025; pp. 70367–70413. [Google Scholar]
- Kolmogorov, A.N. On the representation of continuous functions of many variables by superposition of continuous functions of one variable and addition. Dokl. Akad. Nauk SSSR 1957, 114, 953–956. [Google Scholar]
- Cabral, T.W.; Gomes, F.V.; de Lima, E.R.; Filho, J.C.S.S.; Meloni, L.G.P. Kolmogorov-Arnold network in the fault diagnosis of oil-immersed power transformers. Sensors 2024, 24, 7585. [Google Scholar] [PubMed]
- Rigas, S.; Papachristou, M.; Sotiropoulos, I.; Alexandridis, G. Explainable fault classification and severity diagnosis in rotating machinery using Kolmogorov-Arnold Networks. Entropy 2025, 27, 403. [Google Scholar] [PubMed]
- Morales, L.; Generoso, A.L.; Attux, R. Exploring Kolmogorov-Arnold networks for unsupervised anomaly detection in industrial processes. Processes 2025, 13, 3672. [Google Scholar]
- Yin, S.; Ding, S.X.; Haghani, A.; Hao, H.; Zhang, P. A comparison study of basic data-driven fault diagnosis and process monitoring methods on the benchmark Tennessee Eastman process. J. Process Control 2012, 22, 1567–1581. [Google Scholar] [CrossRef]
- Zhang, Z.; Zhao, J. A deep belief network based fault diagnosis model for complex chemical processes. Comput. Chem. Eng. 2017, 107, 395–407. [Google Scholar] [CrossRef]
- Chadha, G.S.; Schwung, A. Comparison of deep neural network architectures for fault detection in the Tennessee Eastman process. In Proceedings of the IEEE International Conference on Emerging Technologies and Factory Automation (ETFA); IEEE: Piscataway, NJ, USA, 2017; pp. 1–8. [Google Scholar]
- Siddique, M.F.; Zaman, W.; Khalid, M.; Hamdan, B.; Kim, J.M. A multistage transfer learning framework for intelligent fault diagnosis of rotating machinery under variable operating conditions. Sci. Rep. 2026, 16, 18489. [Google Scholar] [CrossRef] [PubMed]
- Lomov, I.; Lyubimov, M.; Makarov, I.; Zhukov, L.E. Fault detection in Tennessee Eastman process with temporal deep learning models. J. Ind. Inf. Integr. 2021, 23, 100216. [Google Scholar] [CrossRef]
- Khan, A.; Al Farid, F.; Junaid, A.; Siddique, M.F.; Iqbal, A.; Siddique, M.S.; Uddin, J.; Abdul Karim, H.; Husnain, G. Early-warning industrial fault detection based on physics-guided residual learning and calibrated CRNNs. Sci. Rep. 2026, 16, 17488. [Google Scholar] [PubMed]
- Wang, J.; Dong, Z.; Zhang, S. KAN-HyperMP: An enhanced fault diagnosis model for rolling bearings in noisy environments. Sensors 2024, 24, 6448. [Google Scholar] [PubMed]
- Livieris, I.E. C-KAN: A new approach for integrating convolutional layers with Kolmogorov-Arnold networks for time-series forecasting. Mathematics 2024, 12, 3022. [Google Scholar]
- Kingma, D.P.; Ba, J. Adam: A method for stochastic optimization. In Proceedings of the ICLR; ICLR: Rio de Janeiro, Brazil, 2015; pp. 1–15. [Google Scholar]
- Loshchilov, I.; Hutter, F. SGDR: Stochastic gradient descent with warm restarts. In Proceedings of the ICLR; ICLR: Rio de Janeiro, Brazil, 2017; pp. 1–15. [Google Scholar]
- van der Maaten, L.; Hinton, G. Visualizing data using t-SNE. J. Mach. Learn. Res. 2008, 9, 2579–2605. [Google Scholar]













| Method Family | Local Features | Global Dependencies | Inspectable Head | TE Benchmark | Key Limitation |
|---|---|---|---|---|---|
| CNN1D/CAE [13,14] | Yes | Limited | No | Yes | Small receptive field misses long-range dependencies; opaque MLP head |
| RNN/BiLSTM [15,16] | Limited | Yes | No | Yes | Sequential and slow to train, vanishing gradients; opaque head |
| Transformer [18,19,20,21,39] | Weak | Yes | No | Yes | High accuracy but opaque head; attention is not classification-layer transparency |
| CNN–LSTM hybrid [22] | Yes | Yes | No | Yes | Combines CNN and RNN but inherits recurrent training cost; opaque head |
| KAN-based [28,29,30,31,32] | Varies | Usually none | Yes | Rarely | Used in isolation; not combined with a CNN + Transformer encoder for supervised TE classification |
| CTKAN (this work) | Yes | Yes | Yes | Yes | Unifies local and global encoding with an inspectable KAN classification head |
| Process Variables | Time Steps Per Sample | Number of Classes | Training Samples | Validation Samples | Test Samples |
|---|---|---|---|---|---|
| 50 | 20 | 21 (1 normal + 20 faults) | ~104,500 | ~26,200 | ~32,700 |
| Model | Architecture | Hidden Dim | Other Key Hyperparams | Params |
|---|---|---|---|---|
| CNN1D | CNN encoder (2 residual blocks) + MLP head | 128 | kernel = 3, MLP hidden = 128 | 193,941 |
| RNN (BiLSTM) | Two-layer bidirectional LSTM + linear classifier | 64 | 2 layers, bidirectional | 164,117 |
| Transformer | Linear projection + Transformer encoder + linear classifier | 64 | h = 8 heads, 2 layers, FFN dim = 2d | 71,701 |
| CTKAN (Proposed) | CNN encoder + Transformer encoder + 2-layer KAN head | 128 | h = 8, 2 layers, KAN G = 5, head hidden = 64 | 526,101 |
| Model | Params | Accuracy (%) | Macro-F1 (%) | Precision (%) | Recall (%) |
|---|---|---|---|---|---|
| CNN1D | 193,941 | 89.89 ± 0.48 | 90.23 ± 0.47 | 90.99 ± 0.61 | 90.11 ± 0.49 |
| RNN (BiLSTM) | 164,117 | 89.14 ± 0.42 | 89.75 ± 0.42 | 90.29 ± 0.39 | 89.42 ± 0.42 |
| Transformer | 71,701 | 90.74 ± 0.28 | 91.28 ± 0.26 | 91.76 ± 0.29 | 90.99 ± 0.25 |
| CTKAN (Proposed) | 526,101 | 90.85 ± 0.20 | 91.38 ± 0.26 | 91.97 ± 0.25 | 91.11 ± 0.21 |
| Configuration | Trans. | KAN | Params | Macro-F1 (%) |
|---|---|---|---|---|
| CNN + MLP (CNN1D) | ✗ | ✗ | 193,941 | 90.23 ± 0.47 |
| CNN + KAN | ✗ | ✓ | 260,885 | 90.48 ± 0.38 |
| CNN + Trans. + MLP | ✓ | ✗ | 459,413 | 91.21 ± 0.32 |
| CNN + Trans. + KAN (CTKAN) | ✓ | ✓ | 526,101 | 91.38 ± 0.26 |
| Comparison (vs. CTKAN) | ΔMacro-F1 (pp) | 95% CI (pp) | Paired t-Test p | Wilcoxon p | Significant (α = 0.05) |
|---|---|---|---|---|---|
| CNN1D | +1.15 | [0.75, 1.54] | <0.001 | 0.002 | Yes |
| RNN (BiLSTM) | +1.63 | [1.25, 2.01] | <0.001 | 0.002 | Yes |
| Transformer | +0.10 | [−0.14, 0.34] | 0.369 | 0.492 | No |
| CNN + KAN | +0.90 | [0.56, 1.25] | <0.001 | 0.002 | Yes |
| CNN + Trans. + MLP | +0.17 | [−0.09, 0.44] | 0.171 | 0.193 | No |
| CNN + Trans. + MLP (matched) | −0.05 | [−0.45, 0.35] | 0.790 | 0.770 | No |
| Fault Condition | CNN1D | RNN | Transformer | CTKAN | CNN + Trans. + MLP | CNN + KAN |
|---|---|---|---|---|---|---|
| Fault 3/IDV(3) (step change, D feed T) | 91.6 ± 0.7 | 90.4 ± 2.1 | 94.0 ± 0.9 | 94.2 ± 1.0 | 93.6 ± 1.0 | 91.2 ± 1.4 |
| Fault 9/IDV(9) (random variation, D feed T) | 76.2 ± 2.1 | 75.6 ± 2.5 | 81.4 ± 1.5 | 82.5 ± 1.1 | 81.0 ± 1.0 | 75.7 ± 1.3 |
| Fault 15/IDV(15) (sticking valve, condenser cooling) | 28.4 ± 4.3 | 38.1 ± 2.2 | 46.1 ± 1.8 | 39.4 ± 4.0 | 36.6 ± 5.9 | 32.1 ± 4.4 |
| Grid Size G | Params | Macro-F1 (%) | ΔMacro-F1 (pp) | 95% CI (pp) | Paired t-Test p |
|---|---|---|---|---|---|
| 3 | 507,029 | 91.45 ± 0.30 | +0.07 | [−0.21, +0.35] | 0.580 |
| 5 (default) | 526,101 | 91.38 ± 0.26 | — | — | — |
| 7 | 545,173 | 91.37 ± 0.41 | −0.01 | [−0.37, +0.36] | 0.972 |
| 10 | 573,781 | 91.39 ± 0.25 | +0.01 | [−0.17, +0.19] | 0.911 |
| Kernel Size | Params | Macro-F1 (%) | ΔMacro-F1 (pp) | 95% CI (pp) | Paired t-Test p |
|---|---|---|---|---|---|
| 3 (default) | 526,101 | 91.38 ± 0.26 | — | — | — |
| 5 | 637,205 | 91.43 ± 0.37 | +0.05 | [−0.29, +0.39] | 0.732 |
| 7 | 748,309 | 91.16 ± 0.29 | −0.22 | [−0.47, +0.03] | 0.076 |
| Process Variable | Mean Attribution (%) | In Top 5 (of 8 Channels) |
|---|---|---|
| v21 | 21.3 | 8 |
| v49 | 6.0 | 8 |
| v22 | 4.6 | 6 |
| v18 | 3.8 | 6 |
| v11 | 3.4 | 4 |
| v42 | 3.1 | 3 |
| v10 | 2.9 | 1 |
| v48 | 2.8 | 1 |
| v1 | 2.6 | 3 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Wu, Y.; Zhang, M.; Ding, A.; Hua, Y.; Jin, Z.; Dai, Y. CNN–Transformer–KAN: A Hybrid Deep-Learning Framework with an Inspectable KAN Classification Head for Industrial Process Fault Diagnosis. Information 2026, 17, 626. https://doi.org/10.3390/info17070626
Wu Y, Zhang M, Ding A, Hua Y, Jin Z, Dai Y. CNN–Transformer–KAN: A Hybrid Deep-Learning Framework with an Inspectable KAN Classification Head for Industrial Process Fault Diagnosis. Information. 2026; 17(7):626. https://doi.org/10.3390/info17070626
Chicago/Turabian StyleWu, Yujie, Maoyu Zhang, Aoxuan Ding, Yu Hua, Zhehao Jin, and Yiyang Dai. 2026. "CNN–Transformer–KAN: A Hybrid Deep-Learning Framework with an Inspectable KAN Classification Head for Industrial Process Fault Diagnosis" Information 17, no. 7: 626. https://doi.org/10.3390/info17070626
APA StyleWu, Y., Zhang, M., Ding, A., Hua, Y., Jin, Z., & Dai, Y. (2026). CNN–Transformer–KAN: A Hybrid Deep-Learning Framework with an Inspectable KAN Classification Head for Industrial Process Fault Diagnosis. Information, 17(7), 626. https://doi.org/10.3390/info17070626

