CA-MC-Transformer: An Operating Condition-Adaptive and Multi-Scale Convolution-Enhanced Transformer Architecture for Furnace Temperature Prediction
Abstract
1. Introduction
- A hierarchical agglomerative clustering algorithm based on the WDTW distance is developed for operating condition clustering, by which unsupervised fine-grained grouping of historical operational data is achieved, and high-quality physical prior labels are provided for subsequent deep learning networks.
- A multi-scale dilated causal convolution architecture is designed to extract local high-dimensional features across different time resolutions while ensuring causality. In the feature fusion stage, a soft attention mechanism is designed to deeply fuse the extracted multi-scale convolutional features with operating condition embedding vectors. Attention weight distributions specific to different operating conditions can be adaptively learned by this mechanism, and the feature representation and robustness of the method in complex and dynamic scenarios are fundamentally enhanced.
- For industrial regression tasks, the decoder module of the traditional Transformer is discarded, and a prediction network based on an encoder-only architecture is constructed. Not only is the cumulative error introduced by autoregressive decoding avoided by this network, but also the computational complexity is significantly reduced to enhance real-time performance.
2. Process Analysis of Aluminum Melting
3. The Proposed CA-MC-Transformer Model
3.1. Time Feature Extraction Module Based on Multi-Scale Convolution
3.2. Hierarchical Clustering-Based Operating Condition Adaptation Module
3.2.1. Weighted Dynamic Time Warping Distance
- Boundary condition. The dynamic warping path must satisfy that the starting and ending points of the sequences are located at the diagonally opposite corners. That is, and .
- Monotonicity. The path indices must increase monotonically, and the temporal order cannot be reversed, ensuring that the alignment process conforms to temporal causality. That is, given and , they must satisfy and .
- Continuity. Each step of the path can only move to adjacent grid points (assuming the starting point is at the top-left corner, the movement can be to the right, downward, or to the bottom-right), with a maximum step size of 1. That is, .
3.2.2. Hierarchical Agglomerative Clustering Based on the WDTW Distance
3.2.3. Soft Attention Mechanism
3.3. Transformer Encoder Module
4. Industrial Case
4.1. Experimental Setup
4.2. Results and Analysis
4.2.1. Comparative Experiments
4.2.2. Ablation Experiments
- MC-Transformer. The operating condition-aware embedding CA module constructed based on agglomerative hierarchical clustering is eliminated in this variant, and pure temporal data are solely relied on for feature extraction.
- CA-Transformer. In this variant, the front-end multi-scale convolutional layer (MC) is removed, and the original sliding window sequence and operating conditions are directly embedded into the encoder after being stacked.
- CA-MC-Transformer (without SA). In this variant, the soft attention mechanism is removed, and the most basic feature concatenation is employed for feature fusion.
- Transformer (Encoder). The model is degraded into a standard Transformer encoder model in this variant by simultaneously removing three major modules, namely, operating condition adaptation, soft attention, and multi-scale convolution.
4.3. Discussion
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Rohatgi, P.; Weiss, D.; Srivatsan, T.S.; Ghaderi, O.; Zare, M. Solidification Processing of Aluminum Alloy Metal Matrix Composites for Use in Transportation Applications. Metall. Mater. Trans. A 2024, 55, 4867–4881. [Google Scholar] [CrossRef] [Scilit]
- Ogawa, F.; Osada, N.; Itoh, T.; Sakane, M. Creep Deformation and Rupture Life Characteristics of High-Purity Aluminum for High-Power Electronic Devices. J. Mater. Eng. Perform. 2025, 34, 7410–7425. [Google Scholar] [CrossRef] [Scilit]
- Dion, L.; Kiss, L.I.; Poncsák, S.; Lagacé, C.L. Prediction of Low-Voltage Tetrafluoromethane Emissions Based on the Operating Conditions of an Aluminium Electrolysis Cell. JOM 2016, 68, 2472–2482. [Google Scholar] [CrossRef] [Scilit]
- Wang, J.; Xu, P.; Yan, H.; Zhou, J.; Li, S.; Gui, G.; Li, W. Burner effects on melting process of regenerative aluminum melting furnace. Trans. Nonferr. Met. Soc. China 2013, 23, 3125–3136. [Google Scholar] [CrossRef] [Scilit]
- Yan, H.; Xie, H.; Zheng, W.; Liu, L. Numerical simulation of combustion and melting process in an aluminum melting furnace: A study on optimizing stacking mode. Appl. Therm. Eng. 2024, 245, 122840. [Google Scholar] [CrossRef] [Scilit]
- Luo, Y.; Dai, J.; Chen, X.; Liu, Y. A Hybrid Modeling Method for Aluminum Smelting Process Based on a Hybrid Strategy-Based Sparrow Search Algorithm. IEEE Access 2022, 10, 101149–101159. [Google Scholar] [CrossRef] [Scilit]
- Caterini, A.L.; Chang, D.E. Recurrent Neural Networks. In Deep Neural Networks in a Mathematical Framework; Caterini, A.L., Chang, D.E., Eds.; Springer International Publishing: Cham, Switzerland, 2018; pp. 59–79. [Google Scholar] [CrossRef] [Scilit]
- Salehinejad, H.; Sankar, S.; Barfett, J.; Colak, E.; Valaee, S. Recent Advances in Recurrent Neural Networks. arXiv 2018, arXiv:1801.01078. [Google Scholar] [CrossRef] [Scilit]
- Pascanu, R.; Gulcehre, C.; Cho, K.; Bengio, Y. How to Construct Deep Recurrent Neural Networks. arXiv 2014, arXiv:1312.6026. [Google Scholar] [CrossRef] [Scilit]
- Lin, Y.; Koprinska, I.; Rana, M. Temporal Convolutional Attention Neural Networks for Time Series Forecasting. In Proceedings of the 2021 International Joint Conference on Neural Networks (IJCNN), Shenzhen, China, 18–22 July 2021; pp. 1–8. [Google Scholar] [CrossRef] [Scilit]
- Zheng, J.; Ma, L.; Wu, Y.; Ye, L.; Shen, F. Nonlinear Dynamic Soft Sensor Development with a Supervised Hybrid CNN-LSTM Network for Industrial Processes. ACS Omega 2022, 7, 16653–16664. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Liu, J.; He, J.; Tang, Z.; Xie, Y.; Gui, W.; Ma, T.; Jahanshahi, H.; Aly, A.A. Frame-Dilated Convolutional Fusion Network and GRU-Based Self-Attention Dual-Channel Network for Soft-Sensor Modeling of Industrial Process Quality Indexes. IEEE Trans. Syst. Man Cybern. Syst. 2022, 52, 5989–6002. [Google Scholar] [CrossRef] [Scilit]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, L.; Polosukhin, I. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA, 4–9 December 2017; pp. 5998–6008. [Google Scholar]
- Wu, H.; Xu, J.; Wang, J.; Long, M. Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting. arXiv 2022, arXiv:2106.13008. [Google Scholar] [CrossRef] [Scilit]
- Nie, Y.; Nguyen, N.H.; Sinthong, P.; Kalagnanam, J. A Time Series is Worth 64 Words: Long-term Forecasting with Transformers. arXiv 2023, arXiv:2211.14730. [Google Scholar] [CrossRef] [Scilit]
- Yuan, X.; Li, L.; Shardt, Y.A.W.; Wang, Y.; Yang, C. Deep Learning With Spatiotemporal Attention-Based LSTM for Industrial Soft Sensor Model Development. IEEE Trans. Ind. Electron. 2021, 68, 4404–4414. [Google Scholar] [CrossRef] [Scilit]
- Lui, C.F.; Liu, Y.; Xie, M. A Supervised Bidirectional Long Short-Term Memory Network for Data-Driven Dynamic Soft Sensor Modeling. IEEE Trans. Instrum. Meas. 2022, 71, 2504713. [Google Scholar] [CrossRef] [Scilit]
- Xu, Z.; Xu, N.; Wang, K.; Yuan, X.; Wang, Y.; Yang, C.; Gui, W.; Cheng, S.; Ye, L. A sampling interval-adaptive transformer for industrial time sequence modeling with heterogeneou s sampling rates in quality prediction. Eng. Appl. Artif. Intell. 2025, 162, 112374. [Google Scholar] [CrossRef] [Scilit]
- Ma, S.; Li, Y.; Luo, D.; Song, T. Temperature Prediction of Medium Frequency Furnace Based on Transformer Model. In Proceedings of the Neural Computing for Advanced Applications; Zhang, H., Chen, Y., Chu, X., Zhang, Z., Hao, T., Wu, Z., Yang, Y., Eds.; Springer: Singapore, 2022; pp. 463–476. [Google Scholar] [CrossRef] [Scilit]
- Chen, Q.; Cai, C.; Chen, Y.; Zhou, X.; Zhang, D.; Peng, Y. TemproNet: A transformer-based deep learning model for seawater temperature prediction. Ocean Eng. 2024, 293, 116651. [Google Scholar] [CrossRef] [Scilit]
- Wen, Y.; Xu, P.; Li, Z.; Xu, W.; Wang, X. RPConvformer: A novel Transformer-based deep neural networks for traffic flow prediction. Expert Syst. Appl. 2023, 218, 119587. [Google Scholar] [CrossRef] [Scilit]
- He, Y.L.; Bai, Z.H.; Xu, Y.; Zhu, Q.X.; Li, L. Patch-Decomposition-Enhanced TCN with Transformer for Soft Sensor Modeling. IEEE Sens. J. 2025, 25, 42364–42371. [Google Scholar] [CrossRef] [Scilit]
- Yang, X.; Chen, D.; Huang, J.; Wu, X.; Chen, Z.; Li, Q. Remaining Useful Life Prediction Under Multiple Operating Conditions Based on a Novel Dual-Layer Temporal Convolutional Network. IEEE Sens. J. 2025, 25, 1900–1911. [Google Scholar] [CrossRef] [Scilit]
- Hua, C.; Wu, J.; Li, J.; Guan, X. Silicon content prediction and industrial analysis on blast furnace using support vector regression combined with clustering algorithms. Neural Comput. Appl. 2017, 28, 4111–4121. [Google Scholar] [CrossRef] [Scilit]
- Duan, Y.; Dai, J.; Luo, Y.; Chen, G.; Cai, X. A Dynamic Time Warping Based Locally Weighted LSTM Modeling for Temperature Prediction of Recycled Aluminum Smelting. IEEE Access 2023, 11, 36980–36992. [Google Scholar] [CrossRef] [Scilit]
- McQueen, J.B. Some methods of classification and analysis of multivariate observations. In Proceedings of the 5th Berkeley Symposium on Mathematical Statistics and Probability; University of California Press: Berkeley, CA, USA, 1967; pp. 281–297. [Google Scholar]
- Kaufman, L.; Rousseeuw, P.J. Finding Groups in Data: An Introduction to Cluster Analysis; John Wiley & Sons: Hoboken, NJ, USA, 2009. [Google Scholar]
- Paparrizos, J.; Gravano, L. k-Shape: Efficient and Accurate Clustering of Time Series. In Proceedings of the 2015 ACM SIGMOD International Conference on Management of Data, Melbourne, VIC, Australia, 31 May–4 June 2015; pp. 1855–1870. [Google Scholar] [CrossRef] [Scilit]
- Peizhuang, W. Pattern Recognition with Fuzzy Objective Function Algorithms (James C. Bezdek). SIAM Rev. 1983, 25, 442. [Google Scholar] [CrossRef] [Scilit]
- Johnson, S.C. Hierarchical clustering schemes. Psychometrika 1967, 32, 241–254. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hochreiter, S.; Schmidhuber, J. Long Short-Term Memory. Neural Comput. 1997, 9, 1735–1780. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Bai, S.; Kolter, J.Z.; Koltun, V. An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling. arXiv 2018, arXiv:1803.01271. [Google Scholar] [CrossRef] [Scilit]
- Ghimire, S.; Yaseen, Z.M.; Farooque, A.A.; Deo, R.C.; Zhang, J.; Tao, X. Streamflow prediction using an integrated methodology based on convolutional neural network and long short-term memory networks. Sci. Rep. 2021, 11, 17497. [Google Scholar] [CrossRef] [Scilit] [PubMed]








| No. | Variable Name | Value Range |
|---|---|---|
| 1 | Furnace temperature | 102.70∼1148.00 °C |
| 2 | Material temperature | 102.7∼1148 °C |
| 3 | Furnace pressure | −52.77∼129.56 Pa |
| 4 | Combustion air temperature | 14∼43.6 °C |
| 5 | 12# Combustion air flow | 1012.66∼2065.07 Nm3/h |
| 6 | 34# Combustion air flow | 286.18∼7132.40 Nm3/h |
| 7 | 12# Combustion air differential pressure | 1424.41∼5631.80 Pa |
| 8 | 34# Combustion air differential pressure | 10.13∼6123.48 Pa |
| 9 | 12# Gas flow | 0∼443.09 Nm3/h |
| 10 | 34# Gas flow | 0∼503.00 Nm3/h |
| 11 | 12# Air–fuel ratio | 4.36∼31.46 |
| 12 | 34# Air–fuel ratio | 2.14∼58.19 |
| 13 | 12# Burner reversing time | 0∼41.00 s |
| 14 | 34# Burner reversing time | 0∼39.00 s |
| Operating Condition Label | Number of Samples | Proportion |
|---|---|---|
| Slagging | 83 | 9.5% |
| Cooling | 114 | 13.1% |
| Heating | 132 | 15.1% |
| Feeding | 349 | 40.0% |
| Holding | 195 | 22.3% |
| Algorithm Name | Silhouette Coefficient | CH Index | DBI |
|---|---|---|---|
| K-means | 0.425 | 1342.518 | 1.582 |
| K-medoids | 0.461 | 1485.247 | 1.415 |
| K-Shape | 0.584 | 1920.863 | 1.104 |
| FCM | 0.448 | 1405.691 | 1.512 |
| AHC (Euclidean distance) | 0.472 | 1512.355 | 1.386 |
| AHC-WDTW | 0.675 | 2354.129 | 0.842 |
| Model Name | RMSE | MAE | MAPE | R2 | Params FLOPs |
|---|---|---|---|---|---|
| LSTM | 18.0093 ± 0.8715 | 11.6427 ± 0.6934 | 0.0212 ± 0.0016 | 0.9622 ± 0.0012 | 1.21 M 0.22 G |
| TCN | 15.5780 ± 1.0213 | 10.2643 ± 0.4256 | 0.0193 ± 0.0008 | 0.9717 ± 0.0023 | 1.85 M 0.35 G |
| CNN-LSTM | 14.0972 ± 0.5581 | 9.5576 ± 0.7321 | 0.0165 ± 0.0014 | 0.9769 ± 0.0017 | 3.22 M 0.78 G |
| Transformer (Encoder-Decoder) | 11.2671 ± 0.9427 | 7.7080 ± 0.2618 | 0.0142 ± 0.0011 | 0.9852 ± 0.0009 | 14.63 M 2.86 G |
| Autoformer | 10.0065 ± 0.6234 | 7.8894 ± 0.5142 | 0.0127 ± 0.0005 | 0.9883 ± 0.0015 | 11.57 M 2.35 G |
| CA-MC-Transformer | 6.5558 ± 0.3872 | 5.2193 ± 0.2549 | 0.0096 ± 0.0007 | 0.9950 ± 0.0006 | 22.15 M 4.12 G |
| Model Name | RMSE | MAE | MAPE | R2 | Params FLOPs |
|---|---|---|---|---|---|
| MC-Transformer | 10.7590 ± 0.4156 | 7.4371 ± 0.2984 | 0.0137 ± 0.0012 | 0.9865 ± 0.0018 | 12.50 M 2.10 G |
| CA-Transformer | 9.7556 ± 0.3892 | 7.5280 ± 0.4135 | 0.0140 ± 0.0017 | 0.9889 ± 0.0015 | 11.20 M 1.90 G |
| CA-MC-Transformer (without SA) | 8.0643 ± 0.5410 | 6.4025 ± 0.3021 | 0.0117 ± 0.0011 | 0.9924 ± 0.0012 | 15.00 M 2.80 G |
| Transformer (Encoder) | 13.1914 ± 0.3127 | 9.1082 ± 0.2145 | 0.0168 ± 0.0008 | 0.9797 ± 0.0021 | 8.50 M 1.50 G |
| CA-MC-Transformer | 6.5558 ± 0.3872 | 5.2193 ± 0.2549 | 0.0096 ± 0.0007 | 0.9950 ± 0.0006 | 22.15 M 4.12 G |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Dai, J.; Chen, Z.; Li, S.; Wu, T. CA-MC-Transformer: An Operating Condition-Adaptive and Multi-Scale Convolution-Enhanced Transformer Architecture for Furnace Temperature Prediction. Electronics 2026, 15, 3784. https://doi.org/10.3390/electronics15173784
Dai J, Chen Z, Li S, Wu T. CA-MC-Transformer: An Operating Condition-Adaptive and Multi-Scale Convolution-Enhanced Transformer Architecture for Furnace Temperature Prediction. Electronics. 2026; 15(17):3784. https://doi.org/10.3390/electronics15173784
Chicago/Turabian StyleDai, Jiayang, Zhen Chen, Shenwang Li, and Thomas Wu. 2026. "CA-MC-Transformer: An Operating Condition-Adaptive and Multi-Scale Convolution-Enhanced Transformer Architecture for Furnace Temperature Prediction" Electronics 15, no. 17: 3784. https://doi.org/10.3390/electronics15173784
APA StyleDai, J., Chen, Z., Li, S., & Wu, T. (2026). CA-MC-Transformer: An Operating Condition-Adaptive and Multi-Scale Convolution-Enhanced Transformer Architecture for Furnace Temperature Prediction. Electronics, 15(17), 3784. https://doi.org/10.3390/electronics15173784

