VTC-Net: A Semantic Segmentation Network for Ore Particles Integrating Transformer and Convolutional Block Attention Module (CBAM)
Abstract
1. Introduction
2. Methodology for Image Data
2.1. Datasets Collection
2.2. Data Preprocessing
3. Methods
3.1. VTC-Net Architecture
3.1.1. CBAM Attention Mechanism
3.1.2. Transformer Blocks
3.2. Parameter Settings
3.3. Loss Function
3.4. Evaluation Indicators
3.5. Cross-Validation Experiments
4. Discussion of the Results
4.1. Discussion on the Position of Adding CBAM
4.2. Attention Mechanism Comparison Experiment
4.3. Ablation Experiments
4.4. Comparative Experiments
4.5. Feature Visualization and Analysis
5. Conclusions
- (1)
- The VTC-Net model effectively mitigates the common issues of “undersegmentation” and “misjudgment” in traditional methods for complex raw coal images. It achieves optimal segmentation performance (MIoU of 89.90%) on the validation set, significantly outperforming multiple classical segmentation networks. This demonstrates the architecture’s distinct advantage in enhancing segmentation accuracy for multi-scale, highly cohesive coal particles.
- (2)
- Ablation experiments demonstrate that the introduced Transformer module, CBAM module, and BatchNorm layer all significantly enhance model performance. The Transformer’s global context modeling capability and CBAM’s channel-space dual attention mechanism complement each other synergistically, enabling a more comprehensive capture of both long-range dependencies between raw coal samples and local feature details.
- (3)
- The placement of the CBAM module significantly impacts performance. Its optimal placement is at the deep encoder layer (Feat4), where the feature carries richer semantic information, enabling the attention mechanism to more precisely enhance the raw coal target region and weak boundaries.
- (4)
- The feature visualization (Grad-CAM) results demonstrate that, compared to the baseline model, the active regions of VTC-Net align more closely with the actual raw coal contours, with heightened focus on adhered areas, which intuitively explains the performance improvement.
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Wang, W.; Li, Q.; Zhang, D.; Li, H.; Wang, H. A survey of ore image processing based on deep learning. Chin. J. Eng. 2023, 45, 621–631. [Google Scholar] [CrossRef]
- Zhao, S.; Zhan, Y.; Niu, W. EU-Net and ACFS: An effective method for segmenting ore images collected on-site. J. King Saud Univ. Comput. Inf. Sci. 2025, 37, 35. [Google Scholar] [CrossRef] [Scilit]
- Wang, W.; Li, Q.; Xiao, C.; Zhang, D.; Miao, L.; Wang, L. An Improved Boundary-Aware U-Net for Ore Image Semantic Segmentation. Sensors 2021, 21, 2615. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chai, X.; Wu, Z.; Li, W.; Fan, H.; Sun, X.; Xu, J. Image Segmentation Based on the Optimized K-Means Algorithm with the Improved Hybrid Grey Wolf Optimization: Application in Ore Particle Size Detection. Sensors 2025, 25, 2785. [Google Scholar] [CrossRef] [Scilit]
- Long, Y.; Cai, B.; Hu, J.; Hu, W.; Yang, W.; Zhang, W.; Qin, Q. YOLOv8-ORE: An Efficient Ore Segmentation Network based on Adaptive Feature Extraction and Attention-Enhanced Spatial Fusion. Signal Image Video Process. 2025, 19, 1280. [Google Scholar] [CrossRef] [Scilit]
- Zhang, H.; Chen, G.; Li, H. Research on segmentation and reconstruction of overlapping ore contours based on EAM-SOLOv2 and convex hulls. Signal Image Video Process. 2024, 18, 5987–5995. [Google Scholar] [CrossRef] [Scilit]
- Budzan, S.; Buchczik, D.; Pawełczyk, M.; Tůma, J. Combining Segmentation and Edge Detection for Efficient Ore Grain Detection in an Electromagnetic Mill Classification System. Sensors 2019, 19, 1805. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhang, G.; Liu, G.; Zhu, H. Segmentation algorithm of complex ore images based on templates transformation and reconstruction. Int. J. Miner. Metall. Mater. 2011, 18, 385–389. [Google Scholar] [CrossRef] [Scilit]
- Kan, Y. An Image Segmentation Method for Blast Pile Ore in Open-pit Mine Based on U-Net and Improved Watershed Algorithm. Met. Mine 2023, 8, 272. [Google Scholar] [CrossRef]
- Wang, G.; Wang, Z.; Luo, D. Image Segmentation of Adherent Rock Particles Based on FCM and Marked Watershed. J. Sichuan Univ. 2012, 49, 356–360. [Google Scholar] [CrossRef]
- Tang, W.; Wu, Z.; Wang, W.; Pan, Y.; Gan, W. VM-UNet++ research on crack image segmentation based on improved VM-UNet. Sci. Rep. 2025, 15, 8938. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wu, W.; Huang, J.; Zhang, M.; Li, Y.; Yu, Q.; Zhao, Q. MSA-MaxNet: Multi-Scale Attention Enhanced Multi-Axis Vision Transformer Network for Medical Image Segmentation. J. Cell. Mol. Med. 2024, 28, e70315. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Li, F.; Liu, X.; Li, Z. A Two-Stage Framework With Ore-Detect and Segment Anything Model for Ore Particle Segmentation and Size Measurement. IEEE Sens. J. 2025, 25, 11722–11736. [Google Scholar] [CrossRef] [Scilit]
- Liu, Y.; Zhang, Z.; Liu, X.; Wang, L.; Xia, X. Efficient image segmentation based on deep learning for mineral image classification. Adv. Powder Technol. 2021, 32, 3885–3903. [Google Scholar] [CrossRef] [Scilit]
- Xiao, D.; Liu, X.; Le, B.T.; Ji, Z.; Sun, X. An Ore Image Segmentation Method Based on RDU-Net Model. Sensors 2020, 20, 4979. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, W.; Yu, C.; Zhang, T.; Chen, F.; Liu, Y.; Liu, Y.; Wu, Z. Oversized ore segmentation using SAM-enhanced U-Net with self-supervised pre-training and semi-supervised self-training. Expert Syst. Appl. 2025, 285, 127980. [Google Scholar] [CrossRef] [Scilit]
- Fu, Y.; Adams, C. Online particle size analysis on conveyor belts with dense convolutional neural networks. Miner. Eng. 2023, 193, 108019. [Google Scholar] [CrossRef] [Scilit]
- Wang, W.; Li, Q.; Zhang, D.; Fu, J. Image segmentation of adhesive ores based on MSBA-Unet and convex-hull defect detection. Eng. Appl. Artif. Intell. 2023, 123, 106185. [Google Scholar] [CrossRef] [Scilit]
- Liu, X.; Zhang, Y.; Jing, H.; Wang, L.; Zhao, S. Ore image segmentation method using U-Net and ResUnet convolutional networks. RSC Adv. 2020, 10, 9396–9406. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhang, H.; Xiao, D. A high-precision and lightweight ore particle segmentation network for industrial conveyor belt. Expert Syst. Appl. 2025, 273, 126891. [Google Scholar] [CrossRef] [Scilit]
- Wang, W.; Li, Q.; Chen, P.; Zhang, D.; Xiao, C.; Wang, Z. An improved U-Net-based network for multiclass segmentation and category ratio statistics of ore images. Soft Comput. 2024, 28, 4725–4741. [Google Scholar] [CrossRef] [Scilit]
- Liu, X.; Zhang, Y. Ore Image Segmentation Method of Conveyor Belt Based on U-Net and ResUNet Models. J. Northeast. Univ. 2019, 40, 1623–1629. [Google Scholar] [CrossRef]
- Li, H.; Wang, X.; Yang, C.; Xiong, W. Ore image segmentation method based on GAN–UNet. Control Theory Appl. 2021, 38, 1393–1398. [Google Scholar] [CrossRef]
- Zhang, H.; Xiao, D.; He, J.; Wu, D.; Li, Z. RSFA-Net: A High-Performance Lightweight Network for Ore Segmentation and Proportion Detection in Conveyor Belt Images. IEEE Sens. J. 2024, 24, 32508–32518. [Google Scholar] [CrossRef] [Scilit]
- Zhou, C.; Xi, Y.; Sun, X.; Liang, W.; Fang, J.; Wang, G.; Zhang, H. Multiclass Classification of Coal Gangue Under Different Light Sources and Illumination Intensities. Minerals 2025, 15, 921. [Google Scholar] [CrossRef] [Scilit]
- Liu, Y.; Wang, X.; Zhang, Z.; Deng, F. Deep learning based data augmentation for large-scale mineral image recognition and classification. Miner. Eng. 2023, 204, 108411. [Google Scholar] [CrossRef] [Scilit]
- Liang, W.; Sun, X.; Li, Y.; Liu, Y.; Wang, G.; Wang, J.; Zhou, C. Coarse-Grained Ore Distribution on Conveyor Belts With TRCU Neural Networks. IET Image Process. 2025, 19, e70057. [Google Scholar] [CrossRef] [Scilit]
- Li, J.; Wang, X.; Li, J.; Zhang, J.; Ma, G. A generative adversarial learning strategy for spatial inspection of compaction quality. Adv. Eng. Inform. 2024, 62, 102791. [Google Scholar] [CrossRef] [Scilit]
- Wang, S. Effectiveness of traditional augmentation methods for rebar counting using UAV imagery with faster R-CNN and YOLOv10-based transformer architectures. Sci. Rep. 2025, 15, 33702. [Google Scholar] [CrossRef] [Scilit]
- Minh, N.Q.; Huong, N.T.T.; Khanh, P.Q.; Hien, L.P.; Bui, D.T. Impacts of Resampling and Downscaling Digital Elevation Model and Its Morphometric Factors: A Comparison of Hopfield Neural Network, Bilinear, Bicubic, and Kriging Interpolations. Remote Sens. 2024, 16, 819. [Google Scholar] [CrossRef] [Scilit]
- Song, Z.; Yao, H.; Tian, D.; Zhan, G.; Gu, Y. Segmentation method of U-net sheet metal engineering drawing based on CBAM attention mechanism. Artif. Intell. Eng. Des. Anal. Manuf. 2025, 39, e14. [Google Scholar] [CrossRef] [Scilit]
- Yang, Z.; Xu, C.; Li, L. Landslide Detection Based on ResU-Net, Transformer and CBAM Embedding: Two Case Studies in Different Geological Environments. Remote Sens. 2022, 14, 2885. [Google Scholar] [CrossRef] [Scilit]
- Li, Z.; Wan, L.; Wu, Y.; Song, R.; Shao, S.; Wu, H. A Tunnel Secondary Lining Leakage Recognition Model Based on an Improved TransUNet. Appl. Sci. 2025, 15, 10006. [Google Scholar] [CrossRef] [Scilit]
- Zhao, D.; Zhang, W.; Wang, Y. Research on Personnel Image Segmentation Based on MobileNetV2 H-Swish CBAM PSPNet in Search and Rescue Scenarios. Appl. Sci. 2024, 14, 10675. [Google Scholar] [CrossRef] [Scilit]
- Wang, S. Development of approach to an automated acquisition of static street view images using transformer architecture for analysis of Building characteristics. Sci. Rep. 2025, 15, 29062. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhao, Q.; Liu, F.; Song, Y.; Fan, X.; Wang, Y.; Yao, Y.; Mao, Q.; Zhao, Z. Predicting Respiratory Rate from Electrocardiogram and Photoplethysmogram Using a Transformer-Based Model. Bioengineering 2023, 10, 1024. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Mao, A.; Huang, E.; Gan, H.; Parkes, R.S.V.; Xu, W.; Liu, K. Cross-Modality Interaction Network for Equine Activity Recognition Using Imbalanced Multi-Modal Data. Sensors 2021, 21, 5818. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Li, J.; Liu, K.; Hu, Y.; Zhang, H.; Heidari, A.A.; Chen, H.; Zhang, W.; Algarni, A.D.; Elmannai, H. Eres-UNet++: Liver CT image segmentation based on high-efficiency channel attention and Res-UNet+. Comput. Biol. Med. 2022, 158, 106501. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, X.; Feng, M.; Tang, X.; Peng, T.; Li, Z.; Yang, C. Ore image segmentation Based on Multiscale Parallel Efficient Channel Attention U-Network. IFAC Pap. 2024, 58, 101–106. [Google Scholar] [CrossRef] [Scilit]
- Wan, T.; Rao, Y.; Jin, X.; Wang, F.; Zhang, T.; Shu, Y.; Li, S. Improved U-Net for Growth Stage Recognition of In-Field Maize. Agronomy 2023, 13, 1523. [Google Scholar] [CrossRef] [Scilit]
- Tang, H.; Wang, H.; Wang, L.; Cao, C.; Nie, Y.; Liu, S. An Improved Mineral Image Recognition Method Based on Deep Learning. JOM 2023, 75, 2590–2602. [Google Scholar] [CrossRef] [Scilit]
- Zhao, X.; Yang, Z.; Yan, X. Coal Transportation area detection algorithm of belt conveyor based on semantic segmentation. Comput. Appl. Softw. 2024, 41, 56–61. [Google Scholar] [CrossRef]
- Li, F.; Jin, W.; Fan, C.; Zou, L.; Chen, Q.; Li, X.; Jiang, H.; Liu, Y. PSANet: Pyramid Splitting and Aggregation Network for 3D Object Detection in Point Cloud. Sensors 2020, 21, 136. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Xie, E.; Wang, W.; Yu, Z.; Anandkumar, A.; Alvarez, J.M.; Luo, P. SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers. Adv. Neural Inf. Process. Syst. 2021, 34, 12077–12090. [Google Scholar] [CrossRef] [Scilit]
- Zheng, S.; Lu, J.; Zhao, H.; Zhu, X.; Luo, Z.; Wang, Y.; Fu, Y.; Feng, J.; Xiang, T.; Torr, P.H.; et al. Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021; pp. 6877–6886. [Google Scholar] [CrossRef] [Scilit]
- Venkatachalam, C.; Shah, P.; Balajee, A.; Karthikc, R.K.M.; Yogesh, K.S.; Roy, A. Advanced Grape Leaf Disease Diagnosis Using EfficientNetV2L with Data Augmentation and Grad-CAM Visualization in Precision Agriculture. Procedia Comput. Sci. 2025, 260, 332–340. [Google Scholar] [CrossRef] [Scilit]







| Parameters | Values | Parameters | Values |
|---|---|---|---|
| Input image size | 512 | Optimizer | Adam |
| Epochs | 100 | 0.9 | |
| Batch size | 8 | Learning Rate | 1 × 10−4 |
| Fold | MIoU (%) | MPA (%) | Acc (%) |
|---|---|---|---|
| Fold1 | 85.48 | 87.48 | 96.44 |
| Fold2 | 84.82 | 86.82 | 96.27 |
| Fold3 | 83.55 | 85.95 | 95.98 |
| Fold4 | 84.95 | 86.95 | 96.21 |
| Fold5 | 84.54 | 85.54 | 96.20 |
| Average | 84.67 | 86.55 | 96.22 |
| Standard deviation | 0.64 | 0.70 | 0.15 |
| Dataset | MIoU (%) | MPA (%) | Acc (%) |
|---|---|---|---|
| Fold1 | 85.48 | 87.48 | 96.44 |
| Test Set | 86.02 | 88.21 | 97.43 |
| Add Location | Feat1 | Feat2 | Feat3 | Feat4 | MIoU (%) | MPA (%) | Acc (%) |
|---|---|---|---|---|---|---|---|
| 1 | √ | 85.80 | 92.91 | 95.90 | |||
| 2 | √ | 85.95 | 92.53 | 95.62 | |||
| 3 | √ | 87.60 | 93.47 | 96.14 | |||
| 4 | √ | 88.30 | 93.88 | 96.35 | |||
| 5 | √ | √ | 87.61 | 93.41 | 96.13 | ||
| 6 | √ | √ | 88.01 | 93.69 | 96.25 | ||
| 7 | √ | √ | 88.21 | 93.85 | 96.32 | ||
| 8 | √ | √ | √ | 87.98 | 93.64 | 96.25 | |
| 9 | √ | √ | √ | √ | 88.11 | 93.72 | 96.29 |
| Models | MIoU (%) | MPA (%) | Acc (%) |
|---|---|---|---|
| UNet + SE | 87.15 | 93.27 | 96.00 |
| UNet + ECA | 87.08 | 93.12 | 95.98 |
| UNet + CA | 88.01 | 93.64 | 96.26 |
| UNet + CBAM | 88.30 | 93.88 | 96.35 |
| Models | MIoU (%) | MPA (%) | Acc (%) |
|---|---|---|---|
| UNet | 86.35 | 92.67 | 95.76 |
| UNet + BatchNorm | 87.43 | 93.39 | 96.12 |
| UNet + Transformer | 88.64 | 94.34 | 96.42 |
| UNet + CBAM | 88.30 | 93.88 | 96.35 |
| UNet + Transformer + CBAM | 88.73 | 94.18 | 96.46 |
| VTC-Net | 89.90 | 94.78 | 96.80 |
| Models | MIoU (%) | MPA (%) | Acc (%) | Params | GFLOPs | Inference Speed (ms) |
|---|---|---|---|---|---|---|
| UNet | 86.35 | 92.67 | 95.76 | 31.23 | 220.72 | 60.28 |
| DeepLabV3 | 82.81 | 90.61 | 94.23 | 36.07 | 100.88 | 23.50 |
| PSPNet | 85.36 | 91.78 | 94.74 | 49.13 | 94.41 | 30.30 |
| PSANet | 85.33 | 91.94 | 94.75 | 39.69 | 55.69 | 25.06 |
| SegFormer | 82.49 | 87.79 | 92.55 | 3.72 | 6.78 | 19.78 |
| SETR | 68.08 | 77.71 | 88.60 | 64.56 | 230.78 | 69.58 |
| VTC-Net | 89.90 | 94.78 | 96.80 | 50.14 | 225.47 | 64.54 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Wu, Y.; Liang, W.; Fang, J.; Zhou, C.; Sun, X. VTC-Net: A Semantic Segmentation Network for Ore Particles Integrating Transformer and Convolutional Block Attention Module (CBAM). Sensors 2026, 26, 787. https://doi.org/10.3390/s26030787
Wu Y, Liang W, Fang J, Zhou C, Sun X. VTC-Net: A Semantic Segmentation Network for Ore Particles Integrating Transformer and Convolutional Block Attention Module (CBAM). Sensors. 2026; 26(3):787. https://doi.org/10.3390/s26030787
Chicago/Turabian StyleWu, Yijing, Weinong Liang, Jiandong Fang, Chunxia Zhou, and Xiaolu Sun. 2026. "VTC-Net: A Semantic Segmentation Network for Ore Particles Integrating Transformer and Convolutional Block Attention Module (CBAM)" Sensors 26, no. 3: 787. https://doi.org/10.3390/s26030787
APA StyleWu, Y., Liang, W., Fang, J., Zhou, C., & Sun, X. (2026). VTC-Net: A Semantic Segmentation Network for Ore Particles Integrating Transformer and Convolutional Block Attention Module (CBAM). Sensors, 26(3), 787. https://doi.org/10.3390/s26030787

