A Weakly Supervised Segmentation Algorithm Based on Local–Global Class Labelling Comparison
Abstract
1. Introduction
2. Related Work
2.1. State-of-the-Art Research on Fully Supervised Semantic Segmentation Based on Deep Learning
2.2. State-of-the-Art Research on Weakly Supervised Semantic Segmentation Based on Deep Learning
3. Method
3.1. Local–Global Class Labelling Comparison Module
3.2. Class-Aware Stimulus Module
3.3. Feature Fusion Class-Aware Activation Maps
4. Experiments
4.1. Experimental Environment and Dataset
4.2. Experimental Evaluation Indicators
4.3. Comparative Analysis of Experimental Results
4.4. Ablation Experiments
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Yan, L.; Chen, J.; Tang, Y. TSD-CAM: Transformer-based self distillation with CAM similarity for weakly supervised semantic segmentation. J. Electron. Imaging 2024, 33, 023029. [Google Scholar] [CrossRef] [Scilit]
- Long, J.; Shelhamer, E.; Darrell, T. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2015; pp. 3431–3440. [Google Scholar]
- Zhou, B.; Khosla, A.; Lapedriza, A.; Oliva, A.; Torralba, A. Learning deep features for discriminative localization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2016; pp. 2921–2929. [Google Scholar]
- Yoon, S.H.; Kwon, H.; Kim, H.; Yoon, K.J. Class tokens infusion for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2024; pp. 3595–3605. [Google Scholar]
- Wang, J.; Dai, T.; Zhang, B.; Yu, S.; Lim, E.G.; Xiao, J. Class token as proxy: Optimal transport-assisted proxy learning for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2025; pp. 21645–21654. [Google Scholar]
- Hanna, J.; Borth, D. Know your attention maps: Class-specific token masking for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2025; pp. 23763–23772. [Google Scholar]
- Xie, E.; Wang, W.; Yu, Z.; Anandkumar, A.; Alvarez, J.M.; Luo, P. SegFormer: Simple and efficient design for semantic segmentation with transformers. Adv. Neural Inf. Process. Syst. 2021, 34, 12077–12090. [Google Scholar]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. Adv. Neural Inf. Process. Syst. 2017, 30, 30–35. [Google Scholar]
- Zhang, J.; Liu, R.; Shi, H.; Yang, K.; Reiß, S.; Peng, K.; Fu, H.; Wang, K.; Stiefelhagen, R. Delivering arbitrary-modal semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2023; pp. 1136–1147. [Google Scholar]
- Lee, J.; Yi, J.; Shin, C.; Yoon, S. Bbam: Bounding box attribution map for weakly supervised semantic and instance segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2021; pp. 2643–2652. [Google Scholar]
- Khoreva, A.; Benenson, R.; Hosang, J.; Hein, M.; Schiele, B. Simple does it: Weakly supervised instance and semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2017; pp. 876–885. [Google Scholar]
- Song, C.; Huang, Y.; Ouyang, W.; Wang, L. Box-driven class-wise region masking and filling rate guided loss for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2019; pp. 3136–3145. [Google Scholar]
- Bearman, A.; Russakovsky, O.; Ferrari, V.; Fei-Fei, L. What’s the point: Semantic segmentation with point supervision. In ECCV; Springer: Berlin/Heidelberg, Germany, 2016; pp. 549–565. [Google Scholar]
- Wei, J.; Lin, G.; Yap, K.H.; Hung, T.Y.; Xie, L. Multi-path region mining for weakly supervised 3D semantic segmentation on point clouds. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2020; pp. 4384–4393. [Google Scholar]
- Zhang, X.; Zhu, L.; He, H.; Jin, L.; Lu, Y. Scribble hides class: Promoting scribble-based weakly-supervised semantic segmentation with its class label. In Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada, 20–27 February 2024; Volume 38, pp. 7332–7340. [Google Scholar]
- Su, H.; Ye, Y.; Hua, W.; Cheng, L.; Song, M. SASFormer: Transformers for Sparsely Annotated Semantic Segmentation. In Proceedings of the IEEE International Conference on Multimedia and Expo (ICME); IEEE: New York, NY, USA, 2023. [Google Scholar]
- Englebert, A.; Cornu, O.; Vleeschouwer, C.D. Poly-cam: High resolution class activation map for convolutional neural networks. Mach. Vis. Appl. 2024, 35, 89. [Google Scholar] [CrossRef] [Scilit]
- Wei, Y.; Feng, J.; Liang, X.; Cheng, M.M.; Zhao, Y.; Yan, S. Object region mining with adversarial erasing: A simple classification to semantic segmentation approach. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2017; pp. 1568–1576. [Google Scholar]
- Kolesnikov, A.; Lampert, C.H. Seed, expand and constrain: Three principles for weakly-supervised image segmentation. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, 11–14 October 2016, Proceedings, Part IV 14; Springer International Publishing: Cham, Switzerland, 2016; pp. 695–711. [Google Scholar]
- Huang, Z.; Wang, X.; Wang, J.; Liu, W.; Wang, J. Weakly-supervised semantic segmentation network with deep seeded region growing. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2018; pp. 7014–7023. [Google Scholar]
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning, Online, 18–24 July 2021; PmLR: Cambridge, MA, USA, 2021; pp. 8748–8763. [Google Scholar]
- Lee, S.; Lee, M.; Lee, J.; Shim, H. Railroad is not a train: Saliency as pseudo-pixel supervision for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2021; pp. 5495–5505. [Google Scholar]
- Lin, Y.; Chen, M.; Wang, W.; Wu, B.; Li, K.; Lin, B.; Liu, H.; He, X. Clip is also an efficient segmenter: A text-driven approach for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2023; pp. 15305–15314. [Google Scholar]
- Selvaraju, R.R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; Batra, D. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE International Conference on Computer Vision; IEEE: New York, NY, USA, 2017; pp. 618–626. [Google Scholar]
- Murugesan, B.; Hussain, R.; Bhattacharya, R.; Ben Ayed, I.; Dolz, J. Prompting classes: Exploring the power of prompt class learning in weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision; IEEE: New York, NY, USA, 2024; pp. 291–302. [Google Scholar]
- Chen, Z.; Sun, Q. Extracting class activation maps from non-discriminative features as well. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2023; pp. 3135–3144. [Google Scholar]
- Lai, Q.; Vong, C.M.; Chen, C. Weakly Supervised Semantic Segmentation via Dual-Stream Contrastive Learning of Cross-Image Contextual Information. In Proceedings of the IEEE Transactions on Industrial Informatics; IEEE: New York, NY, USA, 2024. [Google Scholar]
- Wu, Y.; Li, X.; Dai, S.; Li, J.; Liu, T.; Xie, S. Hierarchical Semantic Contrast for Weakly Supervised Semantic Segmentation. In Proceedings of the IJCAI—International Joint Conference on Artificial Intelligence, Macao, China, 19–25 August 2023; pp. 1542–1550. [Google Scholar]
- Mao, X.; Qi, G.; Chen, Y.; Li, X.; Duan, R.; Ye, S.; He, Y.; Xue, H. Towards robust vision transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2022; pp. 12042–12051. [Google Scholar]
- Gao, W.; Wan, F.; Pan, X.; Peng, Z.; Tian, Q.; Han, Z.; Zhou, B.; Ye, Q. Ts-cam: Token semantic coupled attention map for weakly supervised object localization. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2021; pp. 2886–2895. [Google Scholar]
- Li, R.; Mai, Z.; Zhang, Z.; Jang, J.; Sanner, S. Transcam: Transformer attention-based cam refinement for weakly supervised semantic segmentation. J. Vis. Commun. Image Represent. 2023, 92, 103800. [Google Scholar] [CrossRef] [Scilit]
- Xu, L.; Ouyang, W.; Bennamoun, M.; Boussaid, F.; Xu, D. Multi-class token transformer for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2022; pp. 4310–4319. [Google Scholar]
- Rossetti, S.; Zappia, D.; Sanzari, M.; Schaerf, M.; Pirri, F. Max pooling with vision transformers reconciles class and shape in weakly supervised semantic segmentation. In European Conference on Computer Vision; Springer Nature: Cham, Switzerland, 2022; pp. 446–463. [Google Scholar]
- Ru, L.; Zhan, Y.; Yu, B.; Du, B. Learning affinity from attention: End-to-end weakly-supervised semantic segmentation with transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2022; pp. 16846–16855. [Google Scholar]
- Ru, L.; Zheng, H.; Zhan, Y.; Du, B. Token contrast for weakly-supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2023; pp. 3093–3102. [Google Scholar]
- Yin, X.; Im, W.; Min, D.; Huo, Y.; Pan, F.; Yoon, S.E. Fine-Grained Background Representation for Weakly Supervised Semantic Segmentation. IEEE Trans. Circuits Syst. Video Technol. 2024, 34, 11739–11750. [Google Scholar] [CrossRef] [Scilit]
- Qin, Y.; Pu, N.; Wu, H.; Sebe, N. Margin-aware Noise-robust Contrastive Learning for Partially View-aligned Problem. ACM Trans. Knowl. Discov. From Data 2025, 19, 1–20. [Google Scholar] [CrossRef] [Scilit]
- Tang, J.; Cheng, K.; Wei, L.; Zhan, Y. Inter-image Token Relation Learning for weakly supervised semantic segmentation. J. Vis. Commun. Image Represent. 2025, 112, 104576. [Google Scholar] [CrossRef] [Scilit]
- David, L.; Pedrini, H.; Dias, Z. Learning Weakly Supervised Semantic Segmentation Through Cross-Supervision and Contrasting of Pixel-Level Pseudo-Labels. In Proceedings of the 20th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications (VISAPP), Porto, Portugal, 26–28 February 2025; pp. 154–165. [Google Scholar]
- Zhu, L.; Li, Y.; Fang, J.; Liu, Y.; Xin, H.; Liu, W.; Wang, X. WeakTr: Exploring Plain Vision Transformer for Weakly-supervised Semantic Segmentation. IEEE Trans. Image Process. 2026, 35, 4425–4439. [Google Scholar] [CrossRef] [Scilit]
- van den Oord, A.; Li, Y.; Vinyals, O. Representation learning with contrastive predictive coding. arXiv 2018, arXiv:1807.03748. [Google Scholar]









| Type | Environmental Conditions |
|---|---|
| System | Ubuntu 16.04 |
| GPU | NVIDIA RTX 3090 |
| RAM | 64 GB |
| CPU | Intel(R) Xeon(R) 8255C |
| Framework | Pytorch 1.13 |
| Method | Backbone | Pseudo-Label mIoU (%) Training | Pseudo-Label mIoU (%) Validation |
|---|---|---|---|
| SEAM | ResNet38 | 63.6 | 62.1 |
| SLRNet | ResNet38 | 67.1 | 66.2 |
| ViT-PCM | ViT-B | 67.7 | 66.0 |
| AFA | MiT-B1 | 68.7 | 66.5 |
| MCTformer | DeiT-S | 69.1 | 68.2 |
| ToCo | ViT-B | 72.2 | 70.5 |
| WeakTr | ViT-B | 72.0 | 70.2 |
| Ours | ViT-B | 73.4 | 72.3 |
| Method | Backbone | Supervision | Segmentation mIoU (%) Training | Segmentation mIoU (%) Validation |
|---|---|---|---|---|
| Multi-stage WSSS | ||||
| NSROM | ResNet101 | 68.3 | 68.5 | |
| EPS | ResNet101 | 70.9 | 70.8 | |
| L2G | ResNet101 | 72.1 | 71.7 | |
| CLIP-ES | ResNet101 | 71.1 | 71.4 | |
| SEAM | ResNet38 | 65 | 65.7 | |
| SIPE | ResNet101 | 68.8 | 69.7 | |
| W-OoD | ResNet38 | 70.7 | 70.1 | |
| MCTformer | DeiT-S | 71.9 | 71.6 | |
| Single-stage WSSS | ||||
| AFA | MiT-B1 | 66.0 | 66.3 | |
| SLRNet | ResNet38 | 67.2 | 67.6 | |
| ToCo | ViT-B | 69.8 | 70.5 | |
| WeakTr | ViT-B | - | 78.4 | |
| Ours | ViT-B | 73.8 | 78.5 | |
| Method | Backbone | Supervision | Validation Segmentation mIoU (%) |
|---|---|---|---|
| Multi-stage WSSS | |||
| EPS | ResNet101 | 35.7 | |
| L2G | ResNet101 | 42 | |
| SEAM | ResNet38 | 31.9 | |
| SIPE | ResNet101 | 40.6 | |
| MCTformer | DeiT-S | 42.0 | |
| Single-stage WSSS | |||
| SLRNet | WResNet38 | 35.0 | |
| AFA | MiT-B1 | 38.9 | |
| ToCo | ViT-B | 41.3 | |
| WeakTr | ViT-B | 50.3 | |
| Ours | ViT-B | 50.9 | |
| Class | SEAM | AFA | TransCAM | ToCo | Ours |
|---|---|---|---|---|---|
| bkg | 88.8 | 89.9 | 91.3 | 89.9 | 96.0 |
| aero | 68.5 | 79.5 | 81.9 | 81.8 | 87.0 |
| bike | 33.3 | 31.2 | 35.4 | 35.4 | 47.2 |
| bird | 85.7 | 80.7 | 87 | 68.1 | 92.8 |
| boat | 40.4 | 67.2 | 67.6 | 62.0 | 74.0 |
| bottle | 67.3 | 61.9 | 67.9 | 76.6 | 88.8 |
| bus | 78.9 | 81.4 | 87.5 | 83.6 | 91.8 |
| car | 76.3 | 65.4 | 80.5 | 80.4 | 85.5 |
| cat | 81.9 | 82.3 | 86.5 | 87.7 | 95.7 |
| chair | 29.1 | 28.7 | 31.4 | 25 | 36.5 |
| cow | 75.5 | 83.4 | 73.9 | 88.1 | 96.7 |
| table | 48.1 | 41.6 | 52.5 | 59 | 62.7 |
| dog | 79.9 | 82.2 | 80 | 87.0 | 91.6 |
| horse | 73.8 | 75.9 | 79 | 80 | 86.9 |
| motor | 71.4 | 70.2 | 76 | 76.0 | 80.8 |
| person | 75.2 | 69.4 | 79.0 | 68.2 | 84.8 |
| plant | 48.9 | 53.0 | 47 | 65.6 | 69.8 |
| sheep | 79.8 | 85.9 | 81 | 85.8 | 92.5 |
| sofa | 40.9 | 41 | 47.0 | 42.4 | 53.2 |
| train | 58.2 | 62 | 78.4 | 57.7 | 68.2 |
| tv | 53.0 | 50.9 | 46.6 | 65.6 | 66.0 |
| mIoU | 65.0 | 66.0 | 69.3 | 69.8 | 78.5 |
| Method | mIoU (%) | Cosine Similarity | MAE | MSE | RMSE |
|---|---|---|---|---|---|
| AFA | 66.3 | 0.824 | 0.137 | 0.048 | 0.219 |
| ToCo | 70.5 | 0.873 | 0.112 | 0.035 | 0.187 |
| WeakTr | 78.4 | 0.908 | 0.083 | 0.023 | 0.152 |
| Ours | 78.5 | 0.914 | 0.079 | 0.021 | 0.145 |
| Method | Backbone | Params (M) | Inference Time (ms) |
|---|---|---|---|
| SEAM | ResNet38 | 41.2 | 48.6 |
| ToCo | ViT-B | 86.3 | 72.5 |
| MCTformer | DeiT-S | 44.7 | 61.3 |
| WeakTr | ViT-B | 86.3 | 74.1 |
| Ours | ViT-B | 88.1 | 76.8 |
| Method | LTG | CSM | FFCAM | CAM (%) | Seg. (%) |
|---|---|---|---|---|---|
| Baseline | - | - | - | 49.7 | 47.8 |
| Ours | - | - | 63 | 63.6 | |
| - | 69.4 | 69.7 | |||
| 73.3 | 78.5 |
| Size | CAM mIoU (%) | Segmentation mIoU (%) |
|---|---|---|
| 69.6 | 67.1 | |
| 71.2 | 69.6 | |
| 73.3 | 78.5 | |
| 72.4 | 70.1 |
| Size | CAM mIoU (%) | Segmentation mIoU (%) |
|---|---|---|
| 0 | 70.1 | 68.6 |
| 0.1 | 72.7 | 70.4 |
| 0.9 | 73.3 | 78.5 |
| 0.99 | 71.7 | 69.1 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Guo, B.; Yu, L.; Yang, Y.; Wang, C. A Weakly Supervised Segmentation Algorithm Based on Local–Global Class Labelling Comparison. Appl. Sci. 2026, 16, 8496. https://doi.org/10.3390/app16178496
Guo B, Yu L, Yang Y, Wang C. A Weakly Supervised Segmentation Algorithm Based on Local–Global Class Labelling Comparison. Applied Sciences. 2026; 16(17):8496. https://doi.org/10.3390/app16178496
Chicago/Turabian StyleGuo, Binyu, Laibao Yu, Yiming Yang, and Chunzhi Wang. 2026. "A Weakly Supervised Segmentation Algorithm Based on Local–Global Class Labelling Comparison" Applied Sciences 16, no. 17: 8496. https://doi.org/10.3390/app16178496
APA StyleGuo, B., Yu, L., Yang, Y., & Wang, C. (2026). A Weakly Supervised Segmentation Algorithm Based on Local–Global Class Labelling Comparison. Applied Sciences, 16(17), 8496. https://doi.org/10.3390/app16178496

