Discrepancy-Guided Semantic Segmentation with Boundary Detail Enhancement for Traffic Scenes
Abstract
1. Introduction
- A Gated Collaborative Context Module (GCCM) is proposed to adaptively regulate feature information flow by integrating a gating mechanism, multi-scale feature collaboration, and context completion strategies, thereby effectively alleviating the problems of incomplete semantic representation and insufficient boundary expression in fine-grained target regions.
- A Frequency–Edge Guided Enhancement Module (FEGE) is proposed to combine frequency-domain decomposition with an edge-aware mechanism. Through explicit high-frequency guidance, the module achieves collaborative enhancement of structural preservation and boundary details, thereby improving the model’s responsiveness to high-frequency structural information and the accuracy of boundary representation.
- A Discrepancy-aware Pixel-Adaptive Gating Fusion module (D-PagFM) is designed to adaptively regulate pixel-wise fusion regions by jointly modeling feature similarity and local discrepancy, thereby enhancing the robustness and prediction consistency of feature fusion in boundary regions, small-object areas, and occluded regions.
2. Related Work
2.1. Research on Improving Segmentation Accuracy in Traffic Scenes
2.2. Research on Real-Time Semantic Segmentation for Traffic Scenes
2.3. Research on Cross-Scene Generalization for Traffic Scene Segmentation
3. The Proposed Algorithm
3.1. Network Architecture Overview
3.1.1. Encoder Pipeline
3.1.2. Decoder with Edge Enhancement and Discrepancy-Aware Fusion
3.1.3. Construction of the Semantic Segmentation Head
3.2. Design of the Gated Collaborative Context Module
3.3. Design of the Frequency–Edge Guided Enhancement Module
3.4. Design of the Discrepancy-Aware Pixel-Adaptive Gating Fusion Module
4. Experimental Results and Analysis
4.1. Datasets and Evaluation Metrics
- (1)
- Pixel Accuracy (PA) and Mean Pixel Accuracy (MPA)
- (2)
- Mean Intersection over Union (mIoU)
4.2. Experimental Environment and Parameter Settings
4.3. Ablation Experiments
4.3.1. Inter-Module Ablation Study
4.3.2. Intra-Module Ablation Study of D-PagFM
4.4. Quantitative and Qualitative Analysis of Algorithm Performance
4.4.1. Experimental Results on the Cityscapes Dataset
4.4.2. Experimental Results on the CamVid Dataset
5. Discussion and Conclusions
5.1. Discussion and Limitations
5.2. Limitations and Future Work
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Wang, H.; Chen, Y.; Cai, Y.; Chen, L.; Li, Y.; Sotelo, M.A.; Li, Z. SFNet-N: An improved SFNet algorithm for semantic segmentation of low-light autonomous driving road scenes. IEEE Trans. Intell. Transp. Syst. 2022, 23, 21405–21417. [Google Scholar] [CrossRef]
- Tao, J.; Chen, Z.; Sun, Z.; Guo, H.; Leng, B.; Yu, Z.; Wang, Y.; He, Z.; Lei, X.; Yang, J. Seg-road: A segmentation network for road extraction based on transformer and CNN with connectivity structures. Remote Sens. 2023, 15, 1602. [Google Scholar] [CrossRef]
- Latsaheb, B.; Sharma, S.; Hasija, S. Semantic road segmentation using encoder–decoder architectures. Multimed. Tools Appl. 2025, 84, 5961–5983. [Google Scholar] [CrossRef]
- Wang, H.; Chen, G.; Li, Z.; Wu, Y. Context-aware method for small object segmentation in road scenes. In Proceedings of the International Conference on Advanced Robotics and Mechatronics (ICARM); IEEE: New York, NY, USA, 2022; pp. 238–243. [Google Scholar]
- Xu, T.; Zheng, G. Semantic segmentation of street scene based on multi-scale and attention mechanism. World Sci. Res. J. 2021, 7, 292–302. [Google Scholar]
- Li, K.; Geng, Q.; Zhou, Z. Exploring scale-aware features for real-time semantic segmentation of street scenes. IEEE Trans. Intell. Transp. Syst. 2023, 25, 3575–3587. [Google Scholar] [CrossRef]
- Bai, C.; Zhang, L.; Gao, L.; Peng, L.; Li, P.; Yang, L. Real-time segmentation algorithm of unstructured road scenes based on improved BiSeNet. J. Real-Time Image Process. 2024, 21, 91. [Google Scholar] [CrossRef]
- Xie, G.; Wang, Q.Y.; Xie, X.L.; Wang, J.A. Lightweight transformer-based semantic segmentation algorithm for traffic scenes with multi-scale deep convolution fusion. J. Commun. 2023, 44, 213–225. [Google Scholar]
- Guo, D.; Zhu, L.; Lu, Y.; Yu, H.; Wang, S. Small object sensitive segmentation of urban street scene with spatial adjacency between object classes. IEEE Trans. Image Process. 2018, 28, 2643–2653. [Google Scholar] [CrossRef] [PubMed]
- Gao, C.; Zhao, F.; Zhang, Y.; Wan, M. Research on multitask model of object detection and road segmentation in unstructured road scenes. Meas. Sci. Technol. 2024, 35, 065113. [Google Scholar] [CrossRef]
- Hong, Y.; Pan, H.; Sun, W.; Jia, Y. Deep dual-resolution networks for real-time and accurate semantic segmentation of road scenes. arXiv 2021, arXiv:2101.06085. [Google Scholar]
- Fan, M.; Lai, S.; Huang, J.; Wei, X.; Chai, Z.; Luo, J.; Wei, X. Rethinking BiSeNet for real-time semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2021; pp. 9716–9725. [Google Scholar]
- Xu, J.; Xiong, Z.; Bhattacharyya, S.P. PIDNet: A real-time semantic segmentation network inspired by PID controllers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2023; pp. 19529–19539. [Google Scholar]
- Xu, Z.; Wu, D.; Yu, C.; Chu, X.; Sang, N.; Gao, C. SCTNet: Single-branch CNN with transformer semantic information for real-time segmentation. Proc. AAAI Conf. Artif. Intell. 2024, 38, 6378–6386. [Google Scholar] [CrossRef]
- Ding, J.; Xue, N.; Xia, G.S.; Schiele, B.; Dai, D. HGFormer: Hierarchical grouping transformer for domain generalized semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2023; pp. 15413–15423. [Google Scholar]
- Chen, K.; Chen, H.; Chen, L.; Jin, X.; Jin, Y.; Wei, Z.; Zheng, M. Deliberated domain bridging for domain adaptive semantic segmentation. Adv. Neural Inf. Process. Syst. 2022, 35, 15105–15118. [Google Scholar]
- Gong, Z.; Li, F.; Deng, Y.; Bhattacharjee, D.; Ma, X.; Zhu, X.; Ji, Z. CODA: Instructive chain-of-domain adaptation with severity-aware visual prompt tuning. In Proceedings of the European Conference on Computer Vision (ECCV); Springer: Cham, Switzerland, 2024; pp. 130–148. [Google Scholar]
- Fang, X.; Song, X.; Meng, X.; Fang, X.; Jin, S. MCFNet: Multi-scale covariance feature fusion network for real-time semantic segmentation. arXiv 2023, arXiv:2312.07207. [Google Scholar]
- Mi, B.; Liang, X. Context and boundary guided multi-scale feature fusion network for semantic segmentation. In Proceedings of the IEEE International Conference on Multimedia and Expo (ICME); IEEE: New York, NY, USA, 2022; pp. 1–6. [Google Scholar]
- Li, S.; Wan, L.; Tang, L.; Zhang, Z. MFEAFN: Multi-scale feature enhanced adaptive fusion network for image semantic segmentation. PLoS ONE 2022, 17, e0274249. [Google Scholar] [CrossRef] [PubMed]
- Xiu, C.; Su, H.; Su, X. Semantic segmentation method based on residual and multi-scale feature fusion. In Proceedings of the Chinese Control and Decision Conference (CCDC); IEEE: New York, NY, USA, 2020; pp. 2078–2083. [Google Scholar]
- Yu, C.; Wang, J.; Peng, C.; Gao, C.; Yu, G.; Sang, N. BiSeNet: Bilateral segmentation network for real-time semantic segmentation. In Proceedings of the European Conference on Computer Vision (ECCV); Springer: Cham, Switzerland, 2018; pp. 325–341. [Google Scholar]
- Cordts, M.; Omran, M.; Ramos, S.; Rehfeld, T.; Enzweiler, M.; Benenson, R.; Franke, U.; Roth, S.; Schiele, B. The Cityscapes dataset for semantic urban scene understanding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2016; pp. 3213–3223. [Google Scholar]
- Zhang, Y.; Yang, R.; Wang, J.; Chen, N.; Dai, Q. The impact of parameters on semantic segmentation: A case study on the CamVid dataset. In Proceedings of the IEEE International Conference on High Performance Computing & Communications; IEEE: New York, NY, USA, 2021; pp. 1932–1938. [Google Scholar]
- Deng, J.; Dong, W.; Socher, R.; Li, L.J.; Li, K.; Li, F.-F. ImageNet: A large-scale hierarchical image database. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2009; pp. 248–255. [Google Scholar]
- Chen, L.C.; Zhu, Y.; Papandreou, G.; Schroff, F.; Adam, H. Encoder–decoder with atrous separable convolution for semantic image segmentation. In Proceedings of the European Conference on Computer Vision (ECCV); Springer: Cham, Switzerland, 2018; pp. 801–818. [Google Scholar]
- Xiao, X.; Zhao, Y.; Zhang, F.; Luo, B.; Yu, L.; Chen, B.; Yang, C. BASeg: Boundary-aware semantic segmentation for autonomous driving. Neural Netw. 2023, 157, 460–470. [Google Scholar] [CrossRef] [PubMed]
- Cheng, B.; Misra, I.; Schwing, A.G.; Kirillov, A.; Girdhar, R. Masked-attention mask transformer for universal image segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2022; pp. 1290–1299. [Google Scholar]
- Cheng, M.-M.; Guo, M.-H.; Hou, Q.; Hu, S.-M.; Liu, Z.; Lu, C.-Z. SegNeXt: Rethinking convolutional attention design for semantic segmentation. Adv. Neural Inf. Process. Syst. 2022, 35, 1140–1156. [Google Scholar]
- Xie, E.; Wang, W.; Yu, Z.; Anandkumar, A.; Alvarez, J.M.; Luo, P. SegFormer: Simple and efficient design for semantic segmentation with transformers. Adv. Neural Inf. Process. Syst. 2021, 34, 12077–12090. [Google Scholar]
- Nirkin, Y.; Wolf, L.; Hassner, T. HyperSeg: Patch-wise hypernetwork for real-time semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2021; pp. 4061–4070. [Google Scholar]







| Baseline | GCCM | FEGE | D-PagFM | Params (M) | PA (%) | mPA (%) | mIoU (%) |
|---|---|---|---|---|---|---|---|
| √ | − | − | − | 27.36 | 96.06 | 86.15 | 77.91 |
| √ | − | √ | − | 35.21 | 96.22 | 86.43 | 78.96 |
| √ | √ | − | − | 31.12 | 96.23 | 87.15 | 79.0 |
| √ | − | − | √ | 29.14 | 96.2 | 87.18 | 79.5 |
| √ | √ | √ | − | 38.97 | 96.3 | 87.01 | 79.42 |
| √ | − | √ | √ | 37.00 | 96.33 | 87.5 | 79.68 |
| √ | √ | − | √ | 32.91 | 96.33 | 86.79 | 79.35 |
| √ | √ | √ | √ | 40.76 | 96.37 | 87.95 | 80.08 |
| Baseline | PagFM | D-PagFM | Params (M) | PA (%) | mPA (%) | mIoU (%) |
|---|---|---|---|---|---|---|
| √ | − | − | 27.36 | 96.06 | 86.15 | 77.91 |
| √ | √ | − | 28.70 | 96.24 | 87.05 | 79.35 |
| √ | − | √ | 29.14 | 96.2 | 87.18 | 79.5 |
| Model | Params (M) | PA (%) | mPA (%) | mIoU (%) |
|---|---|---|---|---|
| Sim-PagFM | 29.14 | 96.28 | 86.86 | 79.27 |
| Conf-PagFM | 29.14 | 96.26 | 86.52 | 79.13 |
| Joint-PagFM | 29.14 | 96.26 | 87.04 | 79.41 |
| D-PagFM | 29.14 | 96.2 | 87.18 | 79.50 |
| Model | Params (M) | PA (%) | mPA (%) | mIoU (%) |
|---|---|---|---|---|
| DeepLab-V3+ [26] | 58.75 | 95.67 | 83.49 | 75.04 |
| BASeg [27] | 63.99 | 95.76 | 86.64 | 80.39 |
| Mask2Former-R50 [28] | 62.1 | - | - | 79.4 |
| Mask2Former-Swin-B [28] | 66.1 | - | - | 83.3 |
| SegFormer-B2 [30] | 27.36 | 96.06 | 86.15 | 77.91 |
| SegFormer-B3 [30] | 47.3 | 96.22 | 87.51 | 81.7 |
| STDC2-Seg100 [12] | 12.95 | 95.75 | 84.53 | 76.93 |
| SegNeXt-T [29] | 4.3 | - | - | 79.8 |
| SegNeXt-S [29] | 13.9 | - | - | 81.3 |
| PIDNet-Small [15] | 7.6 | 96.11 | 86.11 | 78.60 |
| Ours | 40.76 | 96.37 | 87.95 | 80.08 |
| Algorithm/Category | SegFormer-B2 | Ours |
|---|---|---|
| road | 98.23 | 98.3 |
| sidewalk | 85.38 | 85.59 |
| building | 92.47 | 93.14 |
| wall | 62.32 | 63.21 |
| fence | 59.94 | 61.91 |
| pole | 63.9 | 69.02 |
| traffic light | 70.53 | 73.79 |
| traffic sign | 78.48 | 81.74 |
| vegetation | 92.21 | 92.75 |
| terrain | 63.56 | 61.6 |
| sky | 94.94 | 95.2 |
| person | 82.29 | 83.68 |
| rider | 61.54 | 64.57 |
| car | 94.95 | 95.36 |
| truck | 82.32 | 84.44 |
| bus | 82.63 | 89.1 |
| train | 71.17 | 80.51 |
| motorcycle | 67.6 | 69.26 |
| bicycle | 76.22 | 78.32 |
| mIoU | 77.93 | 80.08 |
| Model | Params (M) | PA (%) | mPA (%) | mIoU (%) |
|---|---|---|---|---|
| DeepLab-V3+ [26] | 58.75 | 94.22 | 85.16 | 78.03 |
| HyperSeg-L [31] | 67.95 | 94.36 | 85.72 | 77.95 |
| PIDNet-Small [15] | 7.6 | 95.26 | 86.64 | 80.01 |
| Segformer-B2 [30] | 27.36 | 94.22 | 90.85 | 80.65 |
| Ours | 40.76 | 95.06 | 91.33 | 82.97 |
| Algorithm/Category | Segformer-B2 | Ours |
|---|---|---|
| sky | 85.65 | 88.19 |
| building | 89.75 | 91.17 |
| pole | 53.35 | 58.37 |
| road | 97.6 | 97.69 |
| pavement | 89.15 | 89.26 |
| tree | 82.5 | 83.94 |
| signsymbol | 72.42 | 76.93 |
| fence | 76.62 | 79.79 |
| car | 90.0 | 91.82 |
| pedestrain | 69.65 | 73.01 |
| bicyclist | 80.44 | 82.54 |
| mIoU | 80.65 | 82.97 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Yu, C.; Yang, X.; Shen, S. Discrepancy-Guided Semantic Segmentation with Boundary Detail Enhancement for Traffic Scenes. Sensors 2026, 26, 2738. https://doi.org/10.3390/s26092738
Yu C, Yang X, Shen S. Discrepancy-Guided Semantic Segmentation with Boundary Detail Enhancement for Traffic Scenes. Sensors. 2026; 26(9):2738. https://doi.org/10.3390/s26092738
Chicago/Turabian StyleYu, Changshun, Xiujian Yang, and Shiquan Shen. 2026. "Discrepancy-Guided Semantic Segmentation with Boundary Detail Enhancement for Traffic Scenes" Sensors 26, no. 9: 2738. https://doi.org/10.3390/s26092738
APA StyleYu, C., Yang, X., & Shen, S. (2026). Discrepancy-Guided Semantic Segmentation with Boundary Detail Enhancement for Traffic Scenes. Sensors, 26(9), 2738. https://doi.org/10.3390/s26092738

