Exploring the Synergy Between Language Semantic Guidance and Visual Attention in Knowledge Distillation for Semantic Segmentation Under a Limited Field of View
Abstract
1. Introduction
- We devise a SAFD framework that incorporates LLM-derived semantic knowledge into feature distillation without increasing inference complexity.
- We systematically investigate multiple attention-based visual feature distillation strategies and textual representations to analyze how semantic guidance interacts with different forms of visual supervision.
- We show that semantic guidance is especially effective when visual distillation provides limited semantic supervision, while its effectiveness depends on the visual representations transferred from the teacher, the adopted attention strategy, and the dataset characteristics.
- We further demonstrate that semantic guidance produces distinct class-wise confusion patterns across attention-based distillation strategies and consistent performance gains across the evaluated text encoders, while appropriate combinations with each attention strategy can improve generalizability and prediction reliability.
2. Related Work
2.1. Semantic Segmentation in Limited FoV
2.2. Knowledge Distillation
2.3. Language-Guided Semantic Segmentation
3. Visual Attention and Semantic Guidance Strategies
3.1. Visual Attention for Knowledge Transfer
- Channel Attention Module (CAM): Woo et al. [39] proposed two attention modules, the first of which is the Channel Attention Module (CAM) utilizing the inter-channel relationship of features by aggregating spatial information. Specifically, given an input feature map (C, H, and W denote the channel, height, and width dimensions, respectively), the max-pooling and average-pooling operations are applied along the spatial dimensions to generate a channel attention map , highlighting ‘what’ is meaningful in a given image. From a distillation perspective, this approach facilitates the transfer of discriminative channel-wise knowledge to the student model. Following the CAM, the channel-refined feature map is formulated as:where and represent the feature maps resulting from average-pooling and max-pooling operations on the feature map F, respectively. and are the weights of the multi-layer perceptron (MLP) shared for the two pooled feature maps, and is followed by a ReLU activation function. denotes the sigmoid function and ⊗ denotes element-wise multiplication.
- Spatial Attention Module (SAM): The second module is the Spatial Attention Module (SAM), which exploits the spatial relationship within the feature maps. Similar to CAM, two feature maps and are generated by applying average-pooling and max-pooling operations to the feature map F along the channel dimension. These maps are then concatenated and passed through a convolutional layer to generate a spatial attention map highlighting ‘where’ the informative regions are. Following the same principle, the spatially refined feature map of the SAM is computed as:
- Convolutional Block Attention Module (CBAM): CBAM uses CAM and SAM sequentially to complement the features. Specifically, the channel-refined feature map is generated using the channel attention map and is then further refined by the spatial attention map to produce the channel- and spatially refined feature map . The dual attention process is formulated as follows:
3.2. Semantic Knowledge Extraction
All classes in an urban driving scene are "{cls 1}", "{cls 2}", … "{cls N}". Tell me about the {k} general words that describe each class.
- where k is a predefined constant set to 5 during our experiments. Similarly, the prompt template for Sentence format is structured as follows:
All classes in an urban driving scene are "{cls 1}", "{cls 2}", … "{cls N}". Tell me about the general sentences that describe each class.
3.3. Semantic-Guided Attentive Knowledge Distillation
- Visual Feature Attention Module. In the feature transfer process, the student model leverages attentive visual representations from the teacher, which is pre-trained solely on image data and kept frozen. Specifically, intermediate feature maps before the segmentation head are refined using one of the three aforementioned attention strategies. This module enables the student to acquire refined visual knowledge distilled from the teacher.
- Semantic Knowledge Fusion Module. Beyond visual supervision, we incorporate an auxiliary module to compensate for the lack of visual context during training. To improve semantic reasoning based on visual context, we adopt a fusion approach at the logit level of the student model to inject high-level semantic priors. Specifically, the generated linguistic descriptions are first transformed into text embeddings using a pre-trained vision-language text encoder, which aligns textual representations with the visual feature space [26]. The resulting text embeddings are represented as , where c denotes the number of semantic classes and d the embedding dimension. We then employ a learnable projection layer to project the embedding dimension d to the class dimension c. The transformation is defined as:where denotes the learnable projection layer, and represents the projected text embeddings. This projection makes text representations compatible with the student model’s visual logits . Then, to integrate these modalities, we employ a cross-attention mechanism with as the query and as the key and value [45]. The semantic attention feature is computed as:where q, k, and v denote the learnable linear projections for the query, key, and value, respectively. captures the semantic inter-class relationships that the student needs to consider. Finally, is projected back to the prediction space, producing the fused logits . The resulting logits are combined with through a residual connection [46]:where denotes the output projection layer, and represents the final logits used for training. The final logits z are then used to compute the standard cross-entropy loss:where denotes the ground-truth segmentation map.
- Inference Phase. While the student model is trained with both visual and linguistic supervision, the teacher model and the semantic knowledge fusion module are discarded during the inference phase. Therefore, the original student model’s logits serve as the final output for inference. This design maintains the model’s computational efficiency and lightweight advantage while retaining the enhanced contextual knowledge acquired during training.
4. Experiments
4.1. Datasets
4.2. Implementation Details
4.3. Joint Supervision from Attention-Based KD and Semantic Guidance
4.3.1. Comparison with Representative KD Methods
4.3.2. Semantic Guidance with Attention-Based KD
4.4. Ablation and Additional Analysis
4.4.1. Choice of Text Encoder
4.4.2. Hyperparameter Sensitivity
4.4.3. Visualization
- Segmentation Result. Figure 5 presents qualitative semantic segmentation results on the CamVid and KITTI datasets. SAFD-CS is adopted as the representative attention-based distillation strategy, while the Label representation is used to evaluate the effect of semantic guidance.
- Confusion Matrix Difference. While the segmentation results demonstrate spatial improvements, we further analyze the class-wise prediction behavior and inter-class confusions through confusion matrix difference maps shown in Figure 6 and Figure 7. Pixel-level confusion matrices are computed from the predicted and ground-truth labels over all pixels in the test set. Each confusion matrix is row-wise normalized to account for class imbalance, and the difference map is obtained by subtracting the normalized confusion matrix of the baseline method from that of the compared method.
4.4.4. Generalizability Analysis of SAFD
- Generalization under Limited Training Data. To examine the generalization behavior of the framework under limited training data, we vary the proportion of training samples used to train the student to 100%, 50%, and 30% while leaving the validation and test sets unchanged. At each proportion, we compare SAFD-CS (Label) with a standalone student trained using the same subset. The reduced training subsets are nested within the full training set. All other training settings—including data augmentation, the optimizer, batch size, and total training iterations—are kept identical to those described in Section 4.2. Since the number of training iterations is fixed, samples in the reduced subsets are revisited more frequently, providing increasingly challenging conditions with a higher risk of overfitting.
- Reliability and Generalizability. Beyond evaluating performance under limited training data, we further examine whether the predictive probabilities produced by SAFD reliably reflect prediction correctness on unseen test samples. We first visualize the calibration behavior of representative semantic guidance configurations and then quantitatively evaluate all textual representations on the KITTI dataset.
5. Discussion
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Feng, D.; Haase-Schütz, C.; Rosenbaum, L.; Hertlein, H.; Glaeser, C.; Timm, F.; Wiesbeck, W.; Dietmayer, K. Deep multi-modal object detection and semantic segmentation for autonomous driving: Datasets, methods, and challenges. IEEE Trans. Intell. Transp. Syst. 2020, 22, 1341–1360. [Google Scholar] [CrossRef] [Scilit]
- Chen, B.; Gong, C.; Yang, J. Importance-aware semantic segmentation for autonomous vehicles. IEEE Trans. Intell. Transp. Syst. 2018, 20, 137–148. [Google Scholar] [CrossRef] [Scilit]
- Ranftl, R.; Bochkovskiy, A.; Koltun, V. Vision transformers for dense prediction. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2021; pp. 12179–12188. [Google Scholar]
- Vo, X.T.; Nguyen, D.L.; Priadana, A.; Cao, G.; Choi, J.; Jo, K.H. Local self-attention with mixing abstract tokens for urban autonomous driving. IEEE Trans. Ind. Inform. 2025, 21, 5420–5430. [Google Scholar] [CrossRef] [Scilit]
- Deng, L.; Yang, M.; Li, H.; Li, T.; Hu, B.; Wang, C. Restricted deformable convolution-based road scene semantic segmentation using surround view cameras. IEEE Trans. Intell. Transp. Syst. 2019, 21, 4350–4362. [Google Scholar] [CrossRef] [Scilit]
- Yogamani, S.; Unger, D.; Narayanan, V.; Kumar, V.R. DaF-BEVSeg: Distortion-aware Fisheye Camera based Bird’s Eye View Segmentation with Occlusion Reasoning. arXiv 2024, arXiv:2404.06352. [Google Scholar]
- Shi, H.; Li, Y.; Yang, K.; Zhang, J.; Peng, K.; Roitberg, A.; Ye, Y.; Ni, H.; Wang, K.; Stiefelhagen, R. FishDreamer: Towards Fisheye Semantic Completion via Unified Image Outpainting and Segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops; IEEE: New York, NY, USA, 2023; pp. 6434–6444. [Google Scholar]
- Pro, F.; Dionelis, N.; Maiano, L.; Le Saux, B.; Amerini, I. A semantic segmentation-guided approach for ground-to-aerial image matching. In Proceedings of the IGARSS 2024-2024 IEEE International Geoscience and Remote Sensing Symposium; IEEE: New York, NY, USA, 2024; pp. 2630–2635. [Google Scholar]
- Hinton, G.; Vinyals, O.; Dean, J. Distilling the knowledge in a neural network. arXiv 2015, arXiv:1503.02531. [Google Scholar]
- Gou, J.; Yu, B.; Maybank, S.J.; Tao, D. Knowledge distillation: A survey. Int. J. Comput. Vis. 2021, 129, 1789–1819. [Google Scholar] [CrossRef] [Scilit]
- Zagoruyko, S.; Komodakis, N. Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer. In Proceedings of the International Conference on Learning Representations; OpenReview: Alameda, CA, USA, 2017; pp. 1–13. [Google Scholar]
- Ji, M.; Heo, B.; Park, S. Show, Attend and Distill: Knowledge Distillation via Attention-based Feature Matching. Proc. AAAI Conf. Artif. Intell. 2021, 35, 7945–7952. [Google Scholar] [CrossRef] [Scilit]
- Shin, S.; Lee, J.; Lee, J.; Yu, Y.; Lee, K. Teaching where to look: Attention similarity knowledge distillation for low resolution face recognition. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2022; pp. 631–647. [Google Scholar]
- Shu, C.; Liu, Y.; Gao, J.; Yan, Z.; Shen, C. Channel-wise knowledge distillation for dense prediction. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2021; pp. 5311–5320. [Google Scholar]
- Mansourian, A.; Jalali, A.; Ahmadi, R.; Kasaei, S. Attention as Geometric Transformation: Revisiting Feature Distillation for Semantic Segmentation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision; IEEE: New York, NY, USA, 2026; pp. 1287–1297. [Google Scholar]
- Palmer, S.E. The effects of contextual scenes on the identification of objects. Mem. Cogn. 1975, 3, 519–526. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Oliva, A.; Torralba, A. The role of context in object recognition. Trends Cogn. Sci. 2007, 11, 520–527. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Biederman, I. On the semantics of a glance at a scene. In Perceptual Organization; Routledge: London, UK, 2017; pp. 213–253. [Google Scholar]
- Du, T.; Wang, Z.; Wang, Y.; Ma, M.; Li, W. Biologically Inspired Medical Multi-Modal Dataset Distillation via Contrast-Aware Alignment and Memory Compression. Biomimetics 2026, 11, 314. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wu, Y.; Mi, Q.; Gao, T. A comprehensive review of multimodal emotion recognition: Techniques, challenges, and future directions. Biomimetics 2025, 10, 418. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Li, L.H.; Zhang, P.; Zhang, H.; Yang, J.; Li, C.; Zhong, Y.; Wang, L.; Yuan, L.; Zhang, L.; Hwang, J.N.; et al. Grounded language-image pre-training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2022; pp. 10965–10975. [Google Scholar]
- Xu, J.; De Mello, S.; Liu, S.; Byeon, W.; Breuel, T.; Kautz, J.; Wang, X. GroupVit: Semantic segmentation emerges from text supervision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2022; pp. 18134–18144. [Google Scholar]
- Li, B.; Weinberger, K.Q.; Belongie, S.; Koltun, V.; Ranftl, R. Language-driven Semantic Segmentation. In Proceedings of the International Conference on Learning Representations; OpenReview: Alameda, CA, USA, 2022. [Google Scholar]
- Menon, S.; Vondrick, C. Visual Classification via Description from Large Language Models. In Proceedings of the International Conference on Learning Representations; OpenReview: Alameda, CA, USA, 2023. [Google Scholar]
- Pratt, S.; Covert, I.; Liu, R.; Farhadi, A. What does a platypus look like? generating customized prompts for zero-shot image classification. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2023; pp. 15691–15701. [Google Scholar]
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2021; pp. 8748–8763. [Google Scholar]
- Rao, Y.; Zhao, W.; Chen, G.; Tang, Y.; Zhu, Z.; Huang, G.; Zhou, J.; Lu, J. DenseCLIP: Language-guided dense prediction with context-aware prompting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2022; pp. 18082–18091. [Google Scholar]
- Hoyer, L.; Tan, D.J.; Naeem, M.F.; Van Gool, L.; Tombari, F. SemiVL: Semi-supervised semantic segmentation with vision-language guidance. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2024; pp. 257–275. [Google Scholar]
- Wang, Y.; Wang, Y.; Dai, R.; Wang, Y.; Liu, K.; Chu, X.; Li, Y. Urban Socio-Semantic Segmentation with Vision-Language Reasoning. In Proceedings of the International Conference on Learning Representations; OpenReview: Alameda, CA, USA, 2026. [Google Scholar]
- Ma, Y.; Ma, J.; Zhou, M.; Chen, Q.; Ge, T.; Jiang, Y.; Lin, T. Boosting image outpainting with semantic layout prediction. arXiv 2021, arXiv:2110.09267. [Google Scholar]
- Seifi, S.; Tuytelaars, T. Attend and segment: Attention guided active semantic segmentation. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2020; pp. 305–321. [Google Scholar]
- Ba, L.J.; Caruana, R. Do deep nets really need to be deep? Adv. Neural Inf. Process. Syst. 2014, 27. [Google Scholar]
- Heo, B.; Kim, J.; Yun, S.; Park, H.; Kwak, N.; Choi, J.Y. A comprehensive overhaul of feature distillation. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2019; pp. 1921–1930. [Google Scholar]
- Chen, P.; Liu, S.; Zhao, H.; Jia, J. Distilling knowledge via knowledge review. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2021; pp. 5008–5017. [Google Scholar]
- Tung, F.; Mori, G. Similarity-preserving knowledge distillation. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2019; pp. 1365–1374. [Google Scholar]
- Park, W.; Kim, D.; Lu, Y.; Cho, M. Relational knowledge distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2019; pp. 3967–3976. [Google Scholar]
- Chen, L.; Wang, D.; Gan, Z.; Liu, J.; Henao, R.; Carin, L. Wasserstein contrastive representation distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2021; pp. 16296–16305. [Google Scholar]
- Zhou, Z.; Zhuge, C.; Guan, X.; Liu, W. Channel distillation: Channel-wise attention for knowledge distillation. arXiv 2020, arXiv:2006.01683. [Google Scholar]
- Woo, S.; Park, J.; Lee, J.Y.; Kweon, I.S. CBAM: Convolutional Block Attention Module. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2018; pp. 3–19. [Google Scholar]
- Anthropic. Claude 4 Model Family Updates. 2026. Available online: https://www.anthropic.com/claude (accessed on 19 July 2026).
- Google DeepMind. Gemini 3.5: Frontier Intelligence with Action. 2026. Available online: https://deepmind.google/models/gemini/ (accessed on 19 July 2026).
- OpenAI. GPT-5 System Card Large Language Model. 2025. Available online: https://openai.com/index/gpt-5-system-card/ (accessed on 7 August 2025).
- Roth, K.; Vinyals, O.; Akata, Z. Integrating language guidance into vision-based deep metric learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2022; pp. 16177–16189. [Google Scholar]
- El Banani, M.; Desai, K.; Johnson, J. Learning visual representations via language-guided sampling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2023; pp. 19208–19220. [Google Scholar]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, L.; Polosukhin, I. Attention is all you need. Adv. Neural Inf. Process. Syst. 2017, 30. [Google Scholar]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2016; pp. 770–778. [Google Scholar]
- Brostow, G.J.; Shotton, J.; Fauqueur, J.; Cipolla, R. Segmentation and recognition using structure from motion point clouds. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2008; pp. 44–57. [Google Scholar]
- Geiger, A.; Lenz, P.; Stiller, C.; Urtasun, R. Vision meets robotics: The KITTI dataset. Int. J. Robot. Res. 2013, 32, 1231–1237. [Google Scholar] [CrossRef] [Scilit]
- Kohavi, R. A study of cross-validation and bootstrap for accuracy estimation and model selection. In Proceedings of the IJCAI, Montreal, QC, Canada, 20–25 August 1995; Volume 2, pp. 1137–1143. [Google Scholar]
- Xie, E.; Wang, W.; Yu, Z.; Anandkumar, A.; Alvarez, J.M.; Luo, P. SegFormer: Simple and efficient design for semantic segmentation with transformers. Adv. Neural Inf. Process. Syst. 2021, 34, 12077–12090. [Google Scholar]
- Cherti, M.; Beaumont, R.; Wightman, R.; Wortsman, M.; Ilharco, G.; Gordon, C.; Schuhmann, C.; Schmidt, L.; Jitsev, J. Reproducible scaling laws for contrastive language-image learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2023; pp. 2818–2829. [Google Scholar]
- Zheng, B.; Cheng, R. Rethinking Decoupled Knowledge Distillation: A Predictive Distribution Perspective. In Proceedings of the IEEE Transactions on Neural Networks and Learning Systems; IEEE: New York, NY, USA, 2025. [Google Scholar]
- Zhai, X.; Mustafa, B.; Kolesnikov, A.; Beyer, L. Sigmoid loss for language image pre-training. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2023; pp. 11941–11952. [Google Scholar]
- Devlin, J.; Chang, M.W.; Lee, K.; Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers); Association for Computational Linguistics: Stroudsburg, PA, USA, 2019; pp. 4171–4186. [Google Scholar]
- Guo, C.; Pleiss, G.; Sun, Y.; Weinberger, K.Q. On calibration of modern neural networks. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2017; pp. 1321–1330. [Google Scholar]
- Naeini, M.P.; Cooper, G.; Hauskrecht, M. Obtaining well calibrated probabilities using Bayesian binning. Proc. AAAI Conf. Artif. Intell. 2015, 29, 2901–2907. [Google Scholar] [CrossRef] [Scilit]









| Class | Keywords () |
|---|---|
| Sky | Sky, Atmosphere, Horizon, Clouds, Airspace |
| Building | Building, Architecture, Facade, Structure, Windows |
| Road | Road, Asphalt, Lane, Driveway, Trafficway |
| Sidewalk | Sidewalk, Curb, Walkway, Pedway, Paving |
| Fence | Fence, Boundary, Pickets, Barrier, Perimeter |
| Tree | Tree, Foliage, Trunk, Branches, Canopy |
| Pole | Pole, Slender, Vertical, Cable, Metallic |
| Car | Car, Vehicle, Windshield, Engine, Parking |
| SignSymbol | SignSymbol, Traffic sign, Pictogram, Instruction, Regulation |
| Pedestrian | Pedestrian, Crosswalk, Footsteps, Walker, Striding |
| Bicyclist | Bicyclist, Helmet, Pedals, Wheels, Cycling |
| Class | Sentence |
| Sky | The area above the earth, usually blue during the day with clouds. |
| Building | A tall structure made of concrete and glass that houses offices and apartments. |
| Road | A paved surface for vehicles to drive on, often marked with lanes. |
| Sidewalk | A paved path beside the road for people to walk safely. |
| Fence | A wooden or metal barrier that encloses a yard or property. |
| Tree | A tall plant with a trunk, branches, and leaves providing shade. |
| Pole | A vertical post used to support electrical wires or street lamps. |
| Car | A motor vehicle with four wheels used for transportation. |
| SignSymbol | A traffic sign or symbol used to guide or warn drivers and pedestrians. |
| Pedestrian | A person walking on the sidewalk or crossing the street. |
| Bicyclist | A person wearing a helmet riding a bicycle on the street. |
| T: SegFormer-B2 | T: SegFormer-B4 | ||||
|---|---|---|---|---|---|
| DB | Method | mPA | mIoU | mPA | mIoU |
| CamVid | Teacher | 72.85 | 65.47 | 73.62 | 66.52 |
| Student | 67.79 | 60.17 | 67.79 | 60.17 | |
| KD [9] | 66.52 ± 0.11 | 58.96 ± 0.17 | 67.95 ± 0.20 | 60.28 ± 0.17 | |
| AT [11] | 68.05 ± 0.09 | 60.42 ± 0.14 | 67.97 ± 0.16 | 60.49 ± 0.12 | |
| SP [35] | 67.97 ± 0.09 | 60.28 ± 0.17 | 67.84 ± 0.09 | 60.30 ± 0.09 | |
| RKD [36] | 69.19 ± 0.13 | 61.66 ± 0.12 | 68.93 ± 0.13 | 61.40 ± 0.17 | |
| ReviewKD [34] | 68.41 ± 0.26 | 60.91 ± 0.24 | 68.47 ± 0.08 | 60.93 ± 0.13 | |
| GDKD [52] | 68.05 ± 0.05 | 60.28 ± 0.16 | 67.84 ± 0.16 | 60.21 ± 0.17 | |
| AttnFD [15] | 69.72 ± 0.09 | 62.47 ± 0.11 | 70.71 ± 0.03 | 63.38 ± 0.06 | |
| SAFD-CS (Label) | 70.46 ± 0.17 | 62.96 ± 0.10 | 70.72 ± 0.04 | 63.44 ± 0.18 | |
| KITTI | Teacher | 61.93 | 54.82 | 64.01 | 56.91 |
| Student | 55.97 | 48.54 | 55.97 | 48.54 | |
| KD [9] | 54.94 ± 0.16 | 48.23 ± 0.10 | 56.94 ± 0.40 | 49.12 ± 0.37 | |
| AT [11] | 56.51 ± 0.49 | 48.87 ± 0.40 | 56.80 ± 0.24 | 49.10 ± 0.26 | |
| SP [35] | 56.67 ± 0.42 | 48.95 ± 0.37 | 56.61 ± 0.44 | 48.91 ± 0.43 | |
| RKD [36] | 59.38 ± 0.50 | 51.31 ± 0.33 | 58.77 ± 0.46 | 50.97 ± 0.41 | |
| ReviewKD [34] | 57.29 ± 0.32 | 49.60 ± 0.29 | 57.33 ± 0.07 | 49.63 ± 0.15 | |
| GDKD [52] | 56.60 ± 0.33 | 48.94 ± 0.35 | 56.71 ± 0.36 | 49.01 ± 0.35 | |
| AttnFD [15] | 59.59 ± 0.14 | 52.13 ± 0.08 | 59.88 ± 0.37 | 52.18 ± 0.26 | |
| SAFD-CS (Label) | 61.03 ± 0.14 | 52.93 ± 0.13 | 60.40 ± 0.29 | 52.50 ± 0.16 | |
| T: SegFormer-B2 | T: SegFormer-B4 | |||
|---|---|---|---|---|
| DB | Method | Format | mIoU | mIoU |
| CamVid | SAFD-C | w/o SG | 62.65 ± 0.10 | 63.46 ± 0.11 |
| Label | 62.85 ± 0.09 (+0.20) | 63.42 ± 0.11 (−0.04) | ||
| Keywords | 62.90 ± 0.05 (+0.25) | 63.42 ± 0.03 (−0.04) | ||
| Sentence | 62.95 ± 0.07 (+0.30) | 63.33 ± 0.17 (−0.13) | ||
| SAFD-S | w/o SG | 62.59 ± 0.19 | 63.42 ± 0.02 | |
| Label | 62.93 ± 0.14 (+0.34) | 63.53 ± 0.29 (+0.11) | ||
| Keywords | 62.85 ± 0.21 (+0.26) | 63.44 ± 0.08 (+0.02) | ||
| Sentence | 62.74 ± 0.13 (+0.15) | 63.47 ± 0.21 (+0.05) | ||
| SAFD-CS | w/o SG | 62.47 ± 0.11 | 63.38 ± 0.06 | |
| Label | 62.96 ± 0.10 (+0.49) | 63.44 ± 0.18 (+0.06) | ||
| Keywords | 62.87 ± 0.13 (+0.40) | 63.45 ± 0.16 (+0.07) | ||
| Sentence | 62.69 ± 0.13 (+0.22) | 63.33 ± 0.24 (−0.05) | ||
| KITTI | SAFD-C | w/o SG | 52.51 ± 0.09 | 52.45 ± 0.33 |
| Label | 52.52 ± 0.21 (+0.01) | 52.41 ± 0.16 (−0.04) | ||
| Keywords | 52.58 ± 0.25 (+0.07) | 53.09 ± 0.31 (+0.64) | ||
| Sentence | 52.63 ± 0.40 (+0.12) | 52.72 ± 0.53 (+0.27) | ||
| SAFD-S | w/o SG | 52.24 ± 0.13 | 52.09 ± 0.18 | |
| Label | 52.29 ± 0.36 (+0.05) | 52.28 ± 0.22 (+0.19) | ||
| Keywords | 52.62 ± 0.13 (+0.38) | 52.84 ± 0.00 (+0.75) | ||
| Sentence | 52.58 ± 0.27 (+0.34) | 52.79 ± 0.16 (+0.70) | ||
| SAFD-CS | w/o SG | 52.13 ± 0.08 | 52.18 ± 0.26 | |
| Label | 52.93 ± 0.13 (+0.80) | 52.50 ± 0.16 (+0.32) | ||
| Keywords | 52.47 ± 0.23 (+0.34) | 52.02 ± 0.17 (−0.16) | ||
| Sentence | 52.54 ± 0.06 (+0.41) | 52.37 ± 0.22 (+0.19) |
| Method | Text Encoder | Configuration | Pretraining | CamVid | KITTI |
|---|---|---|---|---|---|
| Student | - | - | 60.17 | 48.54 | |
| SAFD-CS | (w/o SG) | - | - | 62.47 ± 0.11 | 52.13 ± 0.08 |
| OpenCLIP [51] | ViT-B/16 | LAION-2B | 62.96 ± 0.10 | 52.93 ± 0.13 | |
| ViT-B/32 | LAION-2B | 62.75 ± 0.07 | 52.57 ± 0.03 | ||
| CLIP [26] | ViT-B/16 | OpenAI | 62.91 ± 0.26 | 52.66 ± 0.03 | |
| ViT-B/32 | OpenAI | 62.96 ± 0.06 | 52.71 ± 0.08 | ||
| SigLIP [53] | ViT-B/16 | WebLI | 62.92 ± 0.21 | 52.69 ± 0.14 | |
| BERT [54] | Base (uncased) | BooksCorpus + Wikipedia | 62.79 ± 0.16 | 52.50 ± 0.11 | |
| DB | Proportion | Train Loss | Validation Loss ↓ | Train–Val Loss Gap ()↓ | Test mIoU (%) ↑ |
|---|---|---|---|---|---|
| CamVid | 100% | 0.1489 | 0.2338 | 0.0849 | 62.96 (+2.79) |
| 50% | 0.1305 | 0.2596 | 0.1291 | 61.39 (+3.05) | |
| 30% | 0.1165 | 0.2954 | 0.1789 | 60.23 (+3.43) | |
| KITTI | 100% | 0.2248 | 0.4510 | 0.2262 | 52.93 (+4.39) |
| 50% | 0.1984 | 0.4901 | 0.2917 | 51.59 (+5.70) | |
| 30% | 0.1159 | 0.5595 | 0.4436 | 49.60 (+6.46) |
| w/o SG | Label | Keywords | Sentence | |||||
|---|---|---|---|---|---|---|---|---|
| Method | ECE | NLL | ECE | NLL | ECE | NLL | ECE | NLL |
| SAFD-C | 3.81 | 0.4132 | 3.84 | 0.4146 | 3.44 | 0.4118 | 3.47 | 0.4140 |
| SAFD-S | 3.77 | 0.4150 | 3.76 | 0.4186 | 3.58 | 0.4146 | 3.41 | 0.4129 |
| SAFD-CS | 3.83 | 0.4140 | 3.14 | 0.4110 | 3.32 | 0.4136 | 3.55 | 0.4142 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Ryu, J.H.; Cha, S.W.; Jeon, E.S. Exploring the Synergy Between Language Semantic Guidance and Visual Attention in Knowledge Distillation for Semantic Segmentation Under a Limited Field of View. Biomimetics 2026, 11, 665. https://doi.org/10.3390/biomimetics11090665
Ryu JH, Cha SW, Jeon ES. Exploring the Synergy Between Language Semantic Guidance and Visual Attention in Knowledge Distillation for Semantic Segmentation Under a Limited Field of View. Biomimetics. 2026; 11(9):665. https://doi.org/10.3390/biomimetics11090665
Chicago/Turabian StyleRyu, Jin Hyeok, Seung Woo Cha, and Eun Som Jeon. 2026. "Exploring the Synergy Between Language Semantic Guidance and Visual Attention in Knowledge Distillation for Semantic Segmentation Under a Limited Field of View" Biomimetics 11, no. 9: 665. https://doi.org/10.3390/biomimetics11090665
APA StyleRyu, J. H., Cha, S. W., & Jeon, E. S. (2026). Exploring the Synergy Between Language Semantic Guidance and Visual Attention in Knowledge Distillation for Semantic Segmentation Under a Limited Field of View. Biomimetics, 11(9), 665. https://doi.org/10.3390/biomimetics11090665

