Lightweight Semantic Segmentation Algorithm Based on Gated Visual State Space Models
Abstract
1. Introduction
- We propose RainMamba, a novel multi-modal semantic segmentation framework that transcends conventional module-splicing by deeply coupling physics-based noise priors with the linear-time Visual State Space Model (VMamba), tailored specifically for adverse weather perception.
- We design a Physical Perception Feature Embedding Module that explicitly integrates the Beer–Lambert Law into the network input. By capturing the physical mapping between distance, intensity attenuation, and environmental noise, it effectively suppresses rain and fog interference at the source.
- We introduce an Uncertainty-Weighted Cross-Modal Correction mechanism. Rather than using computationally expensive cross-attention, we utilize a lightweight gating strategy to dynamically align and fuse RGB texture features, compensating for geometric details lost in 2D spherical projection while strictly preserving linear complexity.
- Extensive evaluations demonstrate exceptional Sim-to-Real generalization. Zero-shot tests on the real-world SemanticSTF dataset prove that RainMamba achieves highly competitive robustness comparable to heavy-weight, state-of-the-art architectures (e.g., PTv3) while operating at real-time speeds with a significantly lower parameter footprint.
2. Proposed Method
2.1. Overall Architecture
2.2. Physical Perception Feature Embedding Module
2.2.1. Spherical Projection and Feature Tensor
2.2.2. Feature Extraction Based on Asymmetric Convolution
- A convolution kernel is used to extract continuous geometric features in the horizontal direction;
- A convolution kernel is used to suppress discrete environmental noise in the vertical direction.
2.3. Uncertainty-Weighted Cross-Modal Correction Mechanism
2.3.1. Multi-Source Data Spatial Alignment
2.3.2. Gated Fusion Network and Dynamic Correction
2.4. RainMamba Context Inference Engine
2.4.1. Linear Complexity and Real-Time Guarantee
2.4.2. SS2D Scanning and Panoramic Perception Verification
2.4.3. Post-Processing and 3D Reprojection
3. System Implementation
3.1. Hardware and Simulation Environment
- Hardware Configuration: To comprehensively evaluate both theoretical accuracy and edge-deployment feasibility, we employed a heterogeneous computing setup. Model training and the accuracy evaluation of resource-intensive baselines (e.g., PTv3) were conducted on a high-performance workstation equipped with an NVIDIA GeForce RTX 3090 GPU (24 GB VRAM) to prevent out-of-memory bottlenecks. However, to rigorously validate the real-time inference speed (FPS) and memory footprint under strict automotive hardware constraints, all inference profiling was intentionally executed on a mobile-grade platform driven by an AMD Ryzen 7 7840H CPU and an NVIDIA GeForce RTX 4060 Laptop GPU (8 GB VRAM) running CUDA 11.8. This constrained setup closely mimics the computing boundaries of real-world embedded nodes, providing a highly persuasive environment for validating our system’s 83 FPS real-time capability.
- Virtual Sensor Injection: To address the lack of real sensor data in non-vehicle environments, a ROS (Humble Hawksbill) 2 data playback node was developed. It parses the SemanticKITTI dataset into real-time PointCloud2 streams, injecting rain/fog noise via timestamp synchronization to decouple the perception algorithm from physical hardware.
3.2. Software Architecture and Optimization
- Modular Design: The system consists of three independent nodes: a Data Preprocessing Node (Spherical Projection), an Inference Engine Node (TensorRT Acceleration), and a Post-processing Node (KNN Smoothing).
- Communication Optimization: A zero-copy transmission strategy based on shared memory is deployed for intra-node communication, avoiding serialization overhead and reducing end-to-end latency.
- Real-time Optimization: The system employs mixed-precision inference (FP16), reducing memory usage by approximately 40%. Additionally, an asynchronous pipeline overlaps CPU data loading with GPU inference, maximizing hardware saturation.
3.3. Visualization Interface
4. Experiments and Results
4.1. Experimental Setup
- mIoU (mean Intersection over Union): Evaluates the segmentation accuracy.
- FPS (Frames Per Second): Evaluates the data throughput and real-time inference capability.
- Params and Memory Usage: Evaluates the computational footprint and resource consumption, which is critical for comprehensively comparing lightweight edge-oriented models against state-of-the-art heavy-weight architectures (e.g., PTv3 [8]).
4.2. Quantitative Validation of Physical Noise Distribution
4.3. Quantitative Performance Analysis
4.4. Module Effectiveness Ablation Studies
- Denoising Effect of Physical Perception: Comparing Experiments A and B, the introduction of the PE module improved mIoU by 3.7%. This significant gain is mainly attributed to the effective filtering of vertical rain noise by asymmetric convolutions, resulting in sharper edges for ground and buildings.
- Geometric Completion of Multi-modality: Comparing Experiments A and C, the CMF module brought a 4.3% improvement. In particular, in areas where LiDAR signals are missing due to specular reflection, RGB texture information successfully guided the correct inference of semantic categories.
- Synergistic Effect: The complete model (Experiment D) achieved the best performance of 58.2%, proving that “physical denoising” and “visual completion” are highly complementary when processing adverse weather data.
4.5. Robustness Analysis at Different Distances
4.6. Class-Wise Performance Analysis
4.7. Resource Consumption Analysis
4.8. Qualitative Analysis
4.9. Zero-Shot Generalization on Real-World SemanticSTF
4.9.1. Robustness Against Optical Degradation
4.9.2. Efficiency and Safety-Critical Analysis
5. Conclusions
Author Contributions
Funding
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Behley, J.; Garbade, M.; Milioto, A.; Quenzel, J.; Behnke, S.; Stachniss, C.; Gall, J. SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR Sequences. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 9296–9306. [Google Scholar]
- Zhu, X.; Zhou, H.; Wang, T.; Hong, F.; Ma, Y.; Li, W.; Li, H.; Lin, D. Cylindrical and Asymmetrical 3D Convolution Networks for LiDAR Segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021; pp. 9934–9943. [Google Scholar]
- Wu, X.; Hou, Y.; Huang, X.; Geng, S.; Hofstetter, H. TASeg: Temporal Aggregation Network for LiDAR Semantic Segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 17–21 June 2024; pp. 15311–15320. [Google Scholar]
- Milioto, A.; Vizzo, I.; Behley, J.; Stachniss, C. RangeNet++: Fast and Accurate LiDAR Semantic Segmentation. In Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Macau, China, 4–8 November 2019; pp. 4213–4220. [Google Scholar]
- Xu, C.; Wu, B.; Wang, Z.; Zhan, W.; Vajda, P.; Keutzer, K.; Tomizuka, M. SqueezeSegV3: Spatially-Adaptive Convolution for Efficient Point-Cloud Segmentation. In Proceedings of the European Conference on Computer Vision (ECCV), Glasgow, UK, 23–28 August 2020; pp. 1–16. [Google Scholar]
- Zhang, Y.; Zhou, Z.; David, P.; Yue, X.; Xi, Z.; Gong, B.; Foroosh, H. PolarNet: An Improved Grid Representation for Online LiDAR Point Clouds Semantic Segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 9601–9610. [Google Scholar]
- Gu, A.; Dao, T. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. arXiv 2023, arXiv:2312.00752. [Google Scholar] [CrossRef]
- Wu, X.; Jiang, L.; Wang, P.-S.; Liu, Z.; Liu, X.; Qiao, Y.; Ouyang, W.; He, T.; Zhao, H. Point Transformer V3: Simpler, Faster, Stronger. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 17–21 June 2024; pp. 4840–4851. [Google Scholar]
- Liu, Y.; Tian, Y.; Zhao, Y.; Yu, H.; Xie, L.; Wang, Y.; Ye, Q.; Liu, Y. VMamba: Visual State Space Model. In Advances in Neural Information Processing Systems 37; Curran Associates, Inc.: Red Hook, NY, USA, 2024; pp. 103031–103063. [Google Scholar]
- Park, J.; Kim, K.; Shim, H. Rethinking Data Augmentation for Robust LiDAR Semantic Segmentation in Adverse Weather. In Proceedings of the European Conference on Computer Vision (ECCV), Cham, Switzerland, 29 September–4 October 2024; pp. 320–336. [Google Scholar]
- Hahner, M.; Sakaridis, C.; Dai, D.; Van Gool, L. Fog Simulation on Real LiDAR Point Clouds for 3D Object Detection in Adverse Weather. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 10–17 October 2021; pp. 15263–15272. [Google Scholar]
- Xiao, A.; Huang, J.; Xuan, W.; Ren, R.; Liu, K.; Guan, D.; El Saddik, A.; Lu, S.; Xing, E. 3D Semantic Segmentation in the Wild: Learning Generalized Models for Adverse-Condition Point Clouds. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2023; pp. 9382–9392. [Google Scholar]
- Romera, E.; Alvarez, J.M.; Bergasa, L.M.; Arroyo, R. ERFNet: Efficient Residual Factorized ConvNet for Real-Time Semantic Segmentation. IEEE Trans. Intell. Transp. Syst. 2018, 19, 263–272. [Google Scholar] [CrossRef]
- Vora, S.; Lang, A.H.; Helou, B.; Beijbom, O. PointPainting: Sequential Fusion for 3D Object Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 4603–4611. [Google Scholar]
- Li, S.; Chen, X.; Liu, Y.; Dai, D.; Stachniss, C.; Gall, J. Multi-scale Interaction for Real-time LiDAR Data Segmentation on an Embedded Platform. arXiv 2021, arXiv:2108.09242. [Google Scholar] [CrossRef]
- Cortinhal, T.; Tzelepis, G.; Aksoy, E.E. SalsaNext: Fast, Uncertainty-aware Semantic Segmentation of LiDAR Point Clouds for Autonomous Driving. arXiv 2020, arXiv:2003.03653. [Google Scholar]
- Cheng, H.; Han, X.; Xiao, G. Cenet: Toward Concise and Efficient Lidar Semantic Segmentation for Autonomous Driving. In Proceedings of the IEEE International Conference on Multimedia and Expo (ICME), Taipei, Taiwan, 18–22 July 2022; pp. 1–6. [Google Scholar]







| Method | Params (M) | GFLOPs | FPS | Car IoU | Road IoU | mIoU (%) |
|---|---|---|---|---|---|---|
| RangeNet++ | 26.5 | 185.3 | 90 | 82.1 | 88.5 | 52.4 |
| SqueezeSegV3 | 19.2 | 145.8 | 110 | 84.5 | 89.1 | 48.6 |
| PolarNet | 13.7 | 112.5 | 65 | 89.2 | 90.5 | 54.3 |
| Cenet | 12.5 | 98.4 | 72 | 90.1 | 91.0 | 55.6 |
| PTv3 | 24.8 | 210.6 | 14 | 96.2 | 94.8 | 63.8 |
| RainMamba (Ours) | 9.8 | 89.5 | 83 | 92.4 | 91.8 | 58.2 |
| Exp. ID | VMamba Backbone | Physical Module (PE) | Cross-Modal Fusion (CMF) | mIoU (%) | Gain |
|---|---|---|---|---|---|
| A | ✔ | 50.8 | - | ||
| B | ✔ | ✔ | 54.5 | +3.7% | |
| C | ✔ | ✔ | 55.1 | +4.3% | |
| D (Ours) | ✔ | ✔ | ✔ | 58.2 | +7.4% |
| Method | Overall | Near (0–15 m) | Medium (15–30 m) | Far (30–50 m) |
|---|---|---|---|---|
| RangeNet++ | 52.4 | 68.5 | 45.2 | 31.6 |
| PolarNet | 54.3 | 70.1 | 48.5 | 35.2 |
| RainMamba (Ours) | 58.2 | 74.3 | 53.8 | 42.1 |
| Method | Overall mIoU | Weather Breakdown (mIoU) | |||
|---|---|---|---|---|---|
| Light Fog | Dense Fog | Rain | Snow | ||
| RangeNet++ | 31.5 | 35.2 | 24.8 | 30.1 | 32.4 |
| PolarNet | 34.8 | 38.5 | 28.1 | 33.5 | 35.2 |
| CENet | 36.2 | 40.1 | 30.5 | 35.8 | 37.1 |
| PTv3 | 44.8 | 48.5 | 45.1 | 43.2 | 42.5 |
| RainMamba (Ours) | 43.9 | 47.8 | 44.2 | 41.5 | 40.8 |
| Method | Params (M) | Speed (FPS) | Safety-Critical Classes (IoU) | ||
|---|---|---|---|---|---|
| Car | Person | Road | |||
| RangeNet++ | 26.5 | 90 | 45.2 | 12.5 | 68.1 |
| PolarNet | 13.7 | 65 | 52.4 | 18.2 | 72.5 |
| CENet | 12.5 | 72 | 55.6 | 21.4 | 74.8 |
| PTv3 | 24.8 | 14 | 72.8 | 40.5 | 82.1 |
| RainMamba (Ours) | 9.8 | 83 | 72.1 | 39.2 | 80.5 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Di, K.; Cheng, J.; Zhang, L.; Bao, Y. Lightweight Semantic Segmentation Algorithm Based on Gated Visual State Space Models. Electronics 2026, 15, 1175. https://doi.org/10.3390/electronics15061175
Di K, Cheng J, Zhang L, Bao Y. Lightweight Semantic Segmentation Algorithm Based on Gated Visual State Space Models. Electronics. 2026; 15(6):1175. https://doi.org/10.3390/electronics15061175
Chicago/Turabian StyleDi, Kui, Jinming Cheng, Lili Zhang, and Yubin Bao. 2026. "Lightweight Semantic Segmentation Algorithm Based on Gated Visual State Space Models" Electronics 15, no. 6: 1175. https://doi.org/10.3390/electronics15061175
APA StyleDi, K., Cheng, J., Zhang, L., & Bao, Y. (2026). Lightweight Semantic Segmentation Algorithm Based on Gated Visual State Space Models. Electronics, 15(6), 1175. https://doi.org/10.3390/electronics15061175
