Lidar–Vision Depth Fusion for Robust Loop Closure Detection in SLAM Systems
Abstract
1. Introduction
- Spatiotemporal Alignment and Depth Completion Framework: A unified process is designed to achieve spatial alignment of LiDAR point clouds and camera images via extrinsic calibration, handle sampling rate discrepancies through timestamp interpolation, and complete depth estimation of visual keypoints using neighborhood search. This enables the construction of 3D visual features as a foundation for cross-modal fusion.
- Hybrid Feature Descriptor Construction: A fusion descriptor combining LiDAR geometric triangle descriptors (representing spatial topology) and visual BRIEF descriptors (representing local texture) is developed. Efficient loop candidate retrieval is achieved via hash indexing, balancing matching efficiency and discriminability. An improved RANSAC-based geometric verification method is further introduced to suppress noise and reduce false matches.
- Comprehensive Evaluation on Public Datasets: Extensive experiments are conducted on the KITTI and NCLT datasets to validate the proposed algorithm’s effectiveness across urban, mixed indoor–outdoor, and seasonally varying environments, demonstrating its potential as a robust solution for long-term autonomous navigation in SLAM systems.
2. Related Work
2.1. Loop Closure Detection Based on Single Modality
2.2. LiDAR-Vision Fusion for Loop Closure Detection
3. Method
3.1. Visual Feature Extraction and Depth Completion
3.2. Multimodal Feature Fusion and Matching
3.3. Improved RANSAC-Based Geometric Verification
4. Experiments
4.1. Experimental Setup and Datasets
4.2. Results on the KITTI Dataset
4.3. Results on the NCLT Dataset
4.4. Runtime Analysis
5. Conclusions
- A tightly coupled LiDAR–vision fusion framework is developed, which achieves effective geometry–texture association through spatiotemporal alignment, depth completion, hybrid feature fusion, and improved RANSAC-based geometric verification.
- Quantitative evaluations on public datasets demonstrate the effectiveness of the proposed approach. On the KITTI dataset, the proposed method achieves an average F1-score of 85.28%, outperforming LiDAR-only, vision-only baselines, and representative multimodal methods.
- On the more challenging NCLT dataset, which features long-term operation, seasonal changes, and mixed indoor–outdoor environments, the proposed method attains an average F1-score of 77.63%, demonstrating improved robustness in scenarios where geometric or visual cues alone become unreliable.
- Investigating dynamic feature filtering strategies to explicitly handle highly dynamic environments.
- Exploring adaptive depth fusion mechanisms to improve robustness under sparse or noisy depth observations.
- Optimizing computational efficiency to further enhance real-time performance and facilitate deployment on embedded robotic platforms.
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Mur-Artal, R.; Montiel, J.M.M.; Tardós, J.D. ORB-SLAM: A Versatile and Accurate Monocular SLAM System. IEEE Trans. Robot. 2015, 31, 1147–1163. [Google Scholar] [CrossRef]
- Qin, T.; Li, P.; Shen, S. VINS-Mono: A Robust and Versatile Monocular Visual-Inertial State Estimator. IEEE Trans. Robot. 2018, 34, 1004–1020. [Google Scholar] [CrossRef]
- Peng, R.; Gong, C.; Zhao, S. Multi-Sensor Information Fusion with Multi-Scale Adaptive Graph Convolutional Networks for Abnormal Vibration Diagnosis of Rolling Mill. Machines 2025, 13, 30. [Google Scholar] [CrossRef]
- Wen, S.; Long, Y.; Li, P.; Wang, B.; Qiu, T.Z. Semantic Constellation Place Recognition Algorithm Based on Scene Text. IEEE Trans. Instrum. Meas. 2025, 74, 1–9. [Google Scholar] [CrossRef]
- Kim, G.; Kim, A. Scan Context: Egocentric Spatial Descriptor for Place Recognition Within 3D Point Cloud Map. In Proceedings of the 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Madrid, Spain, 1–5 October 2018; pp. 4802–4809. [Google Scholar]
- Kim, G.; Choi, S.; Kim, A. Scan Context++: Structural Place Recognition Robust to Rotation and Lateral Variations in Urban Environments. IEEE Trans. Robot. 2022, 38, 1856–1874. [Google Scholar] [CrossRef]
- Jiang, B.; Shen, S. Contour Context: Abstract Structural Distribution for 3D LiDAR Loop Detection and Metric Pose Estimation. In Proceedings of the 2023 IEEE International Conference on Robotics and Automation (ICRA), London, UK, 29 May–2 June 2023; pp. 8386–8392. [Google Scholar]
- Cui, Y.; Chen, X.; Zhang, Y.; Dong, J.; Wu, Q.; Zhu, F. BoW3D: Bag of Words for Real-Time Loop Closing in 3D LiDAR SLAM. IEEE Robot. Autom. Lett. 2023, 8, 2828–2835. [Google Scholar] [CrossRef]
- Galvez-López, D.; Tardos, J.D. Bags of Binary Words for Fast Place Recognition in Image Sequences. IEEE Trans. Robot. 2012, 28, 1188–1197. [Google Scholar] [CrossRef]
- Wen, S.; Tao, S.; Liu, X.; Babiarz, A.; Yu, F.R. CD-SLAM: A Real-Time Stereo Visual–Inertial SLAM for Complex Dynamic Environments With Semantic and Geometric Information. IEEE Trans. Instrum. Meas. 2024, 73, 1–8. [Google Scholar] [CrossRef]
- Liu, Z.; Tang, H.; Amini, A.; Yang, X.; Mao, H.; Rus, D.L.; Han, S. BEVFusion: Multi-Task Multi-Sensor Fusion with Unified Bird’s-Eye View Representation. In Proceedings of the 2023 IEEE International Conference on Robotics and Automation (ICRA), London, UK, 29 May–2 June 2023; pp. 2774–2781. [Google Scholar]
- Pan, Y.; Xu, X.; Li, W.; Cui, Y.; Wang, Y.; Xiong, R. CORAL: Colored structural representation for bi-modal place recognition. In Proceedings of the 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Prague, Czech Republic, 27 September–1 October 2021; pp. 2084–2091. [Google Scholar]
- Vora, S.; Lang, A.H.; Helou, B.; Beijbom, O. PointPainting: Sequential Fusion for 3D Object Detection. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 4603–4611. [Google Scholar]
- Zhao, L.; Zhou, H.; Zhu, X.; Song, X.; Li, H.; Tao, W. LIF-Seg: LiDAR and Camera Image Fusion for 3D LiDAR Semantic Segmentation. IEEE Trans. Multimed. 2024, 26, 1158–1168. [Google Scholar] [CrossRef]
- Bai, X.; Hu, Z.; Zhu, X.; Huang, Q.; Chen, Y.; Fu, H.; Tai, C.L. TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with Transformers. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 1080–1089. [Google Scholar] [CrossRef]
- Lv, X.; He, Z.; Yang, Y.; Nie, J.; Dong, Z.; Wang, S.; Gao, M. MSF-SLAM: Multi-Sensor-Fusion-Based Simultaneous Localization and Mapping for Complex Dynamic Environments. IEEE Trans. Intell. Transp. Syst. 2024, 25, 19699–19713. [Google Scholar] [CrossRef]
- Zhao, X.; Wen, C.; Manoj Prakhya, S.; Yin, H.; Zhou, R.; Sun, Y.; Xu, J.; Bai, H.; Wang, Y. Multimodal Features and Accurate Place Recognition With Robust Optimization for Lidar–Visual–Inertial SLAM. IEEE Trans. Instrum. Meas. 2024, 73, 1–16. [Google Scholar] [CrossRef]
- Besl, P.; McKay, N.D. A method for registration of 3-D shapes. IEEE Trans. Pattern Anal. Mach. Intell. 1992, 14, 239–256. [Google Scholar] [CrossRef]
- Biber, P.; Strasser, W. The normal distributions transform: A new approach to laser scan matching. In Proceedings of the Proceedings 2003 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2003) (Cat. No.03CH37453), Las Vegas, Nevada, USA, 27 October–1 November 2003; Volume 3, pp. 2743–2748. [Google Scholar]
- Cui, Y.; Zhang, Y.; Dong, J.; Sun, H.; Chen, X.; Zhu, F. LinK3D: Linear Keypoints Representation for 3D LiDAR Point Cloud. IEEE Robot. Autom. Lett. 2024, 9, 2128–2135. [Google Scholar]
- Gupta, S.; Guadagnino, T.; Mersch, B.; Vizzo, I.; Stachniss, C. Effectively Detecting Loop Closures using Point Cloud Density Maps. In Proceedings of the 2024 IEEE International Conference on Robotics and Automation (ICRA), Yokohama, Japan, 13–17 May 2024; pp. 10260–10266. [Google Scholar]
- Pirotti, F.; Ravanelli, R.; Fissore, F.; Masiero, A. Implementation and assessment of two density-based outlier detection methods over large spatial point clouds. Open Geospat. Data Softw. Stand. 2018, 3, 14. [Google Scholar] [CrossRef]
- Bay, H.; Tuytelaars, T.; Van Gool, L. SURF: Speeded Up Robust Features. In Proceedings of the Computer Vision—ECCV 2006; Leonardis, A., Bischof, H., Pinz, A., Eds.; Springer: Berlin/Heidelberg, Germany, 2006; pp. 404–417. [Google Scholar]
- Dalal, N.; Triggs, B. Histograms of oriented gradients for human detection. In Proceedings of the 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05), San Diego, CA, USA, 20–26 June 2005; Volume 1, pp. 886–893. [Google Scholar]
- Bai, Y.; Guo, L.; Jin, L.; Huang, Q. A novel feature extraction method using Pyramid Histogram of Orientation Gradients for smile recognition. In Proceedings of the 2009 16th IEEE International Conference on Image Processing (ICIP), Cairo, Egypt, 7–10 November 2009; pp. 3305–3308. [Google Scholar]
- Cummins, M.; Newman, P. Appearance-only SLAM at large scale with FAB-MAP 2.0. Int. J. Robot. Res. 2011, 30, 1100–1123. [Google Scholar]
- Zou, Z.; Zheng, C.; Yuan, C.; Zhou, S.; Xue, K.; Zhang, F. iBTC: An Image-Assisting Binary and Triangle Combined Descriptor for Place Recognition by Fusing LiDAR and Camera Measurements. IEEE Robot. Autom. Lett. 2024, 9, 10858–10865. [Google Scholar] [CrossRef]
- Komorowski, J.; Wysoczańska, M.; Trzcinski, T. MinkLoc++: Lidar and Monocular Image Fusion for Place Recognition. In Proceedings of the 2021 International Joint Conference on Neural Networks (IJCNN), Shenzhen, China, 18–22 July 2021; pp. 1–8. [Google Scholar]
- Zeng, Y.; Zhang, D.; Wang, C.; Miao, Z.; Liu, T.; Zhan, X.; Hao, D.; Ma, C. LIFT: Learning 4D LiDAR Image Fusion Transformer for 3D Object Detection. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 17151–17160. [Google Scholar]
- Harris, C.G.; Stephens, M.J. A Combined Corner and Edge Detector. In Proceedings of the Alvey Vision Conference, Manchester, UK, 31 August–2 September 1988. [Google Scholar]
- Fischler, M.A.; Bolles, R.C. Random sample consensus: A paradigm for model fitting with applications to image analysis and automated cartography. Commun. ACM 1981, 24, 381–395. [Google Scholar] [CrossRef]
- Geiger, A.; Lenz, P.; Stiller, C.; Urtasun, R. Vision meets robotics: The KITTI dataset. Int. J. Rob. Res. 2013, 32, 1231–1237. [Google Scholar]
- Carlevaris-Bianco, N.; Ushani, A.K.; Eustice, R.M. University of Michigan North Campus long-term vision and lidar dataset. Int. J. Robot. Res. 2015, 35, 1023–1035. [Google Scholar] [CrossRef]









| Method | Modality | Representation |
|---|---|---|
| SC/SC++ | LiDAR | Global geometric descriptor |
| BoW3D | LiDAR | BoW LiDAR fertures |
| Map Closure | LiADR | Density image BEV representation |
| HOG/PHOG | Vision | Statistics of gradient orientation distribution in images |
| DBoW2 | Vision | BoW visual features |
| iBTC | LiDAR + Vision | Binary geometric descriptor |
| CoRAL/MinkLoc++ | LiDAR + Vision | CNN-based method |
| BEVFusion | LiDAR + Vision | Fused BEV features |
| Proposed | LiDAR + Vision | Geometry–texture hybrid |
| Method | KITTI00 | KITTI02 | KITTI05 | Average |
|---|---|---|---|---|
| Cont2 | 0.7622 | 0.6840 | 0.7678 | 0.7380 |
| Map Closure | 0.8137 | 0.8015 | 0.8052 | 0.8068 |
| DBoW2 | 0.6537 | 0.6497 | 0.7597 | 0.6877 |
| iBTC | 0.8128 | 0.8250 | 0.8056 | 0.8145 |
| Proposed | 0.8561 | 0.8415 | 0.8608 | 0.8528 |
| Dataset Sequence | Proposed Method /Without Visual | Proposed Method | ||
|---|---|---|---|---|
|
Detected
Loops |
Max
Similarity |
Detected
Loops |
Max
Similarity | |
| KITTI00 | 351 | 0.972 | 359 | 0.979 |
| KITTI02 | 154 | 0.953 | 167 | 0.956 |
| KITTI05 | 252 | 0.948 | 255 | 0.951 |
| KITTI08 | 167 | 0.961 | 169 | 0.961 |
| Max | Mean | Median | Min | RMSE | |
|---|---|---|---|---|---|
| Before Correction | 0.832 | 0.728 | 0.725 | 0.462 | 0.667 |
| After Correction | 0.787 | 0.691 | 0.629 | 0.467 | 0.630 |
| Method | NCLT1 | NCLT2 | Average |
|---|---|---|---|
| Cont2 | 0.6408 | 0.6706 | 0.6557 |
| Map Closure | 0.6567 | 0.6927 | 0.6747 |
| DBoW2 | 0.1908 | 0.4069 | 0.2989 |
| iBTC | 0.7031 | 0.7772 | 0.7402 |
| Proposed | 0.7642 | 0.7884 | 0.7763 |
| Dataset Sequence | Proposed Method /Without Visual | Proposed Method | ||
|---|---|---|---|---|
|
Detected
Loops |
Max
Similarity |
Detected
Loops |
Max
Similarity | |
| NCLT1 | 821 | 0.905 | 887 | 0.907 |
| NCLT2 | 623 | 0.891 | 661 | 0.905 |
| NCLT3 | 388 | 0.913 | 412 | 0.919 |
| NCLT4 | 411 | 0.910 | 435 | 0.929 |
| Max | Mean | Median | Min | RMSE | |
|---|---|---|---|---|---|
| Before Correction | 1.352 | 1.254 | 1.282 | 0.991 | 1.092 |
| After Correction | 1.215 | 1.058 | 1.037 | 0.942 | 0.928 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Liu, B.; Wu, P.; Chen, R.; Zheng, Y.; Li, M. Lidar–Vision Depth Fusion for Robust Loop Closure Detection in SLAM Systems. Machines 2026, 14, 282. https://doi.org/10.3390/machines14030282
Liu B, Wu P, Chen R, Zheng Y, Li M. Lidar–Vision Depth Fusion for Robust Loop Closure Detection in SLAM Systems. Machines. 2026; 14(3):282. https://doi.org/10.3390/machines14030282
Chicago/Turabian StyleLiu, Bingzhuo, Panlong Wu, Rongting Chen, Yidan Zheng, and Mengyu Li. 2026. "Lidar–Vision Depth Fusion for Robust Loop Closure Detection in SLAM Systems" Machines 14, no. 3: 282. https://doi.org/10.3390/machines14030282
APA StyleLiu, B., Wu, P., Chen, R., Zheng, Y., & Li, M. (2026). Lidar–Vision Depth Fusion for Robust Loop Closure Detection in SLAM Systems. Machines, 14(3), 282. https://doi.org/10.3390/machines14030282

