Mobile Robot Localization and SLAM: A Critical Review of Sensors, Multi-Sensor Fusion, and Neural Representations
Abstract
1. Introduction
1.1. Contributions of This Survey
- Unified cross-domain perspective. We review localization and SLAM across ground, aerial, marine, and legged robots, highlighting shared challenges and domain-specific constraints.
- Integration of classical and learning-based pipelines. We provide a structured comparison between probabilistic estimation, geometric SLAM, and emerging neural and foundation-model-based approaches.
- Comprehensive sensor taxonomy. We present an updated overview of sensing modalities, including event cameras, 4D radar, UWB, and multi-modal fusion strategies.
- Analysis of neural SLAM and implicit mapping. We review recent advances in neural radiance fields, differentiable rendering, and learned localization pipelines.
- Benchmark and dataset consolidation. We summarize widely used datasets and evaluation protocols across multiple robotic domains.
- Critical discussion and open research challenges. We identify limitations of current systems and outline future research directions toward robust, scalable, and certifiable localization.
1.2. Operational Definitions
- Real-time: A localization system is considered real-time if it produces a pose estimate at a rate sufficient to close the control loop of the robot without inducing instability. For most ground and aerial platforms, this implies a minimum update rate of 10 Hz; for high-speed UAVs it implies ≥ 50 Hz. Systems requiring offline post-processing, batch optimization, or GPU inference exceeding one frame latency are not classified as real-time.
- Robustness: The ability of a system to maintain bounded localization error across a defined range of environmental perturbations (illumination changes, dynamic objects, adverse weather, and sensor degradation) without manual intervention or re-initialization. Robustness is environment- and sensor-specific; a system may be robust indoors but fragile outdoors.
- Generalization: The capacity of a learned model to produce reliable outputs in environments, sensor configurations, or conditions not represented in the training distribution. Generalization is evaluated by testing on held-out scenes or cross-dataset protocols.
- Deployment-ready: A system is deployment-ready if it has been demonstrated on a physical robot platform under representative operational conditions (not only in controlled laboratory settings or on offline datasets), with reported performance metrics including failure rates.
- Long-term autonomy: Continuous, unsupervised operation over a timescale of days to months in environments that undergo structural or appearance change, without human-initiated map resets or relocalization interventions.
2. Methodology of the Review
2.1. Search Strategy
(“SLAM” OR “simultaneous localization and mapping” OR “mobile robot localization”) AND (“LiDAR” OR “visual odometry” OR “IMU fusion” OR “neural SLAM” OR “NeRF” OR “Gaussian splatting” OR “deep learning” OR “sensor fusion” OR “place recognition” OR “radar odometry”)
2.2. Inclusion Criteria
- Reported at least one quantitative localization or mapping metric (ATE, RPE, RMSE, Recall@N, or equivalent);
- Evaluated on a platform that physically moved through an environment (ground vehicle, UAV, AUV, legged robot, or hand-held rig);
- Were peer-reviewed or, for preprints, met the citation threshold defined in Section 2.1.
2.3. Exclusion Criteria
- Non-peer-reviewed blog posts or tutorials;
- Papers without experimental validation;
- Works focused exclusively on hardware design without a localization or mapping contribution.
2.4. Quality Assessment and Risk of Bias
- Reproducibility (weight ): Source code or evaluation data openly released, or results independently replicated on a public benchmark by a third party.
- Dataset independence (weight ): Primary evaluation performed on at least one public benchmark not used for algorithm design or hyperparameter tuning.
- Ablation completeness (weight ): Key design choices individually ablated, or equivalent sensitivity analysis provided.
- Real-world validation (weight ): System tested on a physical robot platform under conditions beyond a controlled laboratory.
- Uncertainty reporting (weight ): Trajectory error or performance metrics accompanied by variance or confidence intervals across multiple runs or evaluation sequences.
2.5. Taxonomy Construction
- Sensor modality;
- Algorithmic paradigm;
- Deployment scale (single vs multi-robot).
3. Comparison with Existing Surveys
4. Sensor Suites Across Robotic Domains
4.1. Proprioceptive Sensors
4.2. Exteroceptive Sensors
4.2.1. Vision Sensors (Cameras)
4.2.2. LiDAR (Light Detection and Ranging)
4.2.3. Radar
4.2.4. Acoustic and Sonar Sensors
4.2.5. Thermal Cameras
4.3. Global and Infrastructure-Based Sensors
4.4. Sensor Comparison Summary
5. Benchmark Datasets and Evaluation
6. Classical Probabilistic State Estimation
- is the belief—the posterior probability distribution over robot state given all past observations and controls;
- is a scalar normalization constant ensuring the posterior integrates to unity;
- is the observation likelihood (sensor or measurement model), giving the probability of receiving measurement if the true state is ;
- is the motion model (state-transition density), encoding the probability of transitioning from to under control input .
6.1. Gaussian Filters
6.2. Particle Filters (Monte Carlo Localization)
6.3. Graph-Based Optimization (Factor Graphs)
6.4. Key Insights: Classical Estimation
- Classical probabilistic methods provide the principled uncertainty quantification that learning-based systems largely lack, with serious consequences for safety-critical deployment.
- No purely classical method has demonstrated reliable long-term autonomy in environments that change over time. This limitation motivates the multi-sensor and learning-augmented SLAM systems reviewed in the following section.
7. Modern Simultaneous Localization and Mapping (SLAM)
7.1. Visual SLAM (V-SLAM)
7.2. Visual-Inertial Odometry and SLAM (VIO/VI-SLAM)
7.3. LiDAR SLAM and LiDAR-Inertial Odometry (LIO)
7.4. Radar SLAM
7.5. Multi-Modal SLAM
7.6. Key Insights: SLAM Paradigms
- Tightly coupled multi-sensor fusion (LiDAR+camera+IMU) represents the current state of the art in robustness, but introduces calibration complexity and single points of failure at the fusion interface.
- Despite decades of research, no existing SLAM system has demonstrated sustained reliable operation over weeks or months in environments that undergo structural change.
8. End-to-End Neural Mapping and SLAM Pipelines
8.1. NeRF-Based SLAM
8.2. 3D Gaussian Splatting SLAM
8.3. Foundation Models for Geometry and SLAM
8.4. Key Insights and Subcategory Comparison
- NeRF-based SLAM (iMAP [87], NICE-SLAM [88], Co-SLAM [89], ESLAM [90]): Uses a continuous volumetric function (MLP or hash grid) optimized via differentiable volume rendering. Strengths include smooth geometry completion and hallucination of unobserved regions through learned priors. Weaknesses include the following: (1) convergence time of minutes to hours per scene; (2) loop closure requires partial MLP retraining, introducing 2–10 s latency spikes; (3) per-pixel ray-marching is memory-bandwidth-bound and infeasible on embedded hardware. These systems exclusively target small-scale (≤50 m2) indoor reconstruction from RGB-D input.
- 3DGS-based SLAM (SplaTAM [91], MonoGS [92], GS-SLAM [93], Photo-SLAM [94]): Represents scenes as collections of anisotropic 3D Gaussians rendered via differentiable rasterization. Advantages over NeRF include explicit scene representation (editable primitives), faster rendering (>100 FPS for novel-view synthesis on desktop GPU), and easier geometric pruning. Remaining weaknesses: (1) Gaussian count grows unboundedly, leading to memory saturation in large scenes; (2) tracking quality is degraded by rapid viewpoint changes; (3) loop closure with Gaussian pruning is an open problem. Current systems achieve 1–3 FPS end-to-end joint tracking+mapping on an RTX 3090 (350 W) at 640 × 480 RGB-D input streamed at 30 Hz, excluding loop-closure latency, compared to <0.3 FPS for NeRF-SLAM under the same input, hardware, and measurement conditions (see the runtime reporting frame in Section 9).
- Foundation model integration (DUSt3R [95], MASt3R [96]): These systems perform geometry estimation from uncalibrated image pairs but produce no persistent map and operate at 1–2 FPS, making them unsuitable as drop-in SLAM front-ends without further engineering. Their primary value is in providing zero-shot initialization for classical backends.
9. Deep Learning Modules for Relocalization and Feature Extraction
9.1. Absolute Pose Regression (APR)
9.2. Scene Coordinate Regression (SCR)
9.3. Learned Feature Extraction and Matching
9.4. Hierarchical Visual Localization
9.5. Visual Place Recognition
9.6. Semantic Localization
9.7. Classical vs. Learning-Based SLAM
10. Multi-Robot and Collaborative Localization
- Centralized Approaches: A central server collects data from all robots and performs joint optimization. Kimera-Multi [115] extends Kimera to multi-robot metric-semantic SLAM with a centralized server performing distributed place recognition and robust inter-robot loop closure via pairwise consistency maximization.
- Decentralized Approaches: Robots communicate peer-to-peer without a central coordinator. DOOR-SLAM [116] uses pairwise consistency maximization (PCM) to reject outlier inter-robot loop closures in a fully decentralized architecture, ensuring robustness to perceptual aliasing. More broadly, centralized star topologies maximize consistency but create a single point of failure, decentralized peer-to-peer meshes scale more gracefully at the cost of distributed outlier-rejection overhead, and hierarchical cluster-based designs represent an intermediate point between the two.
- Communication-Efficient Methods: Bandwidth constraints in real-world multi-robot systems have motivated compact map representations for sharing. Approaches include exchanging compressed visual descriptors, lightweight 3D descriptors (e.g., Scan Context [117]), or differentially encoded submaps. Swarm-SLAM [118] demonstrates communication-efficient multi-robot SLAM by prioritizing inter-robot loop closures, an effective bandwidth-reduction strategy.
- Relative Localization: In swarm robotics, UWB ranging combined with visual detection enables robots to estimate their relative poses without requiring a shared global map, supporting coordination tasks such as formation flying and cooperative manipulation [45].
- Cybersecurity and Adversarial Robustness: As multi-robot SLAM reaches open deployments, adversarial robustness becomes a first-class concern: compromised agents, spoofed sensors, and denial-of-service attacks can silently corrupt the collective map. While TEASER [119] provides single-robot certifiable guarantees under outlier contamination, analogous distributed guarantees remain absent. Security-by-design principles (Byzantine-fault-tolerant consensus and cryptographic observation authentication) are gaining research attention but are not yet integrated into mainstream collaborative SLAM frameworks.
Key Insights: Multi-Robot Localization
- Multi-robot SLAM provides coverage and resilience inaccessible to single-robot systems, but incorrect inter-robot data association can corrupt the entire shared map.
- Adversarial robustness is an underexplored dimension: safety-critical deployments require Byzantine-fault-tolerant estimation and cryptographically authenticated inter-robot observations, capabilities absent from mainstream systems.
11. Deployment Domains and Real-World Challenges
11.1. Autonomous Driving
11.2. Unmanned Aerial Vehicles (UAVs)
11.3. Autonomous Underwater Vehicles (AUVs)
11.4. Legged and Humanoid Robots
11.5. Consumer and Service Robotics
11.6. Agricultural and Orchard Robotics
11.7. Computational Platforms for On-Robot SLAM
Embedded GPU, ARM, FPGA, and Neuromorphic Platforms
11.8. Cross-Domain Deployment Challenges
12. Limitations of Current Approaches
12.1. Dynamic Objects Corrupt Photometric and Geometric Consistency
12.2. Sensor-Specific Failure Modes Lack Systematic Characterization
12.3. Neural SLAM Is Orders of Magnitude Too Slow for Embedded Deployment
12.4. Generalization of Learning-Based Methods Degrades Precipitously out of Distribution
12.5. Inconsistent Evaluation Metrics Conceal True Performance Variability
12.6. Semantic–Geometric Coupling Remains Superficial
12.7. Safety Certification Is Structurally Unaddressed
13. Open Challenges and Future Directions
13.1. Dynamic-Object-Aware SLAM Without Semantic Priors
13.2. All-Weather Multi-Modal Fusion with Calibrated Uncertainty
13.3. Sub-10-Watt Real-Time Neural Mapping
13.4. Foundation Models as Zero-Shot SLAM Front-Ends
13.5. Reproducible Benchmarking with Mandatory Completeness Metrics
- Segmentation. Partition into maximal continuous tracking segments , where a new segment begins whenever the system re-initializes after a tracking loss (each re-initialization generally establishes a new, unrelated estimator frame).
- Per-segment alignment. For each segment containing at least frames (we recommend , i.e., one second at 30 Hz), compute an SE(3) Umeyama alignment [31] between the estimated and ground-truth poses of that segment only, and evaluate for under that segment’s alignment. Estimated frames are thus never mixed with failed frames in the alignment, and no single global alignment can be biased by post-failure re-initializations. For monocular systems without metric scale, Sim(3) alignment may be substituted, and this must be declared.
- Failed-frame assignment. Every frame , and every frame belonging to a segment shorter than (too short for a well-conditioned alignment), is assigned and therefore scores 0 in the completeness sum below.
13.6. Lifelong Mapping Under Continual Distributional Shift
13.7. Certifiable Pose Bounds for Safety-Critical Systems
13.8. Active SLAM: Closing the Loop Between Estimation and Control
14. Conclusions
Supplementary Materials
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Thrun, S.; Burgard, W.; Fox, D. Probabilistic Robotics; MIT Press: Cambridge, MA, USA, 2005. [Google Scholar]
- Cadena, C.; Carlone, L.; Carrillo, H.; Latif, Y.; Scaramuzza, D.; Neira, J.; Reid, I.; Leonard, J.J. Past, present, and future of simultaneous localization and mapping: Toward the robust-perception age. IEEE Trans. Robot. 2016, 32, 1309–1332. [Google Scholar]
- Borenstein, J.; Everett, H.R.; Feng, L. Where Am I? Sensors and Methods for Mobile Robot Positioning; University of Michigan Technical Report; University of Michigan: Ann Arbor, MI, USA, 1996. [Google Scholar]
- Smith, R.; Self, M.; Cheeseman, P. Estimating uncertain spatial relationships in robotics. In Autonomous Robot Vehicles; Springer: New York, NY, USA, 1990; pp. 167–193. [Google Scholar]
- Leonard, J.J.; Durrant-Whyte, H.F. Simultaneous map building and localization for an autonomous mobile robot. In Proceedings of the IEEE/RSJ International Workshop on Intelligent Robots and Systems (IROS ’91), Osaka, Japan, 3–5 November 1991; pp. 1442–1447. [Google Scholar]
- Dellaert, F.; Fox, D.; Burgard, W.; Thrun, S. Monte Carlo localization for mobile robots. In Proceedings of the Proceedings 1999 IEEE International Conference on Robotics and Automation (Cat. No.99CH36288C), Detroit, MI, USA, 10–15 May 1999; pp. 1322–1328. [Google Scholar]
- Fox, D. Adapting the sample size in particle filters through KLD-sampling. Int. J. Robot. Res. 2003, 22, 985–1003. [Google Scholar] [CrossRef] [Scilit]
- Nistér, D.; Naroditsky, O.; Bergen, J. Visual odometry. In Proceedings of the 2004 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2004, CVPR 2004, Washington, DC, USA, 27 June–2 July 2004; Volume 1, pp. 1–652. [Google Scholar]
- Klein, G.; Murray, D. Parallel tracking and mapping for small AR workspaces. In Proceedings of the 2007 6th IEEE and ACM International Symposium on Mixed and Augmented Reality, Nara, Japan, 13–16 November 2007; pp. 225–234. [Google Scholar]
- Mur-Artal, R.; Montiel, J.M.M.; Tardós, J.D. ORB-SLAM: A versatile and accurate monocular SLAM system. IEEE Trans. Robot. 2015, 31, 1147–1163. [Google Scholar] [CrossRef] [Scilit]
- Shan, T.; Englot, B.; Meyers, D.; Wang, W.; Ratti, C.; Rus, D. LIO-SAM: Tightly-coupled lidar inertial odometry via smoothing and mapping. In Proceedings of the 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Las Vegas, NV, USA, 24 October 2020–24 January 2021; pp. 5135–5142. [Google Scholar]
- Xu, W.; Cai, Y.; He, D.; Lin, J.; Zhang, F. FAST-LIO2: Fast direct LiDAR-inertial odometry. IEEE Trans. Robot. 2022, 38, 2053–2073. [Google Scholar] [CrossRef] [Scilit]
- DeTone, D.; Malisiewicz, T.; Rabinovich, A. SuperPoint: Self-supervised interest point detection and description. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Salt Lake City, UT, USA, 18–22 June 2018; pp. 224–236. [Google Scholar]
- Sarlin, P.E.; DeTone, D.; Malisiewicz, T.; Rabinovich, A. SuperGlue: Learning feature matching with graph neural networks. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 4938–4947. [Google Scholar]
- Kendall, A.; Grimes, M.; Cipolla, R. PoseNet: A convolutional network for real-time 6-DOF camera relocalization. In Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile, 7–13 December 2015; pp. 2938–2946. [Google Scholar]
- Mildenhall, B.; Srinivasan, P.P.; Tancik, M.; Barron, J.T.; Ramamoorthi, R.; Ng, R. NeRF: Representing scenes as neural radiance fields for view synthesis. In Computer Vision—ECCV 2020 Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2020; Volumn 12346. [Google Scholar]
- Kerbl, B.; Kopanas, G.; Leimkühler, T.; Drettakis, G. 3D Gaussian splatting for real-time radiance field rendering. ACM Trans. Graph. 2023, 42, 139. [Google Scholar] [CrossRef] [Scilit]
- Tosi, F.; Zhang, Y.; Gong, Z.; Sandström, E.; Mattoccia, S.; Oswald, M.R.; Poggi, M. How NeRFs and 3D Gaussian splatting are reshaping SLAM: A survey. IEEE Trans. Robot. 2026, 42, 1405–1427. [Google Scholar] [CrossRef] [Scilit]
- Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ 2021, 372, n71. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Whiting, P.; Rutjes, A.W.; Westwood, M.E.; Mallett, S.; Deeks, J.J.; Reitsma, J.B.; Leeflang, M.M.; Sterne, J.A.; Bossuyt, P.M. QUADAS-2: A revised tool for the quality assessment of diagnostic accuracy studies. Ann. Intern. Med. 2011, 155, 529–536. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Amrhein, V.; Greenland, S.; McShane, B. Scientists rise up against statistical significance. Nature 2019, 567, 305–307. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Landis, J.R.; Koch, G.G. The measurement of observer agreement for categorical data. Biometrics 1977, 33, 159–174. [Google Scholar] [CrossRef] [Scilit]
- Bresson, G.; Alsayed, Z.; Yu, L.; Glaser, S. Simultaneous localization and mapping: A survey of current trends in autonomous driving. IEEE Trans. Intell. Veh. 2017, 2, 194–220. [Google Scholar] [CrossRef] [Scilit]
- Fuentes-Pacheco, J.; Ruiz-Ascencio, J.; Rendón-Mancha, J.M. Visual simultaneous localization and mapping: A survey. Artif. Intell. Rev. 2015, 43, 55–81. [Google Scholar]
- Chen, C.; Wang, B.; Lu, C.X.; Trigoni, N.; Markham, A. A survey on deep learning for localization and mapping: Towards the age of spatial machine intelligence. arXiv 2020, arXiv:2006.12567. [Google Scholar]
- Bloesch, M.; Hutter, M.; Hoepflinger, M.A.; Leutenegger, S.; Gehring, C.; Remy, C.D.; Siegwart, R. State estimation for legged robots—Consistent fusion of leg kinematics and IMU. Robot. Sci. Syst. Conf. 2013, 17, 17–24. [Google Scholar] [CrossRef] [Scilit]
- Woodman, O.J. An Introduction to Inertial Navigation; Technical Report UCAM-CL-TR-696; University of Cambridge, Computer Laboratory: Cambridge, UK, 2007. [Google Scholar]
- Forster, C.; Carlone, L.; Dellaert, F.; Scaramuzza, D. On-manifold preintegration for real-time visual-inertial odometry. IEEE Trans. Robot. 2017, 33, 1–21. [Google Scholar] [CrossRef] [Scilit]
- Engel, J.; Koltun, V.; Cremers, D. Direct sparse odometry. IEEE Trans. Pattern Anal. Mach. Intell. 2018, 40, 611–625. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Mur-Artal, R.; Tardós, J.D. ORB-SLAM2: An open-source SLAM system for monocular, stereo, and RGB-D cameras. IEEE Trans. Robot. 2017, 33, 1255–1262. [Google Scholar] [CrossRef] [Scilit]
- Sturm, J.; Engelhard, N.; Endres, F.; Burgard, W.; Cremers, D. A benchmark for the evaluation of RGB-D SLAM systems. In Proceedings of the 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, Vilamoura-Algarve, Portugal, 7–12 October 2012; pp. 573–580. [Google Scholar]
- Campos, C.; Elvira, R.; Rodríguez, J.J.G.; Montiel, J.M.M.; Tardós, J.D. ORB-SLAM3: An accurate open-source library for visual, visual-inertial, and multimap SLAM. IEEE Trans. Robot. 2021, 37, 1874–1890. [Google Scholar] [CrossRef] [Scilit]
- Gallego, G.; Delbrück, T.; Orchard, G.; Bartolozzi, C.; Taba, B.; Censi, A.; Leutenegger, S.; Davison, A.J.; Conradt, J.; Daniilidis, K.; et al. Event-based vision: A survey. IEEE Trans. Pattern Anal. Mach. Intell. 2020, 44, 154–180. [Google Scholar]
- Rebecq, H.; Horstschäfer, T.; Scaramuzza, D. Real-time visual-inertial odometry for event cameras using keyframe-based nonlinear optimization. In Proceedings of the British Machine Vision Conference (BMVC), London, UK, 4–7 September 2017. [Google Scholar]
- Vidal, A.R.; Rebecq, H.; Horstschäfer, T.; Scaramuzza, D. Ultimate SLAM? Combining events, images, and IMU for robust visual SLAM in HDR and high-speed scenarios. IEEE Robot. Autom. Lett. 2018, 3, 994–1001. [Google Scholar] [CrossRef] [Scilit]
- Zhang, J.; Singh, S. LOAM: Lidar odometry and mapping in real-time. In Proceedings of the Robotics: Science and Systems (RSS), Berkeley, CA, USA, 12–16 July 2014. [Google Scholar]
- Harlow, K.; Jang, H.; Barfoot, T.D.; Kim, A.; Heckman, C. A new wave in robotics: Survey on recent mmWave radar applications in robotics. IEEE Trans. Robot. 2024, 40, 4544–4560. [Google Scholar] [CrossRef] [Scilit]
- Zhuang, J.; Wang, C.; Wang, L. 4DRadarSLAM: A 4D imaging radar SLAM system for large-scale environments. In Proceedings of the 2023 IEEE International Conference on Robotics and Automation (ICRA), London, UK, 29 May–2 June 2023; pp. 8221–8227. [Google Scholar]
- Gadd, M.; De Martini, D.; Newman, P. Contrastive learning for robust radar place recognition. IEEE Robot. Autom. Lett. 2024, 9, 1447–1454. [Google Scholar]
- Paull, L.; Saeedi, S.; Seto, M.; Li, H. AUV navigation and localization: A review. IEEE J. Ocean. Eng. 2014, 39, 131–149. [Google Scholar] [CrossRef] [Scilit]
- Kinsey, J.C.; Eustice, R.M.; Whitcomb, L.L. A survey of underwater vehicle navigation: Recent advances and new challenges. In Proceedings of the IFAC Conference on Manoeuvring and Control of Marine Craft, Lisbon, Portugal, 20–22 September 2006; Volume 39, pp. 1–12. [Google Scholar]
- Hover, F.S.; Eustice, R.M.; Kim, A.; Englot, B.; Johannsson, H.; Kaess, M.; Leonard, J.J. Advanced perception, navigation and planning for autonomous in-water ship hull inspection. Int. J. Robot. Res. 2012, 31, 1445–1464. [Google Scholar] [CrossRef] [Scilit]
- Shin, Y.S.; Kim, A. Sparse depth enhanced direct thermal-infrared SLAM beyond the visible spectrum. RA-L 2019, 4, 2918–2925. [Google Scholar] [CrossRef] [Scilit]
- Groves, P.D. Principles of GNSS, Inertial, and Multisensor Integrated Navigation Systems, 2nd ed.; Artech House: Norwood, MA, USA, 2013. [Google Scholar]
- Nguyen, T.H.; Nguyen, T.M.; Xie, L. Tightly-coupled ultra-wideband-aided monocular visual SLAM with degenerate anchor configurations. Auton. Robot. 2020, 44, 1519–1534. [Google Scholar] [CrossRef] [Scilit]
- Yassin, A.; Nasser, Y.; Awad, M.; Al-Dubai, A.; Liu, R.; Yuen, C.; Raulefs, R.; Aboutanios, E. Recent advances in indoor localization: A survey on theoretical approaches and applications. IEEE Commun. Surv. Tutor. 2017, 19, 1327–1346. [Google Scholar] [CrossRef] [Scilit]
- Xia, H.; Wang, Z.; Jiang, Z.; Zhang, Q. Indoor localization via magnetic fingerprinting using LSTM networks. IEEE Sens. J. 2022, 22, 9176–9185. [Google Scholar]
- Geiger, A.; Lenz, P.; Urtasun, R. Are we ready for autonomous driving? The KITTI vision benchmark suite. In Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA, 16–21 June 2012; pp. 3354–3361. [Google Scholar]
- Burri, M.; Nikolic, J.; Gohl, P.; Schneider, T.; Rehder, J.; Omari, S.; Achtelik, M.W.; Siegwart, R. The EuRoC micro aerial vehicle datasets. Int. J. Robot. Res. 2016, 35, 1157–1163. [Google Scholar] [CrossRef] [Scilit]
- Schubert, D.; Goll, T.; Demmel, M.; Usenko, V.; Stückler, J.; Cremers, D. The TUM VI benchmark for evaluating visual-inertial odometry. In Proceedings of the 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Madrid, Spain, 1–5 October 2018; pp. 1680–1687. [Google Scholar]
- Caesar, H.; Bankiti, V.; Lang, A.H.; Vora, S.; Liong, V.E.; Xu, Q.; Krishnan, A.; Pan, Y.; Baldan, G.; Beijbom, O. nuScenes: A multimodal dataset for autonomous driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 11618–11628. [Google Scholar]
- Helmberger, M.; Morin, K.; Berner, B.; Kumar, N.; Cioffi, G.; Scaramuzza, D. The Hilti SLAM challenge: Mapping an indoor construction site. IEEE Robot. Autom. Lett. 2022, 7, 7518–7525. [Google Scholar] [CrossRef] [Scilit]
- Yin, J.; Li, A.; Li, T.; Yu, W.; Zou, D. M2DGR: A multi-sensor and multi-scenario SLAM dataset for ground robots. IEEE Robot. Autom. Lett. 2022, 7, 2266–2273. [Google Scholar] [CrossRef] [Scilit]
- Maddern, W.; Pascoe, G.; Linegar, C.; Newman, P. 1 year, 1000 km: The Oxford RobotCar dataset. Int. J. Robot. Res. 2017, 36, 3–15. [Google Scholar] [CrossRef] [Scilit]
- Ramezani, M.; Wang, Y.; Camurri, M.; Wisth, D.; Mattamala, M.; Fallon, M. The Newer College dataset: Handheld LiDAR, inertial and vision with ground truth. In Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Las Vegas, NV, USA, 25–29 October 2020; pp. 4353–4360. [Google Scholar]
- Zhao, S.; Gao, Y.; Wu, T.; Singh, D.; Jiang, R.; Sun, H.; Sarawata, M.; Qiu, Y.; Whittaker, W.; Higgins, I.; et al. SubT-MRS: Pushing SLAM towards all-weather environments. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024; pp. 22647–22657. [Google Scholar]
- Grupp, M. evo: Python Package for the Evaluation of Odometry and SLAM. 2017. Available online: https://github.com/MichaelGrupp/evo (accessed on 23 May 2026).
- Julier, S.J.; Uhlmann, J.K. Unscented filtering and nonlinear estimation. Proc. IEEE 2004, 92, 401–422. [Google Scholar] [CrossRef] [Scilit]
- Solà, J. Quaternion kinematics for the error-state Kalman filter. arXiv 2017, arXiv:1711.02508. [Google Scholar]
- Montemerlo, M.; Thrun, S.; Koller, D.; Wegbreit, B. FastSLAM: A factored solution to the simultaneous localization and mapping problem. In Proceedings of the Eighteenth national conference on Artificial intelligence, Edmonton, AB, Canada, 28 July–1 August 2002; pp. 593–598. [Google Scholar]
- Grisetti, G.; Stachniss, C.; Burgard, W. Improved techniques for grid mapping with Rao-Blackwellized particle filters. IEEE Trans. Robot. 2007, 23, 34–46. [Google Scholar] [CrossRef] [Scilit]
- Grisetti, G.; Kümmerle, R.; Stachniss, C.; Burgard, W. A tutorial on graph-based SLAM. IEEE Intell. Transp. Syst. Mag. 2010, 2, 31–43. [Google Scholar] [CrossRef] [Scilit]
- Kaess, M.; Johannsson, H.; Roberts, R.; Ila, V.; Leonard, J.J.; Dellaert, F. iSAM2: Incremental smoothing and mapping using the Bayes tree. Int. J. Robot. Res. 2012, 31, 216–235. [Google Scholar] [CrossRef] [Scilit]
- Dellaert, F. Factor Graphs and GTSAM: A Hands-on Introduction; Technical Report GT-RIM-CP&R-2012-002; Georgia Institute of Technology: Atlanta, Georgia, 2012. [Google Scholar]
- Agarwal, S.; Mierle, K. Ceres Solver (Software). 2012. Available online: http://ceres-solver.org (accessed on 23 May 2026).
- Kümmerle, R.; Grisetti, G.; Strasdat, H.; Konolige, K.; Burgard, W. G2o: A general framework for graph optimization. In Proceedings of the 2011 IEEE International Conference on Robotics and Automation, Shanghai, China, 9–13 May 2011; pp. 3607–3613. [Google Scholar]
- Davison, A.J.; Reid, I.D.; Molton, N.D.; Stasse, O. MonoSLAM: Real-time single camera SLAM. IEEE Trans. Pattern Anal. Mach. Intell. 2007, 29, 1052–1067. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Gálvez-López, D.; Tardós, J.D. Bags of binary words for fast place recognition in image sequences. IEEE Trans. Robot. 2012, 28, 1188–1197. [Google Scholar] [CrossRef] [Scilit]
- Engel, J.; Schöps, T.; Cremers, D. LSD-SLAM: Large-scale direct monocular SLAM. In Computer Vision–ECCV 2014; Springer: Cham, Switzerland, 2014; pp. 834–849. [Google Scholar]
- Wang, R.; Schwörer, M.; Cremers, D. Stereo DSO: Large-scale direct sparse visual odometry with stereo cameras. In Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017; pp. 3903–3911. [Google Scholar]
- Forster, C.; Zhang, Z.; Gassner, M.; Werlberger, M.; Scaramuzza, D. SVO: Semidirect visual odometry for monocular and multicamera systems. IEEE Trans. Robot. 2017, 33, 249–265. [Google Scholar] [CrossRef] [Scilit]
- Mourikis, A.I.; Roumeliotis, S.I. A multi-state constraint Kalman filter for vision-aided inertial navigation. In Proceedings of the Proceedings 2007 IEEE International Conference on Robotics and Automation, Rome, Italy, 10–14 April 2007; pp. 3565–3572. [Google Scholar]
- Geneva, P.; Eckenhoff, K.; Lee, W.; Yang, Y.; Huang, G. OpenVINS: A research platform for visual-inertial state estimation. In Proceedings of the 2020 IEEE International Conference on Robotics and Automation (ICRA), Paris, France, 31 May–31 August 2020; pp. 4666–4672. [Google Scholar]
- Leutenegger, S.; Lynen, S.; Bosse, M.; Siegwart, R.; Furgale, P. Keyframe-based visual-inertial odometry using nonlinear optimization. Int. J. Robot. Res. 2015, 34, 314–334. [Google Scholar] [CrossRef] [Scilit]
- Qin, T.; Li, P.; Shen, S. VINS-Mono: A robust and versatile monocular visual-inertial state estimator. IEEE Trans. Robot. 2018, 34, 1004–1020. [Google Scholar] [CrossRef] [Scilit]
- Qin, T.; Cao, S.; Shen, S. A general optimization-based framework for global pose estimation with multiple sensors. IET Cyber Syst. Robot. 2025, 7, e70023. [Google Scholar] [CrossRef] [Scilit]
- Rosinol, A.; Abate, M.; Chang, Y.; Carlone, L. Kimera: An open-source library for real-time metric-semantic localization and mapping. In Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), Paris, France, 31 May–31 August 2020; pp. 1689–1696. [Google Scholar]
- Usenko, V.; Demmel, N.; Schubert, D.; Stückler, J.; Cremers, D. Visual-inertial mapping with non-linear factor recovery. IEEE Robot. Autom. Lett. 2020, 5, 422–429. [Google Scholar] [CrossRef] [Scilit]
- Hess, W.; Kohler, D.; Rapp, H.; Andor, D. Real-time loop closure in 2D LiDAR SLAM. In Proceedings of the 2016 IEEE International Conference on Robotics and Automation (ICRA), Stockholm, Sweden, 16–21 May 2016; pp. 1271–1278. [Google Scholar]
- Shan, T.; Englot, B. LeGO-LOAM: Lightweight and ground-optimized lidar odometry and mapping on variable terrain. In Proceedings of the 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Madrid, Spain, 1–5 October 2018; pp. 4758–4765. [Google Scholar]
- Xu, W.; Zhang, F. FAST-LIO: A fast, robust LiDAR-inertial odometry package by tightly-coupled iterated Kalman filter. IEEE Robot. Autom. Lett. 2021, 6, 3317–3324. [Google Scholar] [CrossRef] [Scilit]
- He, D.; Xu, W.; Chen, N.; Kong, F.; Yuan, C.; Zhang, F. Point-LIO: Robust high-bandwidth LiDAR-inertial odometry. Adv. Intell. Syst. 2023, 5, 2200459. [Google Scholar] [CrossRef] [Scilit]
- Vizzo, I.; Guadagnino, T.; Mersch, B.; Wiesmann, L.; Behley, J.; Stachniss, C. KISS-ICP: In defense of point-to-point ICP—Simple, accurate, and robust registration with no learning. IEEE Robot. Autom. Lett. 2023, 8, 1029–1036. [Google Scholar] [CrossRef] [Scilit]
- Shan, T.; Englot, B.; Ratti, C.; Rus, D. LVI-SAM: Tightly-coupled lidar-visual-inertial odometry via smoothing and mapping. In Proceedings of the 2021 IEEE International Conference on Robotics and Automation (ICRA), Xi’an, China, 30 May–5 June 2021; pp. 5692–5698. [Google Scholar]
- Lin, J.; Zhang, F. R3LIVE: A robust, real-time, RGB-colored, LiDAR-inertial-visual tightly-coupled state estimation and mapping package. In Proceedings of the 2022 International Conference on Robotics and Automation (ICRA), Philadelphia, PA, USA, 23–27 May 2022; pp. 10672–10678. [Google Scholar]
- Zheng, C.; Zhu, Q.; Xu, W.; Liu, X.; Li, Q.; Zhang, F. FAST-LIVO: Fast and tightly-coupled sparse-direct LiDAR-inertial-visual odometry. In Proceedings of the 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Kyoto, Japan, 23–27 October 2022; pp. 4167–4173. [Google Scholar]
- Sucar, E.; Liu, S.; Ortiz, J.; Davison, A.J. iMAP: Implicit mapping and positioning in real-time. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 11–17 October 2021; pp. 6209–6218. [Google Scholar]
- Zhu, Z.; Peng, S.; Larsson, V.; Xu, W.; Bao, H.; Cui, Z.; Oswald, M.R.; Pollefeys, M. NICE-SLAM: Neural implicit scalable encoding for SLAM. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 12776–12786. [Google Scholar]
- Wang, H.; Wang, J.; Agapito, L. Co-SLAM: Joint coordinate and sparse parametric encodings for neural real-time SLAM. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 17–24 June 2023; pp. 13293–13302. [Google Scholar]
- Johari, M.M.; Carta, C.; Fleuret, F. ESLAM: Efficient dense SLAM system based on hybrid representation of signed distance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 17–24 June 2023; pp. 17408–17419. [Google Scholar]
- Keetha, N.; Karhade, J.; Jatavallabhula, K.M.; Yang, G.; Scherer, S.; Ramanan, D.; Luiten, J. SplaTAM: Splat, track & map 3D Gaussians for dense RGB-D SLAM. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024; pp. 21357–21366. [Google Scholar]
- Matsuki, H.; Murai, R.; Kelly, P.H.J.; Davison, A.J. Gaussian splatting SLAM. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024; pp. 18039–18048. [Google Scholar]
- Yan, C.; Qu, D.; Xu, D.; Zhao, B.; Wang, Z.; Wang, D.; Liang, X. GS-SLAM: Dense visual SLAM with 3D Gaussian splatting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024; pp. 19595–19604. [Google Scholar]
- Huang, H.; Li, L.; Cheng, H.; Yeung, S.K. Photo-SLAM: Real-time simultaneous localization and photorealistic mapping for monocular, stereo, and RGB-D cameras. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024; pp. 21584–21593. [Google Scholar]
- Wang, S.; Leroy, V.; Cabon, Y.; Chidlovskii, B.; Revaud, J. DUSt3R: Geometric 3D vision made easy. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024; pp. 20697–20709. [Google Scholar]
- Leroy, V.; Cabon, Y.; Revaud, J. Grounding image matching in 3D with MASt3R. In Computer Vision—ECCV 2024; Leonardis, A., Ricci, E., Roth, S., Russakovsky, O., Sattler, T., Varol, G., Eds.; Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2024; Volume 15130. [Google Scholar]
- Sattler, T.; Zhou, Q.; Pollefeys, M.; Leal-Taixé, L. Understanding the limitations of CNN-based absolute camera pose regression. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019; pp. 3297–3307. [Google Scholar]
- Brachmann, E.; Humenberger, M.; Rother, C.; Sattler, T. Accelerated coordinate encoding: Learning to relocalize in minutes using RGB and poses. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 17–24 June 2023; pp. 5044–5053. [Google Scholar]
- Brahmbhatt, S.; Gu, J.; Kim, K.; Hays, J.; Kautz, J. Geometry-aware learning of maps for camera localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 2616–2625. [Google Scholar]
- Wang, B.; Chen, C.; Lu, C.X.; Zhao, P.; Trigoni, N.; Markham, A. AtLoc: Attention guided camera localization. In Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA, 7–12 February 2020; pp. 10393–10401. [Google Scholar]
- Brachmann, E.; Rother, C. Learning less is more—6D camera relocalization via 3D surface regression. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–23 June 2018; pp. 4654–4662. [Google Scholar]
- Revaud, J.; De Souza, C.; Humenberger, M.; Weinzaepfel, P. R2D2: Reliable and repeatable detector and descriptor. In Proceedings of the 33rd International Conference on Neural Information Processing Systems; ACM: New York, NY, USA, 2019; pp. 12414–12424. [Google Scholar]
- Tyszkiewicz, M.; Fua, P.; Trulls, E. DISK: Learning local features with policy gradient. In NIPS’20: Proceedings of the 34th International Conference on Neural Information Processing Systems, Vancouver, BC, Canada, 6–12 December 2020; ACM: New York, NY, USA, 2020; pp. 14254–14265. [Google Scholar]
- Lindenberger, P.; Sarlin, P.E.; Pollefeys, M. LightGlue: Local feature matching at light speed. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 1–6 October 2023; pp. 17581–17592. [Google Scholar]
- Sun, J.; Shen, Z.; Wang, G.; Bai, X.; Fang, H.; Fu, Q. LoFTR: Detector-free local feature matching with transformers. In Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021; pp. 8922–8931. [Google Scholar]
- Sarlin, P.E.; Cadena, C.; Siegwart, R.; Dymczyk, M. From coarse to fine: Robust hierarchical localization at large scale. In Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019; pp. 12716–12725. [Google Scholar]
- Arandjelović, R.; Gronat, P.; Torii, A.; Pajdla, T.; Sivic, J. NetVLAD: CNN architecture for weakly supervised place recognition. IEEE Trans. Pattern Anal. Mach. Intell. 2018, 40, 1437–1451. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hausler, S.; Garg, S.; Milford, M. Patch-NetVLAD: Multi-scale fusion of locally-global descriptors for place recognition. In Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021; pp. 14141–14152. [Google Scholar]
- Ali-bey, A.; Chaib-draa, B.; Giguère, P. MixVPR: Feature mixing for visual place recognition. In Proceedings of the 2023 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA, 2–7 January 2023; pp. 2998–3007. [Google Scholar]
- Keetha, N.; Mishra, A.; Karhade, J.; Jatavallabhula, K.M.; Scherer, S.; Krishna, M.; Garg, S. AnyLoc: Towards universal visual place recognition. IEEE Robot. Autom. Lett. 2023, 9, 1286–1293. [Google Scholar] [CrossRef] [Scilit]
- Cheng, B.; Misra, I.; Schwing, A.G.; Kirillov, A.; Girdhar, R. Masked-attention mask transformer for universal image segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 1280–1289. [Google Scholar]
- Ravi, N.; Gabeur, V.; Hu, Y.-T.; Hu, R.; Ryali, C.; Ma, T.; Khedr, H.; Rädle, R.; Rolland, C.; Gustafson, L.; et al. SAM 2: Segment Anything in Images and Videos. arXiv 2024, arXiv:2408.00714. [Google Scholar]
- Bescos, B.; Fácil, J.M.; Civera, J.; Neira, J. DynaSLAM: Tracking, mapping, and inpainting in dynamic scenes. IEEE Robot. Autom. Lett. 2018, 3, 4076–4083. [Google Scholar] [CrossRef] [Scilit]
- Nicholson, L.; Milford, M.; Sünderhauf, N. QuadricSLAM: Dual quadrics from object detections as landmarks in object-oriented SLAM. IEEE Robot. Autom. Lett. 2019, 4, 1–8. [Google Scholar] [CrossRef] [Scilit]
- Tian, Y.; Chang, Y.; Arias, F.H.; Nieto-Granda, C.; How, J.P.; Carlone, L. Kimera-Multi: Robust, distributed, dense metric-semantic SLAM for multi-robot systems. IEEE Trans. Robot. 2022, 38, 2022–2038. [Google Scholar] [CrossRef] [Scilit]
- Lajoie, P.Y.; Hu, S.; Beltrame, G. DOOR-SLAM: Distributed, online, and outlier resilient SLAM for robotic teams. IEEE Robot. Autom. Lett. 2020, 5, 1656–1663. [Google Scholar] [CrossRef] [Scilit]
- Kim, G.; Kim, A. Scan Context: Egocentric spatial descriptor for place recognition within 3D point cloud map. In Proceedings of the 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Madrid, Spain, 1–5 October 2018; pp. 4802–4809. [Google Scholar]
- Lajoie, P.Y.; Beltrame, G. Swarm-SLAM: Sparse decentralized collaborative simultaneous localization and mapping framework for multi-robot systems. IEEE Robot. Autom. Lett. 2024, 9, 475–482. [Google Scholar] [CrossRef] [Scilit]
- Yang, H.; Shi, J.; Carlone, L. TEASER: Fast and certifiable point cloud registration. IEEE Trans. Robot. 2020, 37, 314–333. [Google Scholar] [CrossRef] [Scilit]
- Biber, P.; Straßer, W. The normal distributions transform: A new approach to laser scan matching. In Proceedings of the Proceedings 2003 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2003) (Cat. No.03CH37453), Las Vegas, NV, USA, 27–31 October 2003; Volume 3, pp. 2743–2748. [Google Scholar]
- Biber, P.; Duckett, T. Dynamic maps for long-term operation of mobile service robots. In Proceedings of the Robotics: Science and Systems (RSS), Cambridge, MA, USA, 8–11 June 2005; pp. 17–24. [Google Scholar]
- Chen, M.; Tang, Y.; Zou, X.; Huang, Z.; Zhou, H.; Chen, S. 3D global mapping of large-scale unstructured orchard integrating eye-in-hand stereo vision and SLAM. Comput. Electron. Agric. 2021, 187, 106237. [Google Scholar] [CrossRef] [Scilit]
- Chen, M.; Chen, Z.; Luo, L.; Tang, Y.; Cheng, J.; Wei, H.; Wang, J. Dynamic visual servo control methods for continuous operation of a fruit harvesting robot working throughout an orchard. Comput. Electron. Agric. 2024, 219, 108774. [Google Scholar] [CrossRef] [Scilit]
- Davies, M.; Wild, A.; Orchard, G.; Sandamirskaya, Y.; Guerra, G.A.F.; Joshi, P.; Plank, P.; Risbud, S.R. Advancing neuromorphic computing with Loihi: A survey of results and outlook. Proc. IEEE 2021, 109, 911–934. [Google Scholar] [CrossRef] [Scilit]
- Xu, Y.; Jiang, H.; Xiao, Z.; Feng, J.; Zhang, L. DG-SLAM: Robust dynamic Gaussian splatting SLAM with hybrid pose optimization. Adv. Neural Inf. Process. Syst. 2024, 37, 51577–51596. [Google Scholar] [CrossRef] [Scilit]
- Sattler, T.; Maddern, W.; Toft, C.; Torii, A.; Hammarstrand, L.; Stenborg, E.; Safari, D.; Okutomi, M.; Pollefeys, M.; Sivic, J.; et al. Benchmarking 6DOF outdoor visual localization in changing conditions. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 8601–8610. [Google Scholar]
- McCormac, J.; Handa, A.; Davison, A.; Leutenegger, S. SemanticFusion: Dense 3D semantic mapping with convolutional neural networks. In Proceedings of the 2017 IEEE International Conference on Robotics and Automation (ICRA), Singapore, 29 May–3 June 2017; pp. 4628–4635. [Google Scholar]
- Adolfsson, D.; Magnusson, M.; Alhashimi, A.; Lilienthal, A.J.; Andreasson, H. CFEAR radarodometry: Conservative filtering for efficient and accurate radar odometry. In Proceedings of the 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Prague, Czech Republic, 27 September–1 October 2021; pp. 5462–5469. [Google Scholar]
- Campi, M.C.; Garatti, S. The exact feasibility of randomized solutions of uncertain convex programs. SIAM J. Optim. 2008, 19, 1211–1230. [Google Scholar] [CrossRef] [Scilit]
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning transferable visual models from natural language supervision. In Proceedings of the Thirty-Eighth International Conference on Machine Learning (ICML), Virtually, 18–24 July 2021. [Google Scholar]


| Dimension | Visual SLAM | LiDAR SLAM | NeRF-SLAM | 3DGS-SLAM | Multi-robot |
|---|---|---|---|---|---|
| (Moderate/High) | () | () | () | () | () |
| Reproducibility | 14/41, 34% [22–49] | 7/32, 22% [11–39] | 8/15, 53% [30–75] | 5/10, 50% [24–76] | 7/17, 41% [22–64] |
| Real-world validation | 16/41, 39% [26–54] | 7/32, 22% [11–39] | 14/15, 93% [70–99] | 8/10, 80% [49–94] | 10/17, 59% [36–78] |
| Uncertainty reporting | 24/41, 59% [43–72] | 17/32, 53% [36–69] | 14/15, 93% [70–99] | 9/10, 90% [60–98] | 12/17, 71% [47–87] |
| Weighting Scheme | LiDAR | Visual | Multi-Robot | 3DGS | NeRF |
|---|---|---|---|---|---|
| Baseline () | 31.6 | 39.7 | 59.7 | 66.0 | 72.6 |
| Uniform () | 33.1 | 41.0 | 61.2 | 68.0 | 74.6 |
| Reproducibility-dominant () | 30.3 | 39.2 | 56.2 | 63.5 | 69.2 |
| Validation-dominant () | 30.3 | 40.5 | 60.6 | 71.0 | 79.2 |
| Uncertainty-dominant () | 38.1 | 45.3 | 63.5 | 73.5 | 79.2 |
| Survey | Year | Sensors | Classical SLAM | Deep Learning | Multi-Robot | Cross-Domain |
|---|---|---|---|---|---|---|
| Fuentes-Pacheco et al. [24] | 2015 | Vision | Yes | No | No | No |
| Cadena et al. [2] | 2016 | Yes | Yes | No | Yes | No |
| Bresson et al. [23] | 2017 | Yes | Yes | No | No | No |
| Chen et al. [25] | 2020 | Yes | Yes | Yes | No | No |
| Tosi et al. [18] | 2024 | Vision | No | Yes | No | No |
| This survey | 2026 | Yes | Yes | Yes | Yes | Yes |
| Sensor | Range | Precision | Rate | Weight | Env. | Cost (USD) |
|---|---|---|---|---|---|---|
| Wheel Encoder [3] | – | mm | >100 Hz | Low | Ground | $10–200 |
| IMU (MEMS) [27] | – | Variable | 200–1000 Hz | Low | All | $20–200 |
| Monocular Cam. [10,29] | ∞ | px-level | 30–120 Hz | Very Low | Lit | $50–500 |
| Stereo Camera [30] | 0.5–20 m | cm | 30–90 Hz | Low | Lit | $200–800 |
| RGB-D Camera [31] | 0.3–10 m | mm–cm | 30 Hz | Low | Indoor | $100–500 |
| Event Camera [33] | ∞ | px-level | µs | Very Low | All | $3k–15k |
| 2D LiDAR [1] | 0.1–30 m | cm | 10–40 Hz | Low | All | $100–2k |
| 3D LiDAR [36] | 0.3–200 m | cm | 10–20 Hz | Medium | All | $5k–80k |
| Solid-State LiDAR [12] | 0.3–450 m | cm | 10 Hz | Low | All | $1k–10k |
| 4D Imaging Radar [37,38] | 0.2–300 m | dm | 10–20 Hz | Low | All | $2k–20k |
| DVL [40] | 0.5–200 m | mm/s | 1–10 Hz | Medium | Underwater | $5k–30k |
| USBL [41] | 100–10,000 m | 0.1–1% R | 0.1–1 Hz | Medium | Underwater | $5k–60k |
| Thermal Camera [43] | ∞ | px-level | 30–60 Hz | Low | All | $1k–10k |
| GNSS (RTK) [44] | Global | cm | 1–20 Hz | Low | Outdoor | $500–5k |
| UWB [45] | 0–100 m | cm–dm | 10–100 Hz | Very Low | Indoor | $30–200 |
| Dataset | Sensors | Environment | Ground Truth (Accuracy) | Metric/Alignment | Ref. |
|---|---|---|---|---|---|
| KITTI Odometry | Stereo, LiDAR, GPS/IMU | Outdoor driving | RTK-GNSS/INS (∼10 cm) | RTE/RRE, per-segment, no align. | [48] |
| EuRoC MAV | Stereo, IMU | Indoor MAV | Laser tracker/Vicon (mm) | ATE, SE(3) Umeyama | [49] |
| TUM RGB-D | RGB-D | Indoor handheld | Motion capture (mm) | ATE/RPE, SE(3)/Sim(3) | [31] |
| TUM VI | Stereo, IMU | Indoor/outdoor | Mocap (partial coverage) | ATE on mocap segments | [50] |
| nuScenes | Camera, LiDAR, Radar | Urban driving | Map-based loc. + GNSS/INS (dm) | Task-specific (detection-centric) | [51] |
| Hilti Challenge | LiDAR, Camera, IMU | Construction sites | Total station/TLS prisms (mm–cm) | ATE at control points, SE(3) | [52] |
| M2DGR | Multi-modal | Ground robot | RTK-GNSS/mocap/tracker (cm) | ATE, SE(3) Umeyama | [53] |
| Oxford RobotCar | Camera, LiDAR, GPS | Urban long-term | GPS/INS (m-level, drifting) | RTE/place-recognition recall | [54] |
| Newer College | LiDAR, Camera, IMU | Outdoor handheld | ICP vs. TLS prior map (cm) | ATE, SE(3) Umeyama | [55] |
| SubT-MRS | Multi-modal | Subterranean | Total station + FARO scans (cm) | ATE, SE(3); failures logged | [56] |
| System | Sensor | Real-Time | Loop Closure | GPU | Open-Source | Environment |
|---|---|---|---|---|---|---|
| ORB-SLAM3 | Mono/Stereo/VIO | Yes | Yes | No | Yes | Indoor/Outdoor |
| Cartographer | LiDAR | Yes | Yes | No | Yes | Indoor/Outdoor |
| LOAM | LiDAR | Yes | Limited | No | Yes | Outdoor |
| VINS-Fusion | VIO | Yes | Yes | No | Yes | Indoor/Outdoor |
| LIO-SAM | LiDAR+IMU | Yes | Yes | No | Yes | Outdoor |
| Kimera | VIO | Yes | Yes | Optional | Yes | Indoor |
| Approach | Task | Mapping | Online | Sensor | Accuracy | Compute | Real-Time | Train. | Expl. |
|---|---|---|---|---|---|---|---|---|---|
| Classical Filters (EKF/UKF) [58] | Pose tracking | No | Yes | Any | Medium | <1 ms | Yes | No | High |
| Particle Filters (MCL) [7] | Global loc. | Grid | Yes | LiDAR/cam. | Medium | 10–500 ms | Limited | No | High |
| Graph-based SLAM [62,63] | Pose + map | Landmark | Yes | Any | High | 10–100 ms | Yes | No | High |
| Visual SLAM (ORB-SLAM3) [32] | Pose + map | Sparse | Yes | Camera | High | 20–100 ms | Yes | No | Medium |
| LiDAR SLAM (LIO-SAM) [11] | Pose + map | Point cl. | Yes | LiDAR + IMU | Very High | 50–200 ms | Yes | No | High |
| Visual-Inertial (VIO) [75,78] | Pose tracking | Sparse | Yes | Cam. + IMU | High | 10–50 ms | Yes | No | Medium |
| Multi-modal SLAM [84,86] | Pose + map | Dense | Yes | LiDAR + Cam + IMU | Very High | 100–500 ms | Yes | No | Medium |
| Deep Learning (APR) [15,97] | Relocalization | No | Yes | Camera | Low | 5–50 ms † | Yes | Yes | Low |
| Scene Coord. Regress. [98] | Relocalization | Implicit | Limited | Camera | High | 0.1–1 s † | Limited | Yes | Low |
| NeRF-SLAM [87,88] | Pose + dense map | Neural | No | RGB-D | Very High | >30 s † | No | No | Low |
| 3DGS SLAM [91,92] | Pose + dense map | Gaussian | Partial | RGB-D/Stereo | Very High | >1 s † | No | No | Low |
| Platform | Architecture | INT8 TOPS (Peak) | Mem BW (GB/s, Peak) | TDP (W) | TOPS/W | DRAM |
|---|---|---|---|---|---|---|
| STM32H7 b | ARM Cortex-M7, 480 MHz | — | 3.2 | <0.5 | — | 1 MB SRAM |
| Raspberry Pi 4 [12] | ARM Cortex-A72, 1.8 GHz | — | 25.6 | 5–7 | — | 8 GB LPDDR4 |
| Jetson Orin Nano b | Ampere iGPU + A78AE | 40 | 68 | 5–15 | 4.0 | 8 GB LPDDR5 |
| Jetson Orin NX 16 GB b | Ampere iGPU + A78AE | 70 | 102 | 10–25 | 4.0 | 16 GB LPDDR5 |
| Jetson AGX Orin 64GB [12,82] | Ampere iGPU + A78AE | 275 | 204 | 15–60 | 7.3 | 64 GB LPDDR5 |
| Xilinx ZU9EG b | FPGA + ARM A53 | ∼8 (DSP) | 34 | 10–20 | ∼0.5 | 4 GB DDR4 |
| Intel Loihi 2 [124] | Neuromorphic, 128 cores | ∼15 † | 1.8 | <1 | ∼30 † | 128 MB SRAM |
| NVIDIA RTX 3090 [91,92] | Ampere GA102 | 568 | 936 | 350 | 1.6 | 24 GB GDDR6X |
| Platform | Class | Power | Filter SLAM | VIO/LIO | Neural SLAM |
|---|---|---|---|---|---|
| ARM Cortex-M7 (STM32H7) | MCU | <1 W | Dead-reckoning only | No | No |
| Raspberry Pi 4 (Cortex-A72) | ARM SoC | 5 W | Yes | Yes (10–30 Hz) | No |
| Jetson Orin Nano (40 TOPS) | Emb. GPU | 5–15 W | Yes | Yes (>100 Hz) | No |
| Jetson AGX Orin (60 TOPS) | Emb. GPU | 15–60 W | Yes | Yes (>100 Hz) | Partial ★ |
| Xilinx ZU+ FPGA | FPGA | 5–15 W | Yes | Front-end only | No |
| Intel Loihi 2 | Neuromorphic | <1 W | Event-cam odometry | No | No |
| NVIDIA RTX 3090 (desktop) | Desktop GPU | 350 W | Yes | Yes | Yes |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Guerrero Hernández, J.M.; Pérez-Rodríguez, R.; Cely, J.S.; Aguado, E.; Martín Rico, F. Mobile Robot Localization and SLAM: A Critical Review of Sensors, Multi-Sensor Fusion, and Neural Representations. Robotics 2026, 15, 142. https://doi.org/10.3390/robotics15080142
Guerrero Hernández JM, Pérez-Rodríguez R, Cely JS, Aguado E, Martín Rico F. Mobile Robot Localization and SLAM: A Critical Review of Sensors, Multi-Sensor Fusion, and Neural Representations. Robotics. 2026; 15(8):142. https://doi.org/10.3390/robotics15080142
Chicago/Turabian StyleGuerrero Hernández, José Miguel, Rodrigo Pérez-Rodríguez, Juan S. Cely, Esther Aguado, and Francisco Martín Rico. 2026. "Mobile Robot Localization and SLAM: A Critical Review of Sensors, Multi-Sensor Fusion, and Neural Representations" Robotics 15, no. 8: 142. https://doi.org/10.3390/robotics15080142
APA StyleGuerrero Hernández, J. M., Pérez-Rodríguez, R., Cely, J. S., Aguado, E., & Martín Rico, F. (2026). Mobile Robot Localization and SLAM: A Critical Review of Sensors, Multi-Sensor Fusion, and Neural Representations. Robotics, 15(8), 142. https://doi.org/10.3390/robotics15080142

