Next Article in Journal
Real-Time Robotic Navigation with Smooth Trajectory Using Variable Horizon Model Predictive Control
Previous Article in Journal
QoS and Grid-Shifting Ability Guaranteed Optimal Capacity Sizing Method of Battery Swapping Station Considering Seasonal Characteristics
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

Multi-Camera Simultaneous Localization and Mapping for Unmanned Systems: A Survey

1
College of Electronic Science and Technology, National University of Defense Technology, Changsha 410073, China
2
School of Artificial Intelligence and Robotics, Hunan University, Changsha 410012, China
*
Author to whom correspondence should be addressed.
Electronics 2026, 15(3), 602; https://doi.org/10.3390/electronics15030602
Submission received: 20 December 2025 / Revised: 17 January 2026 / Accepted: 21 January 2026 / Published: 29 January 2026

Abstract

Autonomous navigation in unmanned systems increasingly relies on robust perception and mapping capabilities under large-scale, dynamic, and unstructured environments. Multi-camera simultaneous localization and mapping (MCSLAM) has emerged as a promising solution due to its improved field-of-view coverage, redundancy, and robustness compared to single-camera systems. However, the deployment of MCSLAM introduces several technical challenges that remain insufficiently addressed in existing literature. These challenges include the high-dimensional nature of multi-view visual data, the computational cost associated with multi-view geometry and large-scale bundle adjustment, and the strict requirements on camera calibration, temporal synchronization, and geometric consistency across heterogeneous viewpoints. This survey provides a comprehensive review of recent advances in MCSLAM for unmanned systems, categorizing existing approaches based on system configuration, field-of-view overlap, calibration strategies, and optimization frameworks. We further analyze common failure modes, evaluate representative algorithms, and identify emerging research trends toward scalable, real-time, and uncertainty-aware MCSLAM in complex operational environments.

1. Introduction

In the contemporary landscape of intelligent industries, Simultaneous Localization and Mapping (SLAM) has emerged as a critical research frontier, addressing the fundamental challenges of precise localization and environmental mapping in unknown scenarios. Its applications permeate diverse sectors, ranging from autonomous navigation and planetary exploration to immersive technologies like augmented reality (AR) and virtual reality (VR). Particularly, SLAM plays a crucial role in enabling safe operations in high-risk environments and optimizing decision-making in smart manufacturing systems.
Single-camera and single-sensor SLAM systems have demonstrated impressive progress in controlled environments, yet they remain fundamentally constrained when operating in large-scale, unstructured, or visually degraded settings. These systems are susceptible to scale drift, limited field-of-view, sensitivity to occlusion, and performance degradation in texture-sparse or highly dynamic regions. As a result, their perception and localization accuracy often deteriorate when the environment lacks stable geometric or photometric features. In contrast, multi-camera SLAM architectures leverage wider baseline coverage, redundant viewpoints, and increased geometric constraints to improve robustness, maintain tracking under partial occlusions, and enhance map consistency. The availability of multiple synchronized visual streams enables more stable feature association and reduces the risk of catastrophic tracking failures that typically occur in single-sensor configurations. These inherent advantages form the core motivation for adopting MCSLAM in unmanned systems operating across expansive, cluttered, and unpredictable environments.
SLAM algorithms are typically categorized by sensor modalities, including photonic information captured based on LiDAR [1,2,3,4], visual [5,6,7,8], infrared [9,10], MMwave Radar [11], and sonar-based algorithms [12,13], each presenting unique advantages. For example, LiDAR-based SLAM achieves superior accuracy but remains constrained by high equipment costs. Visual SLAM leverages rich texture information but exhibits sensitivity to lighting conditions and texture sparsity. Infrared SLAM operates effectively in low-light conditions despite limited texture discrimination capabilities. Sonar-based SLAM demonstrates specialized utility for underwater robotic applications.
As shown in Table 1 and Figure 1, systematic reviews of SLAM published since 2015 reveal a pronounced emphasis on visual SLAM systems compared to other modalities, reflecting their rapid technological maturation. However, existing surveys exhibit notable gaps in comprehensive analysis of multi-camera SLAM (MCSLAM). This review addresses this deficit by conducting an in-depth examination of MCSLAM literature published over the past two decades, curated from Web of Science, IEEE Xplore, and Google Scholar databases. Building upon the foundational framework of visual SLAM, this paper endeavors to analyze recent advancements in MCSLAM, identify unresolved challenges, and project future research trajectories.
The proliferation of computational resources and advancements in computer vision, particularly in feature extraction [14,15], descriptor matching [16,17], loop closure detection [18,19], and mapping [20,21], have accelerated the adoption of visual SLAM due to its cost-effectiveness, low power consumption, and compact form factor. Fundamentally, these systems operate by using image sensors (such as CMOS or CCD) to capture environmental photons and convert this photonic information into the electrical signals that constitute a digital image. The robustness of a visual SLAM system is therefore highly contingent upon the sensor’s ability to reliably capture and resolve this photon information. Consequently, under adverse lighting conditions, such as insufficient illumination (leading to photon sparsity) or overexposure (resulting in sensor saturation), the performance of feature extraction and matching can be substantially compromised. While monocular cameras benefit from minimal hardware complexity, they suffer from scale ambiguity and require significant inter-frame parallax for accurate depth estimation. Stereo cameras mitigate scale issues but remain vulnerable to baseline limitations in distant object tracking. Critically, both monocular and stereo cameras exhibit inherent limitations in large-scale environments, dynamic lighting conditions, low-texture regions, and highly dynamic scenarios.
Table 1. SLAM Research Surveys since 2015.
Table 1. SLAM Research Surveys since 2015.
ReferenceTitleTypeYear
Kostavelis et al. [22]Semantic Mapping for Mobile Robotics Tasks: A SurveyVisual SLAM2015
Yousif et al. [23]An Overview to Visual Odometry and Visual SLAM:
Applications to Mobile Robotics
2015
Lowry et al. [24]Visual Place Recognition: A Survey2016
Taketomi et al. [25]Visual SLAM Algorithms: A Survey from 2010 to 20162017
Saputra et al. [26]Visual SLAM and Structure from Motion in Dynamic Environments: A Survey2018
Jamiruddin et al. [27]RGB-Depth SLAM Review2018
Chen et al. [28]A Review of V-SLAM2018
Duan et al. [29]Deep Learning for Visual SLAM in Transportation Robotics: A Review.2019
Garg et al. [30]Semantics for Robotic Mapping, Perception and Interaction: A Survey2020
Chen et al. [31]A Survey on Deep Learning for Localization and Mapping:
Towards the Age of Spatial Machine Intelligence
2020
Zeng et al. [32]View Planning in Robot Active Vision:
A Survey of Systems, Algorithms, and Applications
2020
Xia et al. [33]A Survey of Image Semantics-Based Visual Simultaneous Localization and Mapping:
Application-Oriented Solutions to Autonomous Navigation of Mobile Robots
2020
Servières et al. [34]Visual and Visual–Inertial SLAM:
State of the Art, Classification, and Experimental Benchmarking
2021
Tsintotas et al. [35]The Revisiting Problem in Simultaneous Localization and Mapping:
A Survey on Visual Loop Closure Detection
2022
Chen [36]Semantic Visual Simultaneous Localization and Mapping: A Survey2022
Agostinho et al. [37]A Practical Survey on Visual Odometry for Autonomous Driving in
Challenging Scenarios and Conditions
2022
Fabio et al. [38]How NeRFs and 3D Gaussian Splatting are Reshaping SLAM: a Survey2024
Zhuang et al. [39]Visual SLAM for Unmanned Aerial Vehicles: Localization and Perception2024
Wang et al. [40]A Survey of Visual SLAM in Dynamic Environment:
The Evolution From Geometric to Semantic Approaches
2024
Chong et al. [41]Sensor Technologies and Simultaneous Localization and Mapping (SLAM)Multi-sensor SLAM
(including vision)
2015
Cadena et al. [42]Past, Present, and Future of Simultaneous Localization and Mapping:
Toward the Robust-Perception Age
2016
Saeedi et al. [43]Multiple-Robot Simultaneous Localization and Mapping: A Review2016
Zaffar et al. [44]Sensors, Slam and Long-Term Autonomy: A Review2018
Sualeh et al. [45]Simultaneous Localization and Mapping in the Epoch of Semantics: A Survey2019
Huang et al. [46]A Survey of Simultaneous Localization and Mapping with An Envision in 6G Wireless Networks2019
Zhao et al. [47]Review of SLAM Techniques for Autonomous Underwater Vehicles2019
Rosen et al. [48]Advances in Inference and Representation for Simultaneous Localization and Mapping2021
Liu et al. [49]Simultaneous Localization and Mapping Related Datasets: A Comprehensive Survey2021
Hriday et al. [50]From SLAM to Situational Awareness: Challenges and Survey2023
Reza et al. [51]Hardware Implementation of SLAM Algorithms:
A Survey on Implementation Approaches and Platforms
2023
Tezerjani et al. [52]A Survey on Reinforcement Learning Applications in SLAM2024
Debeunne et al. [53]A Review of Visual-LiDAR Fusion Based Simultaneous Localization and MappingVisual-LiDAR SLAM2020
Kolhatkar et al. [54]Review of SLAM Algorithms for Indoor Mobile Robot with LiDAR and RGB-D Camera Technology2020
Yang et al. [55]A Survey of SLAM Research Based on LiDAR SensorsLiDAR SLAM2019
MCSLAM leverages expanded Field-of-View (FOV) and multi-sensor redundancy to achieve superior localization accuracy and environmental resilience [56,57]. By integrating three or more cameras with overlapping FOV, MCSLAM enhances spatial perception in complex terrains while mitigating scale ambiguity through geometric baseline triangulation [58]. This configuration provides inherent advantages for navigation in unstructured environments, such as forests, subsurface voids, and dynamic industrial settings, where traditional visual SLAM struggles with occlusion, texture sparsity, and dynamic lighting conditions. Importantly, existing MCSLAM studies have been validated across a diverse range of real-world datasets, including indoor manufacturing facilities, warehouses, outdoor industrial sites, and large-scale infrastructure environments, demonstrating strong generalizability across different industrial scenarios. Recent innovations in sensor fusion algorithms, real-time optimization pipelines, and deep-learning-based feature extraction have further accelerated MCSLAM’s evolution, enabling robust performance in high-dynamic scenarios and large-scale deployments. These advancements, coupled with adaptability to diverse lighting conditions and geometric complexities, solidify MCSLAM’s role as a cornerstone of next-generation unmanned systems across robotics, planetary exploration, and immersive technologies.
Multi-camera SLAM systems can be broadly categorized into two structural configurations: rigid multi-camera assemblies with overlapping fields of view, and arbitrarily arranged camera setups with non-overlapping FOV. Rigid overlapping systems typically use a fixed mechanical frame with accurately known extrinsics, enabling direct multi-view feature matching, wide-baseline triangulation, and strongly constrained bundle adjustment. This configuration enhances depth accuracy and robustness but requires precise calibration and synchronized acquisition; at the same time, overlapping views provide inherent redundancy, allowing the system to maintain localization performance under partial camera occlusion or temporary sensor failure. In contrast, non-overlapping configurations aim to maximize environmental coverage by distributing cameras across different viewing directions. Due to the absence of shared visual content between cameras, these systems rely on motion priors, cross-frame pose estimates, or map-level fusion to achieve consistent localization, and robustness under camera dropout is typically achieved through adaptive view selection or fusion at the pose-graph level rather than direct feature redundancy. The differing requirements on calibration, data association, optimization strategies, and fault tolerance form a fundamental structural distinction that significantly influences system design and performance in multi-camera SLAM.
Compared with existing SLAM surveys published since 2015, this work provides several unique contributions:
  • A thorough review of the existing research in MCSLAM is undertaken, followed by a systematic categorization of these advancements to provide a structured overview of the field.
  • In order to support researchers in selecting datasets aligned with the characteristics of their algorithms, a detailed analysis and comparison of current MCSLAM-related datasets is presented.
  • The strengths and limitations of representative MCSLAM approaches are examined in depth, and the key challenges as well as future development trends of MCSLAM are thoroughly analyzed.
The remainder of this paper is structured as follows. Following the introduction, foundational research on camera sensors in MCSLAM is reviewed. Building on visual SLAM architecture, key advancements in MCSLAM are systematically discussed, including intrinsic and extrinsic calibration, initialization, feature extraction, pose estimation, depth estimation, and mapping. Section 4 examines recent deep-learning-driven breakthroughs in MCSLAM, while Section 5 outlines the current state of multi-sensor fusion for MCSLAM. Section 6 introduces available multi-camera datasets. Section 7 analyzes challenges and future trends in MCSLAM, and Section 8 concludes the paper.

2. Multi-Camera Systems

Multi-camera systems (MCSs) can be categorized based on lens type into pinhole camera systems and fisheye camera systems. Depending on the geometric arrangement of the installed system, MCSs can be further divided into artificial compound eye systems (ACESs) and arbitrarily configured MCSs. Table 2 summarizes the parameters of MCSs used for SLAM algorithms in papers published since 2018, and the relative proportions are presented graphically in Figure 2.
From Table 2 and Figure 2, the following points can be observed:
  • MCSs are predominantly employed in outdoor settings.
  • Research on ACESs is limited, with arbitrarily configured MCSs being more prevalent. This is mainly due to the high structural design and production demands of ACESs. Arbitrarily configured MCSs are preferred as they do not hinder the exploration of fundamental theories and technologies of MCSs. However, research on such systems may provide limited guidance for the standardization and commercialization of MCSs.
  • Most MCSs consist of fewer than five cameras. Due to the substantial data volume in MCSs, increasing the number of cameras could lead to higher computational expense. Consequently, exploring MCSs with more than five cameras is a direction warranting further research.
  • MCSs, particularly those of arbitrarily configured types, often neglect considerations of camera system synchronization or asynchrony.
In conclusion, within the realm of MCSLAM, numerous laboratories and researchers have developed and designed their own MCSs. However, a unified set of referenceable parameter standards for these sensors has yet to be established. This implies that there is ample room for research in the design and application of such sensors. More work is required to refine relevant standards, ultimately facilitating the development of industrially viable products with referenceable attributes.

3. Research on MCSLAM

The Visual SLAM estimates the camera pose with front-end data, improving the initial value for back-end optimization. Its flowchart is depicted as Figure 3, which has five parts: sensor data, front-end, back-end, mapping, and loop closure.
In accordance with the Visual SLAM flowchart, this section focuses on the research advancements of MCSLAM in terms of aspects such as the intrinsic and extrinsic calibration, initialization, feature extraction, pose and depth estimation, mapping, and research related to MCSLAM incorporating deep learning.

3.1. Intrinsic and Extrinsic Calibration

Multiple sensors have become the fundamental configuration for increasingly sophisticated unmanned systems, with sensor calibration constituting an indispensable task. Precisely calibrating the intrinsic and extrinsic parameters within an MCSLAM framework is a prerequisite and fundamental requirement for better performance in visual SLAM, alongside other positioning, navigation, and object tracking applications.
Given the limited or non-overlapping FOV among cameras, calibrating MCSs through conventional chessboard or dot pattern approaches proves arduous. Consequently, researchers have devised various calibration methods aimed at achieving more streamlined and precise calibration outcomes [87].

3.1.1. Calibration of MCSs with Strictly Non-Overlapping FOV

In the case of non-overlapping FOV, Li and Heng et al. [88] proposed a novel calibration pattern and a Matlab toolbox based on feature descriptors. These tools simplify the intrinsic and extrinsic calibration for MCSs. Their method necessitates that adjacent cameras partially observe the calibration pattern simultaneously, though these observed portions do not overlap. Kumar et al. [89] employed standard calibration methods to determine the intrinsic and extrinsic parameters of the real cameras from their mirrored poses by formulating constraints between them. However, this method requires the entire pattern to be visible in the camera, which limits its applicability. Pollok et al. [90] tackled the extrinsic calibration problem for distributed, non-overlapping MCSs. They reconstructed the scene using information captured by each camera and matched 3D points with SLAM map points. Nevertheless, this algorithm’s limitation is that pose estimation for each camera is performed separately. Zhang et al. [91] introduced an online automatic calibration technique that leverages an enhanced feature descriptor-based calibration pattern. This approach ensures precise intrinsic and extrinsic calibration for non-overlapping MCSs. It utilizes a refined parameter polynomial projection model, which enhances accuracy over a unified projection model. Additionally, it employs a novel optimization approach that extends the calibration procedure by substituting the residual function and collectively refining all parameters. However, a drawback of this method is its low efficiency. The computational burden rises with the polynomial order in the model, and the complexity of feature matching between the calibration pattern image and each image further increases computational demands.
Modern intelligent vehicles are often equipped with surround-view camera systems, which are increasingly popular for external perception. These systems can passively or actively assist drivers, particularly in parking assistance systems. However, the FOV overlap between cameras in such surround-view systems is typically minimal, to the point of being described as non-existent. Ouyang et al. [92] presented a comprehensive online optimization strategy for obtaining exterior orientation parameters, validated across simulated, indoor, and extensive outdoor environments. The authors derived the camera-to-vehicle frame rotations, and the exterior orientations were optimized as part of two-view geometry incidence relationships. Although this method does not require dedicated calibration objects or specialized infrastructure, its limitation is that the vehicle must travel straight on a flat horizontal surface.

3.1.2. Calibration with No Strict Regulations on Overlapping FOV

Traditional online calibration methods like Kalibr [93] fail in limited or non-overlapping FOV scenarios. The CamMap method [94], designed for MCSs, relaxes overlapping FOV requirements. It introduces three operational rules to mitigate error accumulation and a bidirectional reprojection cost function for precise extrinsic parameter estimation. This approach calibrates any number of cameras, regardless of resolution, frame rate, or synchronization, matching Kalibr’s accuracy with simpler operation, faster computation, and broader applicability. However, it requires at least one stereo pair for calibration.
Dexheimer et al. [95] proposed an information-theoretic online calibration framework for MCSs. Their entropy-based keyframe selection avoids complexity scaling with camera count, supporting arbitrary configurations, including overlapping or non-overlapping FOV in MCSLAM. However, this method focuses exclusively on extrinsic parameter calibration.
Su et al. [96] presented a spherical-object-based calibration for RGB-D camera networks, addressing wide camera separations via spherical view correspondence. Outperforming planar-object methods, it enables fast, accurate calibration for partially overlapping FOV but struggles with 3D spatial distortion effects.
Meng et al. [86] extended monocular RGB-D SLAM to MCSLAM with limited or no overlapping FOVs. They introduced two calibration approaches: one tailored for systems equipped with inertial measurement units (IMUs), similar to Ref. [97], and another applicable to pure visual SLAM without supplementary sensors. In the back-end stage, a series of constraints are added to the camera poses, and pose graph optimization is employed to refine the poses. While the individual breakthroughs of the two calibration methods may not be substantial, it is noteworthy that they amalgamate the two techniques to enhance the precision of extrinsic parameter calibration across diverse environments.
Tao et al. [98] proposed a low-cost calibration object using a low-sphericity sphere with dot-pattern-encoded markers, bypassing the need for expensive high-precision manufacturing. Subsequently, a graph-theory-based optimal path transformation algorithm is developed for estimating camera poses and spatial point coordinates in the global coordinate system, and a spherical projection optimization algorithm is designed to spheroidize the spatial coordinates of feature points.

3.1.3. Dynamic Calibration for MCSs

The calibration object utilized within the MCSs mentioned above adopts a fixed installation mode, where cameras are not integrated into a driving mechanism, such as the gimbal commonly employed in commercial drones, as depicted in Figure 4. This configuration of multi-camera ensemble is referred to as Dynamic Camera Cluster (DCC) systems. However, extant calibration methodologies for DCC systems [99,100] exhibit certain limitations:
  • Calibration difficulty grows with increasing degrees of freedom (DOFs);
  • Calibrating two or more cameras with overlapping FOVs but distinct individual FOV specifications necessitates multiple intricate calibration procedures;
  • Calibration procedure is achieved in a restricted range of motion.
Figure 4. A MCS mounted on a drone gimbal.
Figure 4. A MCS mounted on a drone gimbal.
Electronics 15 00602 g004
Rebello et al. [101] introduced a method for time-varying extrinsic calibration between multiple cameras, achieving full configuration space excitation through joint optimization of calibration parameters and unknown joint angles via pose error minimization. This approach eliminates the need for overlapping FOV while addressing degenerate parameter conditions caused by scale ambiguity. However, the method lacks physical experimental validation.
Das and Waslander completed a time-varying extrinsic calibration of DCCs [99], which uses the Denavit–Hartenberg convention [102] to parameterize the actuated mechanism and then determines the calibration parameters which allow for the estimation of the time varying extrinsic transformations between camera frames. The work in [99] was later extended to perform fully automatic viewpoint selection and calibration using a local optimal or suboptimal view method [100]. Although existing DCC methods provide good calibration results, the proposed methods require accurate measurements of joint angles in the mechanism. In many practical situations, accurate measurements of joint angles are not obtainable, necessitating an encoderless joint angle estimation method.
Christopher et al. [103] proposed an encoderless DCC calibration method that does not require measuring joint angles in the driving mechanism. This algorithm can be used on any driving mechanism.
The hand-eye calibration problem can be viewed as an instance of the SLAM. Heng et al. [104] performed hand-eye calibration of a multi-camera system’s external parameters, primarily using odometry to address scale measurement issues. However, as this method relies mainly on trajectory matching rather than aligning feature points or maps, phenomena such as vehicle turning or slipping can lead to inaccurate rotation estimates. Additionally, when the motion trajectory exhibits simple variations, some degrees of freedom in the calibration parameters may be unobservable. Wang et al. [105] extended the formulas in Ref. [106] to the intrinsic and extrinsic calibration of MCSs. However, it requires an external motion capture system to accurately recover the camera’s position during calibration, making the entire method less extensible or suitable for small equipment.
This section emphasizes research advancements in pure MCSs. While substantial work exists on hybrid sensor calibration, including the joint calibration of laser and MCSs [107,108,109,110] and IMU–camera integration [67,111,112,113], these topics fall beyond the current scope.
In summary, calibration methods for MCSs mainly focuses on enhancing feature descriptors, incorporation of reprojection errors, optimizing calibration objects, and other methods to achieve simpler and faster calibration of intrinsic and extrinsic parameters. It should also be pointed out that calibration in static environments is to some extent not limited by time and space. Therefore, achieving simpler, faster, and more accurate calibration of MCSs under online conditions is a great significance. Furthermore, the motion process can cause significant interference to the camera system, so some researchers have turned their attention to online calibration of MCSs and calibration methods for DCC systems.
In arbitrarily configured multi-camera systems, intrinsic and extrinsic calibration are tightly coupled, and their interaction plays a critical role in overall system accuracy, especially when camera fields of view do not overlap. Offline calibration methods are commonly employed to estimate intrinsic parameters and initial extrinsic relationships under controlled conditions, providing a reliable initialization for subsequent SLAM optimization. However, in practical deployments, mechanical tolerances, temperature variations, and long-term operation can introduce deviations that degrade the validity of offline calibration. In such cases, inaccuracies in intrinsic parameters may propagate into extrinsic estimation, leading to scale drift or biased pose optimization in multi-camera bundle adjustment. Online calibration strategies aim to mitigate these issues by refining extrinsic parameters during system operation, often leveraging motion constraints, reprojection residuals, or map consistency criteria. While online methods typically assume fixed intrinsic parameters to maintain observability and computational efficiency, recent studies indicate that partial online refinement can compensate for residual calibration errors in arbitrary, non-overlapping configurations. From a sensitivity perspective, existing studies consistently report that even small extrinsic calibration errors can accumulate over time, resulting in degraded mapping consistency and localization accuracy, highlighting the importance of calibration-aware optimization in MCSLAM systems.

3.2. Initialization

The pipeline of MCSLAM is given in Figure 5. MCSs offer distinct advantages, including extended baselines and expanded FOV, which enhance environmental perception capabilities. However, these systems introduce unique challenges, particularly during the initialization phase of MCSLAM. Inaccurate offline extrinsic calibration makes initialization more difficult. Despite the limited research specifically addressing MCSLAM initialization, its critical role in system robustness deserves dedicated examination.
In addition, image quality assessment and preprocessing play a critical role in ensuring efficient data management in multi-camera SLAM systems. Low-quality frames, caused by blur, noise, or poor illumination, can degrade feature extraction and increase computational burden. To address this, recent approaches employ learning-based filters, such as ResNet-based networks [68], to automatically evaluate and select high-quality frames for subsequent processing. By prioritizing frames with reliable visual information, these methods reduce unnecessary computations, improve feature matching robustness, and enhance overall system efficiency, particularly in large-scale or resource-constrained scenarios.
Existing initialization methods, such as matching features across different cameras [57], usually assume that cameras have a sufficient overlapping of FOV. For MCSs with a limited common FOV and inaccurate extrinsic calibration, Li et al. [58] proposes a robust initialization method. This method takes the inaccurate extrinsic poses as soft constraints to accommodate the calibration errors. By incorporating those soft pose constraints, it becomes possible to avoid false feature matching and triangulation caused by inaccurate extrinsic parameters while maintaining a limited solution when only a few feature correspondences exist. The actual test results show that even if there is a significant difference between the pose prior obtained from offline calibration and the true pose prior, this method can significantly improve the success rate of SLAM initialization.

3.3. Feature Extraction

Leveraging the wide FOV in MCSs, these systems can achieve instantaneous omnidirectional spatial information, and features can be continuously observed for a long time. These advantages of MCSs are highly beneficial for updating the system’s state. The longer the continuous feature observation time, the stronger the error convergence, effectively reducing system uncertainty and positioning errors, thus enhancing system localization accuracy [62,114].
Pedro et al. [115] proposed the idea of using mixed features (points and lines) in MCSs, providing two minimal solvers: (1) generating four solutions using two points and one line; (2) generating eight solutions using two lines and one point. The proposed minimal solvers are beneficial and robust in noisy, dynamic, and challenging road scenarios.
Kei et al. [116] introduced a line-based MCSLAM framework utilizing vehicle-mounted camera sequences. The method considered the prior distribution of line features detected in an urban environment. Then, the prior distribution is defined as a four-component Gaussian Mixture Model (GMM). A novel cost function integrating this prior distribution was formulated, enabling joint optimization of camera poses, positions, and 3D line segments through bundle adjustment. This probabilistic formulation yielded demonstrable improvements in SLAM accuracy by leveraging environmental line regularities.
Jiao et al. [117] presented a novel 2D Random Sample Consensus (RANSAC) framework integrating 3D-2D point–line feature matching for visual localization. The method derives minimal closed-form solutions using either one point with one line or two-point matches in both monocular and MCSs. By leveraging heterogeneous feature types and reducing pose computation requirements, this framework demonstrates superior outlier robustness. Additionally, a learning-based sampling strategy selection mechanism and feature scoring network enable dynamic environmental adaptability across diverse scenes, with resilience to seasonal variations and viewpoint changes. While point–line matching enhances system robustness, it introduces computational complexity. Thus, it is necessity to balance these competing factors through optimized feature selection and algorithmic design.
Marcel et al. [118] proposed efficient 2D–3D features against a pre-built global 3D map, presenting a prioritized feature matching approach tailored for MCSs. This method employs an active search strategy by interleaving prioritized feature matching with camera pose estimation. In contrast to conventional approaches, it does not terminate the search once a fixed number of matches are found; this method dynamically adjusts parameters for matching feature quantities, enhancing its adaptability across diverse scenes and reducing the computational burden of feature matching. Furthermore, it introduces a robust filtering step based on pose priors, contributing to the overall computational efficiency of the system.
Zhang et al. [119] investigated panoramic visual SLAM localization techniques, with key research points including the following: (1) spherical imaging model, representing pixel coordinates on a sphere using latitude and longitude—this method provides a concise and understandable set of equations, facilitating the backend optimization of panoramic SLAM. (2) Spherical image feature extraction and matching methods, including Spherical Oriented FAST and Rotated BRIEF [120] and the Scale-Invariant Feature Transform (SIFT) algorithm [121]—this panoramic visual SLAM algorithm aims to improve the robustness and accuracy of MCSLAM.
Kang et al. [122] introduced the abstraction of the Manhattan world hypothesis in artificial environments into MCSLAM, aiming to improve the algorithm of estimation processes. They proposed a fast method to estimate Manhattan rotation constraints. This method utilizes three-point sampling on a Gaussian sphere to directly generate Manhattan hypotheses. In the final rotation estimation refinement step, a new cost function is introduced to avoid re-associating the inner RANSAC layer in the Manhattan domain direction. The algorithm demonstrates stability in low-texture areas, especially in environments with low lighting and dynamic objects being occluded.
Wang et al. [75] proposed a Multi-camera Visual Inertial Odometry (MCVIO), aiming to improve the robustness of state estimation. They incorporate line features into the state vector of MCVIO. Line features are modeled using Plücker coordinates and a minimal parameterization standard orthogonal representation, ensuring both system accuracy and computational efficiency. Furthermore, a point–line-based initialization method is proposed to obtain accurate and reliable initial poses. Finally, a universal hypergraph optimization method is proposed for accurate and efficient state estimation. This algorithm has the characteristics of accuracy and efficiency, and ensures the computational efficiency.
In challenging environments, such as texture-sparse parking lots, state-of-the-art methods often struggle to obtain sufficiently reliable features for robust localization and consistent mapping. Yu et al. [73] proposed a multi-level visual information hierarchical fusion SLAM algorithm based on an MCS. They fuse multiple views into a surrounding environment view and establish a guidance mechanism between visual information at different levels. Different segmentation methods are used to extract advanced features such as object edges. Under the guidance of high-level information, low-level key-points are selected. The hierarchical fusion structure consists of two parts. High-level visual information is loosely fused with the wheel encoder to estimate initial data association. Then, low-level feature points of multiple views are tightly fused to facilitate robust localization. With such a structure, a consistent map can be built even under the condition that one view is lost.
External factors such as extreme illumination, motion blur, and adverse weather conditions significantly impact the robustness of multi-camera SLAM systems. For instance, overexposure or underexposure can reduce feature contrast, leading to fewer reliable keypoints for tracking. Motion blur, caused by rapid camera or platform movement, degrades feature sharpness and introduces errors in feature matching. Weather conditions such as rain, fog, or snow can occlude key visual cues and change scene appearance, reducing the consistency of multi-view observations. To mitigate these effects, recent studies have incorporated illumination-invariant feature descriptors, robust outlier rejection methods, and multi-level feature fusion strategies that combine information across cameras and temporal frames. Such approaches enhance the system’s resilience, enabling reliable localization and mapping even under challenging environmental conditions.
In summary, in MCSLAM, feature extraction leverages multi-view observations to enhance robustness and accuracy. Methods include mixed point–line features, line-based bundle adjustment, 2D–3D feature matching, and hierarchical multi-view fusion. Techniques also exploit structural priors, panoramic imaging, and visual–inertial integration to handle dynamic, low-texture, or large-scale environments, balancing descriptive richness, computational cost, and cross-view consistency.

3.4. Depth and Pose Estimation

Accurate and efficient depth estimation of feature points, coupled with robust pose estimation for MCSs, constitutes fundamental objectives in MCSLAM. While diverse methodologies exist for pose estimation (e.g., geometric optimization and deep learning approaches), many SLAM implementations demonstrate that system localization accuracy is intrinsically linked to the precision of feature depth calculations. Notably, these two estimation processes exhibit a complementary relationship in advanced SLAM frameworks: pose refinement enhances depth map consistency through bundle adjustment, while improved depth estimation facilitates more stable trajectory optimization. This paper consolidates the relevant content of both into one subsection.

3.4.1. Depth Estimation

For depth estimation in MCSs, Won et al. [123] introduced an improved lightweight neural network for omnidirectional depth estimation, demonstrating superior speed and accuracy compared to conventional architectures. By leveraging the network’s omnidirectional dense depth output, the system enhances inter-view matching and triangulation performance. The estimated depth map enables cross-view reprojection of key points, thereby optimizing feature matching efficiency. These components are subsequently integrated into a truncated signed distance function (TSDF) volume for 3D mapping.
Yang et al. [124] proposed a multi-camera depth estimation system based on semantic segmentation. This system obtains depth maps between adjacent cameras through multi-camera calibration and stereo matching. It then stitches multi-camera images through feature extraction and matching. A semantic segmentation network is used for object recognition and segmentation, establishing an object height library. Finally, based on the principle of similar triangles and the object height library, the system estimates the distance of non-overlapping areas of objects. The depth obtained by this method has certain reference value and meets the distance measurement requirements of SLAM.
Ghosh et al. [125] extended the EVO framework [126] by proposing a novel multi-event-camera 3D reconstruction algorithm. The core innovation lies in Disparity Space Image (DSI) fusion across multiple cameras without explicit data association. The study systematically evaluates various fusion strategies, including event ray summation, while leveraging the Event-based Multi-View Stereo (EMVS) mapping module [127] that enables GPU-free 3D reconstruction through event-based multi-view stereo.
Wei et al. [128] introduced SurroundDepth for self-supervised multi-camera depth estimation, whose core insight revolves around entangling multi-camera information and processing all surrounding views in a concerted manner. The cross-view transformer is executed across multiple scales to seamlessly integrate multi-view features. To attain scale-aware depth predictions, pretraining with Structure-from-Motion (SfM) and joint pose estimation are leveraged to fully exploit the multi-camera extrinsic matrices.
It should be noted that recent advances in deep depth estimation for MCSLAM predominantly employ deep learning methodologies.

3.4.2. Pose Estimation

The evolution of pose estimation in MCSLAM demonstrates significant methodological advancements. Kangni et al. [129] pioneered a panoramic image-based pose estimation method in 2007, utilizing fundamental matrix recovery for scene localization enhanced through bundle adjustment optimization. Subsequently, Rituerto et al. [130] implemented a panoramic visual SLAM based on the Extended Kalman Filter (EKF) algorithm in 2010, validating its positioning accuracy superior over monocular SLAM. Extended research by Valiente et al. [131,132,133] further optimized EKF implementations for panoramic SLAM. In 2015, Gamallo et al. [134] proposed a panoramic visual SLAM algorithm for omnidirectional cameras (OVFastSLAM) designed for heavily occluded environments. Caruso et al. [135] introduced a large-scale direct visual SLAM for omnidirectional cameras based on LSD-SLAM. Liu et al. [136] and Matsuki et al. [137] separately proposed fisheye stereo DSO and omnidirectional DSO based on DSO [138]. Forster et al. [139] and Heng et al. [140] separately proposed multi-camera SVO and fisheye stereo SVO. OpenVSLAM [141] implemented a general visual SLAM framework with high usability and scalability. This system can handle various types of camera models, such as perspective and fisheye.
Some researchers regard the MCS as a single generalized camera. Pless et al. [142] introduced a technique where every pixel in the image signifies a sample of the spatial area within the scene, rather than representing rays interacting with a sensor at a particular position. This approach offers the benefit of computing a generalized model even when the projection centers of individual cameras differ. The paper employs a generalized camera model instead of a predetermined configuration for a single camera. This generalized camera model facilitates the computation of the minimum solution for collectively estimating all camera poses. Similarly, the method proposed by Lee et al. [143] can also find the minimum solution for pose. However, it uses the classical camera model to handle each camera separately instead of using the generalized camera model. Chen and Chang [144] and Nister and Stewenius [145] proposed the GP3P minimal estimator that a multi-camera pose estimation system requires three pairs of 2D-3D correspondences.
In contrast to previous minimal solvers, Kneip et al. [146] introduced UPnP, which is an efficient non-minimal pose estimator derived from a cost function based on minimal reprojection error. Inspired by DLS [147], UPnP reformulates the cost function to depend only on the unit norm quaternions. UPnP finds the optimal rotation by solving a polynomial system [148,149].
Achieving high speed and maintaining high precision in pose and scale estimators are often conflicting goals. To simultaneously achieve both, Fragoso et al. [150] utilized prior knowledge about the solution space, proposing gDLS*. gDLS* is a pose and scale estimator of a generalized camera model that uses scale and gravity priors. The complexity of computing its parameters is O ( n ) , and gDLS* continuously improves pose accuracy in less time.

3.5. Mapping

BEV-SLAM [151] represents a pioneering graph-based SLAM framework that aligns semantically segmented Bird’s Eye View (BEV) predictions derived from monocular camera inputs. The authors introduce an innovative occlusion reasoning methodology within BEV estimation, empirically validating its critical role in enhancing spatial aggregation processes that significantly improve BEV prediction accuracy. This results in a highly adaptable SLAM capable of functioning across diverse multi-camera configurations while ensuring seamless sensor fusion capabilities.
Ochoa et al. [71] provide a new solution for the navigation of remotely operated robots in complex environments. It is an omnidirectional MCS for collision detection and avoidance in underwater vehicles. The system can output warning signals, allowing operators or control systems to easily perform avoidance operations. The system uses MSCKF to accomplish SLAM in order to create a more dense 360° map representation of the local environment, and assess the danger level of surrounding objects in real-time.
Recent advancements in mapping theories for MCSLAM have revealed persistent limitations. The primary challenge stems from the information-intensive nature of MCSLAM, which necessitates prioritizing real-time performance over map density. Consequently, current implementations predominantly generate sparse maps to ensure computational efficiency. However, further research on sparse maps is not of great significance. There is still great room for innovation in map creation research, especially in the direction of dense maps.

4. MCSLAM Based on Deep Learning

Recent years have witnessed significant progress in computer vision tasks such as detection and segmentation, largely driven by advancements in deep learning [152]. This paradigm shift has prompted researchers to explore the integration of deep learning methods into MCSLAM. While the relevant research results are not rich so far, initial findings demonstrate promising effectiveness. Consequently, this paper dedicates a subsection to list MCSLAM algorithms leveraging deep learning, with specific focus on depth estimation detailed in Section 3.4.1.
Currently, several algorithms based on deep learning use multi-view strategies to improve pose estimation. Notable algorithms include [153,154,155,156,157]. These methods typically rely on monocular pose estimation, utilizing calibrated multi-view data primarily for optimization. For example, CosyPose [153] uses a monocular detector but subsequently merges monocular scene reconstructions from multiple frames in order to achieve relative camera transformations and more accurate pose estimation. In contrast, Kaskman et al. [154] proposed a distinct approach by reconstructing sparse 3D structures from image sets lacking relative pose information.
Innovations like SuperGlue [158] leverage deep learning to advance key point detection, description, and matching. Similarly, modern SLAM frameworks exploit deep neural network (DNN)-detected objects as features, as demonstrated by CubeSLAM [159], enhancing robustness in texture-scarce environments like monotonous indoor walls.
For off-road environments, the challenges faced by SLAM are enormous due to factors such as direct sunlight, leaf occlusions, rough roads, sensor failures, and the sparsity of stable and trackable textures. Traditional visual SLAM methods are prone to these factors, leading to reduced stability and reliability. Yang et al. [83] proposed a multi-camera collaborative panoramic visual SLAM. This system rapidly establishes relationships between 3D points and pixels captured in images without distortion calibration. The loop closure detection accuracy in weak dynamic environments and sparsely textured situations are enhanced.
Karpyshev et al. [68] improved the computational efficiency and robustness of visual SLAM for mobile robots with multiple cameras and limited computing capabilities. They introduce an intermediate layer between the camera and SLAM algorithm layers, using a neural network based on ResNet18 to classify images to determine the applicability of images to robots’ localization. After estimating the images’ quality, the system uses data only from the camera with the best data quality, avoiding using image data from all cameras. This method is at least six times faster on CPU than ORB extractors and feature matchers and more than 30 times faster on GPU.
This work clearly illustrates how hardware constraints influence algorithm design choices in practical MCSLAM systems. By explicitly considering limited CPU/GPU resources, the proposed quality-aware filtering strategy significantly reduces redundant computation and data throughput, enabling real-time performance on resource-constrained platforms. Such approaches demonstrate that hardware-aware system design—through selective data utilization, lightweight neural networks, and adaptive processing—can be as critical as algorithmic accuracy for deploying MCSLAM on mobile robots. Consequently, computational efficiency and hardware feasibility should be treated as first-order considerations alongside localization performance in real-world MCSLAM applications.
Li et al. [160] introduced Weakly Supervised Object Pose Estimation (WS-OPE), a method trainable with 2D bounding box labels, object dimensions, and relative camera poses as weak supervision constraints. The paper also introduced a novel rotated Intersection over Union (IoU) loss function based on multi-view perspectives to guide the prediction of 3D bounding boxes and direct 6D pose regression. Additionally, a novel rotation estimation procedure was proposed, combining coarse rotation classification with residual rotation regression. Despite being trained with weak labels, the method is capable of predicting high-quality poses. The direct pose regression approach, without the need for consecutive refinement stages, ensures real-time performance.
In particular, Kaygusuz et al. [161] proposed a depth sensor fusion framework using a CNN-RNN hybrid model to extract spatio-temporal feature representations from a sequence of continuous images. A mixed density network (MDN) independently predicts the motion probability distribution for each camera. This method allows the model to learn the confidence or uncertainty of each camera’s motion and uses a fusion module to estimate the final pose of MCS. This uncertainty can improve the optimization ability of the system and enhance its robustness and generalization ability.
This uncertainty-aware formulation exemplifies a broader trend in MCSLAM toward explicitly modeling measurement confidence at the sensor or camera level. By representing motion estimates as probability distributions rather than deterministic values, probabilistic fusion frameworks can weight multi-view observations according to their reliability, thereby reducing the influence of degraded views caused by occlusion, illumination changes, or motion blur. Such probabilistic modeling is particularly beneficial in arbitrarily configured multi-camera systems, where heterogeneous viewpoints and non-overlapping fields of view introduce uneven measurement quality. Consequently, uncertainty-driven fusion not only improves pose estimation accuracy but also enhances system robustness and generalization across diverse environments.
Despite these advancements, deep-learning-augmented MCSLAM remains constrained by fundamental challenges inherent to the technology, including data dependency, computational resource requirements, and limited generalization capabilities. These limitations can impede real-time performance and practical implementation in MCSLAM applications.

5. Special MCSLAM

The primary focus of the aforementioned content lies in research progress regarding MCSLAM within the vision domain. However, to stimulate greater reader interest in MCSs and offer more valuable resources to related researchers, this section will provide an overview of recent advancements in MCSLAM based on multi-sensor fusion, as well as other special MCSLAM algorithms.

5.1. MCSLAM Based on Multi-Sensor Fusion

Visual sensors are vulnerable to rapid motion or sudden changes in lighting conditions. This weakness can be compensated for by incorporating information from different sensors, such as IMU or LiDAR.
Oskiper et al. [162] proposed a multi-camera Visual–Inertial Odometry, termed VIO-SLAM. The algorithm extracts frame-to-frame motion constraints through three-point RANSAC and fuses these constraints with IMU data using an Extended Kalman Filter. Lee et al. [163] presented a four-point solution based on Generalized Essential Coordinates (GECs) for an MCS on autonomous vehicles, assuming roll and pitch can be directly measured from IMU.
These approaches illustrate how explicit geometric and probabilistic constraints are embedded in multi-camera pose estimation. In VIO-SLAM, the three-point RANSAC formulation enforces epipolar consistency between successive frames, while the Extended Kalman Filter incorporates inertial measurements as motion priors to constrain scale, rotation, and temporal consistency. Similarly, the four-point GEC-based method exploits known roll and pitch angles from IMU measurements to reduce the degrees of freedom in pose estimation, thereby simplifying the solution space and improving numerical stability. By explicitly encoding such mathematical constraints—ranging from minimal geometric solvers to probabilistic filtering—these methods enhance pose observability and robustness in multi-camera systems, particularly in dynamic or large-scale environments.
Jaekel et al. [164] proposed a novel multi-stereo Visual–Inertial Odometry (VIO) framework that can merge arbitrary pairs of non-overlapping stereo FOVs, aiming to enhance the robustness of state estimation during aggressive motion and in visually challenging environments. In the paper, a one-point RANSAC algorithm is proposed, which is able to perform outlier rejection across features from all stereo pairs. The proposed algorithm is able to maintain a state estimation in scenarios where traditional VIO algorithms fail.
Ye et al. [165] proposed a robust and efficient estimation method using multiple cameras, odometry, and a gyroscope. They derived pre-integration of the odometer and gyroscope and estimated sensor biases in a tightly coupled sliding window optimization framework. The proposed approach can robustly and efficiently estimate the motion of vehicles.
Tschopp et al. [166] presented the VersaVIS system. It is an open, versatile multi-camera visual–inertial sensor system designed as an efficient research platform for easy deployment, integration, and extension in various mobile robotic applications. VersaVIS provides a complete open-source hardware, firmware, and software package that enables the time synchronization of multiple cameras through an IMU. It can serve as a platform for simulating and verifying algorithms. The synchronization accuracy of the framework is evaluated in multiple experiments, achieving a timing accuracy of less than 1 ms.
Eckenhoff et al. [76] developed a versatile and adaptable multi-IMU and multi-camera SLAM algorithm capable of seamlessly integrating multi-modal visual–inertial data from an arbitrary number of uncalibrated cameras and IMUs. Operating within an efficient multi-state constraint Kalman filter framework, the proposed algorithm optimally merges asynchronous measurements from all sensors, ensuring smooth, uninterrupted, and precise 3D motion tracking even in the event of sensor failures. The core concept of the proposed MIMC-VINS involves employing high-order on-manifold state interpolation to effectively handle all available visual measurements without escalating computational demands associated with estimating additional sensor poses during asynchronous imaging intervals.
Seok et al. [82] introduced ROVINS (Robust Omnidirectional Visual Inertial Navigation System), a novel framework featuring 360° FOV coverage with stereo vision overlaps. The system achieves seamless integration of inertial measurements into its pose optimization pipeline by incorporating IMU-derived relative motion constraints as soft-pose priors. This formulation significantly enhances both the accuracy and stability of pose estimation within the ROVINS framework.
Müller et al. [167] utilized two pairs of stereo cameras (looking up and down) on micro aerial vehicles (MAVs). The visual output and IMU were loosely fused using an Extended Kalman Filter. However, the application of this achievement is limited to indoor office buildings.
Zhang et al. [74] proposed a factor graph optimization-based multi-camera visual–inertial tracking system. This system tightly fuses tracking features from any number of stereo and monocular cameras and IMU measurements while maintaining a fixed overall feature budget. This paper mainly focuses on motion tracking in challenging environments, such as narrow corridors, dark spaces, and sudden lighting changes. A simple and effective method has been proposed to track features with overlapping FOV to reduce repetitive landmark tracking and improve accuracy. A submatrix feature selection (SFS) scheme was adopted to select the optimal landmark for optimization under a fixed feature budget. Compared to using all available features, this limits the calculation time and achieves better accuracy.
Mixture of Experts-based Visual Odometry (MIXO) [64] introduced a data-driven, machine-learning-based technique that loosely combines odometry outputs from multiple cameras to achieve more accurate and reliable global estimates. In MIXO, each camera (or expert) is independently processed. One of the advantages of MIXO is that it is a lightweight module that can be easily implemented on top of any visual odometry method. MIXO achieves more robust and accurate results than any single camera, reducing the absolute rotation and translation errors by 38% and 15% respectively. Compared to ORB-SLAM2 [168], GPU optimization is not required. But one drawback of MIXO is that it is trained in a supervised manner. Therefore, it requires a large dataset with known reference standards and may not be well generalizable to different scenarios.
Shen et al. [169] introduced an Efficient Multi-Camera-Assisted LiDAR-Inertial Odometry algorithm, utilizing methods such as removing LiDAR noise, setting nearest neighbor search conditions, and replacing kd-Trees with ikd-Trees to improve efficiency and enhance system robustness.
Wang et al. [59] introduced MAVIS, an innovative optimization-based Visual–Inertial SLAM algorithm specifically designed for multiple partially overlapped camera systems. To bolster tracking performance, particularly under conditions of rapid rotational motion and extended integration durations, the authors proposed an enhanced IMU pre-integration formulation grounded in the exponential function of an automorphism of S E 2 ( 3 ) . Additionally, they extended the traditional front-end tracking and back-end optimization modules, initially tailored for monocular or stereo setups, to seamlessly integrate with MCSs.
For asynchronous MCSs, Wang et al. [60] introduce a robust feature-aware multi-camera-IMU state estimator. This estimator consists of parallel front-ends, a front-end coordinator, and a back-end optimization module. It effectively leverages input frames through the use of dynamic feature number allocation and a frame priority coordination strategy. The estimator’s exceptional robustness and performance have been rigorously validated through real-flight experiments in a variety of challenging scenarios.
Overall, the combination of MCSLAM and IMU sensors is more common because visual sensors and IMU sensors have complementary characteristics. Visual sensors have slow frequency but no cumulative error, while IMU has cumulative error but fast frequency. These multi-sensor fusion-based systems leverage the complementary strengths of visual and inertial sensors, leading to more robust and accurate SLAM solutions, especially in challenging environments. The fusion of LiDAR and visual sensors needs further deepening and improvement, both in terms of methodology and theory.

5.2. Others

The enhancement of MCSLAM in this chapter goes beyond the basic framework of visual SLAM. They refine MCSLAM from a completely new perspective or integrate various MCSLAM approaches, which is of significant reference value for the development of MCSLAM.
Kaess [170] introduced an innovative probabilistic methodology for data association, which considers the possibility of features moving between cameras during robot motion. This approach overcomes the challenges of combinatorial data association by employing an incremental expectation maximization algorithm.
Multi-Camera PTAM (MCPTAM) [171,172] modified the viewpoint camera model to a generic polynomial model. This system is applicable to cameras with minimal or no overlapping fields of view.
Kuo et al. [173] made significant changes to the MCSLAM, including the addition of an adaptive initialization scheme, key frame selection algorithm, and voxel mapping. These enhancements allow for increased robustness by creating complex multi-camera configurations. Multicol-SLAM [174] is an improvement over the ORB-SLAM [175], extending it to MCSLAM by implementing multi-keyframes and multi-camera loop closures, along with some performance enhancements. This approach enhances perception quality and robustness in challenging environments.
Existing MCSLAM methods often assume that all camera shutters are synchronized, but this is not always the case in practical scenarios. In this work [176], a generalized formula for MCSLAM is proposed. A continuous-time motion model is integrated into the framework, incorporating correlated information across asynchronous multi-frames during tracking, local mapping, and loop closure. This represents a fully asynchronous continuous-time multi-camera visual SLAM system designed for large outdoor environments. This work highlights a fundamental limitation of many existing MCSLAM systems, namely the strong assumption of strict temporal synchronization across cameras. In real-world deployments, heterogeneous sensors often operate with different frame rates, rolling shutters, or unsynchronized clocks, making perfect synchronization difficult or even infeasible. By adopting a continuous-time motion representation, asynchronous MCSLAM frameworks relax this assumption and allow measurements from different cameras to be fused at their true acquisition times.
Zou et al. [57] investigated the visual SLAM problem in dynamic environments with multiple cameras. Their cooperative SLAM (CoSLAM) implementation achieved real-time operation (38 ms per frame) using inter-camera tracking and mapping. It also employed a method to differentiate static background points from dynamically moving foreground objects.
Yang et al. [62] proposed MCOV-SLAM, which harnesses the observability enhancement of an MCS to achieve more precise pose estimation. Additionally, it capitalizes on the omnidirectional perception characteristics to enable omnidirectional loop-closing, thereby lifting the constraint of a specific sensor direction and presenting more opportunities for global optimization.
Song et al. [61] introduced BundledSLAM, a visual SLAM approach tailored for MCS. This approach integrates image data from various cameras into a unified “bundled frame” structure, enabling real-time pose tracking, local mapping for optimizing poses and map points, and loop closing for ensuring global consistency.
From a methodological perspective, many foundational ideas underlying multi-camera SLAM are deeply rooted in the theory of distributed estimation and multi-agent filtering. In multi-camera settings, particularly when cameras are spatially separated, loosely synchronized, or deployed across multiple platforms, the estimation problem naturally aligns with distributed state estimation paradigms rather than centralized filtering. Early works on distributed Kalman filtering established the theoretical basis for fusing information from multiple sensors while maintaining consistency and scalability, as exemplified by the seminal study on distributed Kalman filtering for sensor networks [177]. Subsequent advances introduced consensus-based information fusion mechanisms that enable distributed agents to achieve global estimation objectives through local communication, as formalized in embedded average consensus frameworks [178]. More recent studies have further analyzed the dynamics and convergence properties of nonlinear distributed filtering and learning in multi-agent systems [179], providing insights that are directly relevant to decentralized SLAM and cooperative perception. Comprehensive bibliographic reviews on distributed Kalman filtering [180] also highlight the evolution of these methods and their applicability to large-scale, sensor-rich systems. Collectively, these works form an important theoretical foundation for understanding multi-camera and multi-agent SLAM, particularly in terms of scalability, robustness, and information consistency across distributed sensing architectures.

6. Datasets

As various visual SLAM algorithms continue to evolve, their accuracy and effectiveness have been steadily improved. Consequently, an increasing number of datasets collected in real-world environments have been proposed to provide performance benchmarks and evaluation references for newly developed methods. In this work, we summarize publicly available datasets for MCSLAM since 2009, with detailed parameters reported in Table 3, Table 4 and Table 5. While these datasets cover a wide range of sensor configurations and environmental conditions, most of them are collected over relatively short time spans and lack long-term temporal or seasonal variations. This limitation restricts the evaluation of long-term robustness and generalization of MCSLAM systems. We further introduce several representative datasets and discuss the specific challenges they are designed to address.
These datasets are selected to reflect the diverse challenges encountered in MCSLAM research. Specifically, some datasets focus on large-scale outdoor environments to evaluate long-term drift and scalability, while others emphasize indoor scenes to assess calibration accuracy and multi-view consistency. Datasets containing dynamic objects, illumination variations, or asynchronous sensor configurations are particularly relevant for benchmarking robustness in feature extraction, data association, and synchronization. By covering these representative challenges, the collected datasets provide a comprehensive basis for evaluating and comparing MCSLAM algorithms under realistic conditions.
Sewtz et al. [63] presented a dataset for indoor environments suitable for low-cost MCSLAM, along with high-precision ground truth motion data. It comprises five different scenes and can be utilized to evaluate the MCSLAM algorithm in both static and changing environments.
M2DGR [69] is a novel large-scale dataset collected by a ground robot, including six fisheye cameras, a sky-facing RGB camera, an infrared camera, an event camera, a visual–inertial sensor, an IMU, a LiDAR, a consumer-grade global navigation satellite system (GNSS) receiver, and a GNSS-IMU navigation system with real-time kinematic (RTK) signal. All these sensors are well-calibrated and synchronized. Ground truth trajectories are obtained using motion capture devices, a laser 3D tracker, and an RTK receiver. The dataset includes 36 sequences (about 1TB) captured in various scenes, including indoor and outdoor environments.
The contribution of the paper [72] is the benchmark dataset that incorporates multiple sensors, including an event-based stereo camera, a conventional stereo camera, multiple depth sensors, and an IMU. This setup is fully hardware-synchronized and undergoes precise external calibration. All data have real ground data captured with high-precision external reference devices such as motion capture systems. The data sequences encompass small and large-scale environments, as well as complex environments including underground spaces and textureless environments, along with environments featuring low or dynamically changing lighting conditions.
Ali et al. [79] introduced a challenging outdoor dataset featuring authentic forest landscapes captured in the outskirts of Tampere, Finland. The sequences encompass diverse environmental conditions encountered during both summer and winter seasons. The dataset comprises semi-structured forest paths within highly similar natural surroundings, varying in factors such as lighting, weather, vegetation, and infrastructure. Moreover, the sequences encompass scenarios representative of typical forestry activities, including stationary scenes, rapid movements, terrain irregularities, slopes, and oscillating motion. Additionally, the dataset includes scenes showcasing common forestry elements such as logs, close-ups of trees, and off-road pathways. Sensors include four RGB cameras, an inertial measurement unit, and a global navigation satellite system receiver. The sensors are synchronized based on non-drift time stamps.
Dataset [187] includes data recorded in scenarios related to autonomous driving and ground robot navigation. Only a few provide LiDAR data, where changing lighting conditions and moving objects pose the main challenges in recording datasets for autonomous driving scenes.
The more challenging UZH-FPV dataset [194] is specifically designed for drone racing and includes a series of aggressive flight trajectories. The drone is equipped with a mini DAVIS346 (346 × 260, event camera and APS frames, six-axis built-in IMU), a wide-angle camera, and a stereo camera with a fisheye lens. The UZH-FPV Drone Racing dataset consists of over 27 sequences, covering more than 10 km of flight distance, captured from the first-person view (FPV) of a racing quadrotor piloted by an expert. These sequences are faster and more challenging in terms of apparent scene motion compared to any existing dataset. While previously presented datasets include only a primary viewing direction, the FOV size can be significantly expanded by deploying multiple sensing devices with differing orientations. However, most representatives of datasets that employ this approach, such as the NCLT dataset [187] and PennCOSYVIO dataset [190], neither include high-precision ground truth information nor a hardwired time synchronization between IMU and the relevant sensors.
Yogamani et al. [195] unveiled the inaugural comprehensive fisheye car dataset, known as WoodScape, featuring four surround-view cameras and encompassing nine tasks such as segmentation, depth estimation, 3D bounding box detection, and dirt detection. WoodScape [195] offers instance-level semantic annotations for 40 classes across more than 10,000 images, along with supplementary annotations for other tasks covering over 100,000 images
Zhang et al. [207] introduced a multi-camera-LiDAR-inertial dataset spanning 4.5 km. All devices have undergone hardware synchronization, which is more accurate than software synchronization. The dataset also provides ground truth poses using LiDAR data. Utilizing LiDAR data as real data, this dataset encompasses small and narrow passages, large open spaces, and vegetated areas. Moreover, some sequences present challenging situations such as sudden changes in lighting, textureless surfaces, and aggressive motions. The only dataset with similar sensors to the data in Ref. [207] is the Hilti SLAM Challenge dataset [211]. The Hilti SLAM Challenge dataset focuses on construction site environments and provides sparse ground truth poses.
The TUM-VIE [209] dataset is captured by a pair of ee Gen4 CD event cameras (1280 × 720, event only), a stereo camera, and a six-axis IMU. These sequences are recorded in different-scale environments and under various motion conditions, such as walking, running, skating, and biking. Ground truth poses are provided by a motion capture system.
Helmberger et al. [211] delineated a novel publicly available dataset captured via a multi-sensor platform encompassing visual, inertial, and LiDAR modalities. This dataset comprises authentic real-world data captured across diverse settings including indoor offices, laboratories, indoor and outdoor architectural environments, and outdoor parking lots. It encompasses challenging featureless regions and diverse lighting conditions. The sensor suite utilized comprises five cameras (including one stereo pair), two LiDAR units, and three IMUs meticulously calibrated in both spatial and temporal domains.
Apolloscape [192] further advance the scale of annotations. ApolloScape contains much large and richer annotaions including holistic semantic dense point cloud for each site, stereo, per-pixel semantic labeling, lanemark labeling, instance segmentation, 3D car instances, and highly accurate location for every frame in various driving videos from multiple sites, cities, and times of day.
MVSEC [193] is considered a modern cross-modal dataset, featuring a rich array of sensors, including a pair of DAVIS m346B sensors (346 × 260, event and APS frames, with a six-axis built-in IMU and an approximately 10 cm baseline), a VI-Sensor with a stereo camera (an approximately 10 cm baseline) and a built-in nine-axis IMU, and a 16-channel LiDAR. The provided sequences can be classified by motion type, as they are recorded by six-axis aircraft, handheld devices, driving cars, and motorcycles. Ground truth for event frames and depth maps is generated by Cartographer [213] and LOAM [214], respectively.
A dataset [215] focusing on agricultural robot scenarios recorded across various types of agricultural environments during autumn is proposed. The sensors utilized include a DVS240 (240 × 180, event only, with a built-in six-axis IMU), a stereo camera, an RGB-D sensor, and a 16-channel LiDAR. All sensors, except for the RGB-D sensor, are hardware synchronized. Due to technical limitations, GPS data is not included in the dataset, and ground truth is approximated using three different LiDAR SLAM algorithms.
The Lafida [189] dataset contains three fish-eye cameras on a helmet, but the maximum recording time of its sequences is too short for long-term evaluation (usually longer than 20 min).
It should be noted that, to test asynchronous MCSLAM, Yang et al. [176] collected the AMV-Bench dataset. It is a challenging new SLAM dataset covering 482 km of driving recorded using their asynchronous multi-camera robot platform. AMV-Bench covers diverse and challenging motions and environments, such as low-light scenes, occlusions, fast driving, and complex maneuvers like three-point turns and reversing.
These datasets are crucial for testing MCSLAM algorithms, each tailored to its specific environment. Researchers can choose and utilize datasets based on their requirements. However, the scenes depicted in these datasets alone are insufficient. It is imperative for more researchers to develop additional datasets to further bolster related research endeavors.

7. The Development Trend of MCSLAM

This paper reviews the MCSLAM algorithms developed in recent years. Despite extensive research and numerous improvement methods proposed for various modules and system levels of MCSLAM, many scientific challenges remain unresolved, preventing these systems from meeting current engineering requirements. To explore the potential of MCSLAM, address practical needs, and achieve more robust, accurate, and efficient results in future SLAM development, we have summarized future open research directions for researchers.
  • MCSLAM based on deep learning: It faces severe challenges in small-sample and dynamic environments. Deep neural networks have shown impressive results in various applications, and they have become an important development trend in the SLAM field. They can serve as reliable feature extractors and have solved many cognitive and learning tasks that rely on human design feature cannot achieve. The advantages of deep learning in object detection, semantic segmentation, dense matching, and other fields have become prominent and significantly superior to traditional methods. However, deep learning methods heavily rely on large datasets and their performance is limited by the size of the dataset. At the same time, the performance of deep learning in dynamic environments remains problematic. Enhancing the adaptability of deep learning in scenarios with small datasets, large-scale environments, and dynamic conditions is of great significance.
  • MCSLAM based on point–line-plane feature matching: Solely relying on point features may not be sufficient for various environments. The combination of point–line-plane features can achieve more robust feature matching, reduce system drift errors, and improve the system’s understanding of the environment.
  • Balancing information retrieval and computation cost: Sparse map computation has lower costs but provides less information. Dense maps, while recording complete scene information, require real-time computing performance. MCSs generate a massive amount of information, and improving the performance of information retrieval will have a significant impact on the entire system.
  • Asynchronous MCSLAM: Current MCSLAM studies are based on synchronous cameras. However, in practical applications, the natural asynchrony of cameras can affect system performance, leading to accuracy impacts.
  • MCSLAM in dynamic environments (challenging scenarios): SLAM’s development has been following the development of the demand for autonomous driving technology. With the gradual improvement in and popularization of autonomous driving technology, the demands and requirements for unmanned systems in dynamic environments have gradually increased. Most traditional SLAM mostly solves problems in static environments. Addressing accurate self-positioning and collision avoidance in dynamic environments are hot topics in current MCSLAM research.
  • Establish and expand dense maps: Dense maps can provide richer environmental information for unmanned systems. The breakthrough in dense map construction for the fusion of multiple sensors, drone cluster systems, and ground air joint navigation is of great significance.

8. Conclusions

This paper presents a comprehensive survey of contemporary multi-camera SLAM (MCSLAM) algorithms. Building upon the general framework of visual SLAM, we systematically classify existing MCSLAM approaches from multiple perspectives, including multi-camera sensing configurations, calibration and initialization strategies, feature extraction, depth and pose estimation, mapping, deep learning integration, multi-modal sensor fusion, and datasets. The advantages, limitations, and open challenges of representative MCSLAM methods are critically reviewed. Furthermore, we discuss current development trends and identify several open research issues for future investigation. By enabling an expanded field of view, multi-camera systems significantly reduce perceptual blind spots, enhance environmental awareness, and improve localization robustness, which directly contributes to higher system reliability and more effective collision avoidance in complex and dynamic environments. Finally, with recent advances in photon-based sensing technologies, event-camera-based multi-camera systems are emerging as a promising direction for achieving robust SLAM under challenging illumination conditions and for supporting high-speed or ultra-low-light applications. This survey aims to provide valuable insights and references for researchers working on MCSLAM and related unmanned system applications.
In addition, the scalability of MCSLAM remains a critical consideration when extending these systems to heterogeneous robotic swarms or large-scale networked platforms. As the number of agents and cameras increases, challenges related to computational load, communication bandwidth, and global consistency become more pronounced. Distributed and decentralized MCSLAM frameworks offer a potential path toward scalable deployment, but they also introduce new issues such as inter-agent synchronization, uncertainty propagation, and heterogeneous sensor fusion. Addressing these challenges is essential for enabling reliable and scalable MCSLAM in cooperative multi-robot systems and large autonomous networks.
Furthermore, emerging sensors such as event-based cameras and high-dynamic-range (HDR) imaging systems provide new opportunities for improving robustness under extreme illumination changes, high-speed motion, and low-light conditions. However, most existing MCSLAM frameworks are still designed around conventional frame-based cameras, and their direct extension to heterogeneous and evolving sensor modalities remains non-trivial. Developing sensor-aware and adaptive MCSLAM architectures that can effectively integrate novel sensing technologies while preserving scalability and real-time performance constitutes an important direction for future research.
Sensor characteristics also play an important role in temporal consistency for MCSLAM systems. Rolling-shutter cameras, which are widely used due to their low cost and power efficiency, may introduce motion-induced distortions under fast motion or vibration, leading to degraded feature tracking and pose estimation. In contrast, global-shutter sensors provide temporally consistent image acquisition and are better suited for highly dynamic or high-speed scenarios, albeit with higher hardware cost and power consumption. Understanding and mitigating the impact of shutter mechanisms is therefore an important consideration for achieving reliable temporal consistency in multi-camera SLAM, especially in dynamic environments.
In addition to architectural and sensing considerations, the evaluation of MCSLAM systems lacks fully standardized benchmarks and metrics. Current studies primarily rely on localization accuracy, trajectory error, and map consistency as core evaluation criteria, while system-level factors such as computational cost, scalability, and robustness under real-world disturbances are often reported in a fragmented manner. Moreover, practical failure modes—including calibration drift, insufficient texture, severe occlusion, sensor asynchrony, and partial sensor failure—remain recurring challenges that significantly affect deployment reliability. A more systematic characterization of these failure cases, together with unified evaluation practices, would facilitate fair comparison and accelerate progress in the field. Overall, by synthesizing algorithmic advances, practical constraints, and open challenges, this survey aims to provide a coherent and structured perspective on the current state and future directions of MCSLAM research.

Author Contributions

G.W.: Writing original draft. L.W.: Writing original draft. J.H.: Conceptualization and revision. Y.J.: Writing original draft. Y.Z.: Document organization. Q.Q.: Document organization. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by the National Natural Science Foundation of China under Grant 62303478, 62401465, 62401588.

Data Availability Statement

All data generated or analyzed during this study are included in this published article.

Conflicts of Interest

The authors declare no conflict of interest.

References

  1. Hussain, K.; Oh, I.Y. Joint radar, communication, and integration of beamforming technology. Electronics 2024, 13, 1531. [Google Scholar] [CrossRef]
  2. Yu, J.; Yu, T.H.; Zhang, Q.Y.; Nguyen, T.T. NI-LIO: A Hybrid Approach Combining ICP and NDT for Improving Simultaneous Localization and Mapping Performance. Electronics 2025, 14, 178. [Google Scholar] [CrossRef]
  3. He, Y.; Li, B.; Ruan, J.; Yu, A.; Hou, B. ZUST Campus: A Lightweight and Practical LiDAR SLAM Dataset for Autonomous Driving Scenarios. Electronics 2024, 13, 1341. [Google Scholar] [CrossRef]
  4. Zou, Y.; Meng, P.; Xiong, J.; Wan, X. SS-LIO: Robust Tightly Coupled Solid-State LiDAR–Inertial Odometry for Indoor Degraded Environments. Electronics 2025, 14, 2951. [Google Scholar] [CrossRef]
  5. Shan, D.; Su, J.; Wang, X.; Liu, Y.; Zhou, T.; Wu, Z. VID-SLAM: Robust pose estimation with RGBD-inertial input for indoor robotic localization. Electronics 2024, 13, 318. [Google Scholar] [CrossRef]
  6. Zhou, K.; Yu, Z.; Zhou, X.; Tan, P.; Yin, Y.; Luo, H. ADEmono-SLAM: Absolute Depth Estimation for Monocular Visual Simultaneous Localization and Mapping in Complex Environments. Electronics 2025, 14, 4126. [Google Scholar] [CrossRef]
  7. Rostum, H.; Vásárhelyi, J. Enhancing Machine Learning Techniques in VSLAM for Robust Autonomous Unmanned Aerial Vehicle Navigation. Electronics 2025, 14, 1440. [Google Scholar] [CrossRef]
  8. Yuan, C.; Wang, D.; Li, Z.; Xu, Y.; Zhang, Z. DJPETE-SLAM: Object-Level SLAM System Based on Distributed Joint Pose Estimation and Texture Editing. Electronics 2025, 14, 1181. [Google Scholar] [CrossRef]
  9. Zhou, L.; Wang, M.; Zhang, X.; Qin, P.; He, B. Adaptive slam methodology based on simulated annealing particle swarm optimization for auv navigation. Electronics 2023, 12, 2372. [Google Scholar] [CrossRef]
  10. Wang, C.; Cheng, C.; Yang, D.; Pan, G.; Zhang, F. Underwater auv navigation dataset in natural scenarios. Electronics 2023, 12, 3788. [Google Scholar] [CrossRef]
  11. Zhu, Z.; Wu, F.; Sun, W.; Wu, Q.; Liang, F.; Zhang, W. Depth Estimation Based on MMwave Radar and Camera Fusion with Attention Mechanisms and Multi-Scale Features for Autonomous Driving Vehicles. Electronics 2025, 14, 300. [Google Scholar] [CrossRef]
  12. Meng, Y.; Pan, X.; Wu, M.; Guo, Y.; Liu, Y.; Jiang, J.; Chen, C. VIFPE: Multimodal Fusion of Visible and Infrared Images for Pose Estimation in Large-Scale Urban Environments. Electronics 2025, 14, 3621. [Google Scholar] [CrossRef]
  13. Han, D.; Choi, Y. Monocular depth estimation from a single infrared image. Electronics 2022, 11, 1729. [Google Scholar] [CrossRef]
  14. Xu, W.; Zhang, P.; Yuan, G.; Xu, S.; Li, L.; Zhang, J.; Li, L.; Li, T.; Wang, Z. DCNN–Transformer Hybrid Network for Robust Feature Extraction in FMCW LiDAR Ranging. Photonics 2025, 12, 995. [Google Scholar] [CrossRef]
  15. Zuo, Z.; Wu, Z.; Wei, J.; Wu, P.; Huang, S.; Cheng, Z. A Checkerboard Corner Detection Method for Infrared Thermal Camera Calibration Based on Physics-Informed Neural Network. Photonics 2025, 12, 847. [Google Scholar] [CrossRef]
  16. You, B.; Chen, H.; Li, J.; Li, C.; Chen, H. Fast point cloud registration algorithm based on 3DNPFH descriptor. Photonics 2022, 9, 414. [Google Scholar] [CrossRef]
  17. Chen, Z.S.; Chrisantonius; Raswa, F.H.; Chen, S.K.; Huang, C.I.; Li, K.C.; Chen, S.L.; Li, Y.H.; Wang, J.C. A Hybrid Deep Learning and Feature Descriptor Approach for Partial Fingerprint Recognition. Electronics 2025, 14, 1807. [Google Scholar] [CrossRef]
  18. Zhang, X.; Zhang, Z.; Wang, Q.; Yang, Y. Using a two-stage method to reject false loop closures and improve the accuracy of collaborative SLAM systems. Electronics 2021, 10, 2638. [Google Scholar] [CrossRef]
  19. Chghaf, M.; Rodríguez Flórez, S.; El Ouardi, A. Extended Study of a Multi-Modal Loop Closure Detection Framework for SLAM Applications. Electronics 2025, 14, 421. [Google Scholar] [CrossRef]
  20. Park, J.; Lee, B.; Lee, S.; Song, S. Stereo-GS: Online 3D Gaussian Splatting Mapping Using Stereo Depth Estimation. Electronics 2025, 14, 4436. [Google Scholar] [CrossRef]
  21. Chen, J.; Jiao, Y.; Jin, F.; Qin, X.; Ning, Y.; Yang, M.; Zhan, Y. Plant Sam Gaussian Reconstruction (PSGR): A High-Precision and Accelerated Strategy for Plant 3D Reconstruction. Electronics 2025, 14, 2291. [Google Scholar] [CrossRef]
  22. Kostavelis, I.; Gasteratos, A. Semantic mapping for mobile robotics tasks: A survey. Robot. Auton. Syst. 2015, 66, 86–103. [Google Scholar] [CrossRef]
  23. Yousif, K.; Bab-Hadiashar, A.; Hoseinnezhad, R. An overview to visual odometry and visual SLAM: Applications to mobile robotics. Intell. Ind. Syst. 2015, 1, 289–311. [Google Scholar] [CrossRef]
  24. Lowry, S.; Sünderhauf, N.; Newman, P.; Leonard, J.J.; Cox, D.; Corke, P.; Milford, M.J. Visual place recognition: A survey. IEEE Trans. Robot. 2015, 32, 1–19. [Google Scholar] [CrossRef]
  25. Taketomi, T.; Uchiyama, H.; Ikeda, S. Visual SLAM algorithms: A survey from 2010 to 2016. IPSJ Trans. Comput. Vis. Appl. 2017, 9, 16. [Google Scholar] [CrossRef]
  26. Saputra, M.R.U.; Markham, A.; Trigoni, N. Visual SLAM and structure from motion in dynamic environments: A survey. ACM Comput. Surv. 2018, 51, 1–36. [Google Scholar] [CrossRef]
  27. Jamiruddin, R.; Sari, A.O.; Shabbir, J.; Anwer, T. RGB-depth SLAM review. arXiv 2018, arXiv:1805.07696. [Google Scholar] [CrossRef]
  28. Chen, Y.; Zhou, Y.; Lv, Q.; Deveerasetty, K.K. A review of v-slam. In Proceedings of the 2018 IEEE International Conference on Information and Automation (ICIA), Wuyishan, China, 11–13 August 2018; IEEE: Piscataway, NJ, USA, 2018; pp. 603–608. [Google Scholar]
  29. Duan, C.; Junginger, S.; Huang, J.; Jin, K.; Thurow, K. Deep learning for visual SLAM in transportation robotics: A review. Transp. Saf. Environ. 2019, 1, 177–184. [Google Scholar] [CrossRef]
  30. Garg, S.; Sünderhauf, N.; Dayoub, F.; Morrison, D.; Cosgun, A.; Carneiro, G.; Wu, Q.; Chin, T.J.; Reid, I.; Gould, S.; et al. Semantics for robotic mapping, perception and interaction: A survey. Found. Trends Robot. 2020, 8, 1–224. [Google Scholar] [CrossRef]
  31. Chen, C.; Wang, B.; Lu, C.X.; Trigoni, N.; Markham, A. A survey on deep learning for localization and mapping: Towards the age of spatial machine intelligence. arXiv 2020, arXiv:2006.12567. [Google Scholar] [CrossRef]
  32. Zeng, R.; Wen, Y.; Zhao, W.; Liu, Y.J. View planning in robot active vision: A survey of systems, algorithms, and applications. Comput. Vis. Media 2020, 6, 225–245. [Google Scholar] [CrossRef]
  33. Xia, L.; Cui, J.; Shen, R.; Xu, X.; Gao, Y.; Li, X. A survey of image semantics-based visual simultaneous localization and mapping: Application-oriented solutions to autonomous navigation of mobile robots. Int. J. Adv. Robot. Syst. 2020, 17, 1729881420919185. [Google Scholar] [CrossRef]
  34. Servières, M.; Renaudin, V.; Dupuis, A.; Antigny, N. Visual and visual-inertial slam: State of the art, classification, and experimental benchmarking. J. Sens. 2021, 2021, 2054828. [Google Scholar] [CrossRef]
  35. Tsintotas, K.A.; Bampis, L.; Gasteratos, A. The revisiting problem in simultaneous localization and mapping: A survey on visual loop closure detection. IEEE Trans. Intell. Transp. Syst. 2022, 23, 19929–19953. [Google Scholar] [CrossRef]
  36. Chen, K.; Zhang, J.; Liu, J.; Tong, Q.; Liu, R.; Chen, S. Semantic Visual Simultaneous Localization and Mapping: A Survey. arXiv 2022, arXiv:2209.06428. [Google Scholar] [CrossRef]
  37. Agostinho, L.R.; Ricardo, N.M.; Pereira, M.I.; Hiolle, A.; Pinto, A.M. A practical survey on visual odometry for autonomous driving in challenging scenarios and conditions. IEEE Access 2022, 10, 72182–72205. [Google Scholar] [CrossRef]
  38. Tosi, F.; Zhang, Y.; Gong, Z.; Sandström, E.; Mattoccia, S.; Oswald, M.R.; Poggi, M. How NeRFs and 3D Gaussian Splatting Are Reshaping SLAM: A Survey. arXiv 2024, arXiv:2402.13255. [Google Scholar] [CrossRef]
  39. Zhuang, L.; Zhong, X.; Xu, L.; Tian, C.; Yu, W. Visual SLAM for Unmanned Aerial Vehicles: Localization and Perception. Sensors 2024, 24, 2980. [Google Scholar] [CrossRef]
  40. Wang, Y.; Tian, Y.; Chen, J.; Xu, K.; Ding, X. A Survey of Visual SLAM in Dynamic Environment: The Evolution From Geometric to Semantic Approaches. IEEE Trans. Instrum. Meas. 2024, 73, 1–21. [Google Scholar] [CrossRef]
  41. Chong, T.; Tang, X.; Leng, C.; Yogeswaran, M.; Ng, O.; Chong, Y. Sensor technologies and simultaneous localization and mapping (SLAM). Procedia Comput. Sci. 2015, 76, 174–179. [Google Scholar] [CrossRef]
  42. Cadena, C.; Carlone, L.; Carrillo, H.; Latif, Y.; Scaramuzza, D.; Neira, J.; Reid, I.; Leonard, J.J. Past, present, and future of simultaneous localization and mapping: Toward the robust-perception age. IEEE Trans. Robot. 2016, 32, 1309–1332. [Google Scholar] [CrossRef]
  43. Saeedi, S.; Trentini, M.; Seto, M.; Li, H. Multiple-robot simultaneous localization and mapping: A review. J. Field Robot. 2016, 33, 3–46. [Google Scholar] [CrossRef]
  44. Zaffar, M.; Ehsan, S.; Stolkin, R.; Maier, K.M. Sensors, slam and long-term autonomy: A review. In Proceedings of the 2018 NASA/ESA Conference on Adaptive Hardware and Systems (AHS), Edinburgh, UK, 6–9 August 2018; IEEE: Piscataway, NJ, USA, 2018; pp. 285–290. [Google Scholar]
  45. Sualeh, M.; Kim, G.W. Simultaneous localization and mapping in the epoch of semantics: A survey. Int. J. Control Autom. Syst. 2019, 17, 729–742. [Google Scholar] [CrossRef]
  46. Huang, B.; Zhao, J.; Liu, J. A survey of simultaneous localization and mapping with an envision in 6G wireless networks. arXiv 2019, arXiv:1909.05214. [Google Scholar]
  47. Zhao, W.; He, T.; Sani, A.Y.M.; Yao, T. Review of slam techniques for autonomous underwater vehicles. In Proceedings of the 2019 International Conference on Robotics, Intelligent Control and Artificial Intelligence, Shanghai, China, 20–22 September 2019; ACM: New York, NY, USA, 2019; pp. 384–389. [Google Scholar]
  48. Rosen, D.M.; Doherty, K.J.; Terán Espinoza, A.; Leonard, J.J. Advances in inference and representation for simultaneous localization and mapping. Annu. Rev. Control Robot. Auton. Syst. 2021, 4, 215–242. [Google Scholar] [CrossRef]
  49. Liu, Y.; Fu, Y.; Chen, F.; Goossens, B.; Tao, W.; Zhao, H. Simultaneous localization and mapping related datasets: A comprehensive survey. arXiv 2021, arXiv:2102.04036. [Google Scholar] [CrossRef]
  50. Bavle, H.; Sanchez-Lopez, J.L.; Cimarelli, C.; Tourani, A.; Voos, H. From slam to situational awareness: Challenges and survey. Sensors 2023, 23, 4849. [Google Scholar] [CrossRef]
  51. Eyvazpour, R.; Shoaran, M.; Karimian, G. Hardware implementation of SLAM algorithms: A survey on implementation approaches and platforms. Artif. Intell. Rev. 2023, 56, 6187–6239. [Google Scholar] [CrossRef]
  52. Tezerjani, M.D.; Khoshnazar, M.; Tangestanizadeh, M.; Kiani, A.; Yang, Q. A Survey on Reinforcement Learning Applications in SLAM. arXiv 2024, arXiv:2408.14518. [Google Scholar] [CrossRef]
  53. Debeunne, C.; Vivet, D. A review of visual-LiDAR fusion based simultaneous localization and mapping. Sensors 2020, 20, 2068. [Google Scholar] [CrossRef]
  54. Kolhatkar, C.; Wagle, K. Review of SLAM algorithms for indoor mobile robot with LIDAR and RGB-D camera technology. In Innovations in Electrical and Electronic Engineering; Springer: Singapore, 2021; pp. 397–409. [Google Scholar]
  55. Yang, J.; Li, Y.; Cao, L.; Jiang, Y.; Sun, L.; Xie, Q. A survey of SLAM research based on LiDAR sensors. Int. J. Sens. 2019, 1, 1003. [Google Scholar]
  56. Zhang, Z.; Rebecq, H.; Forster, C.; Scaramuzza, D. Benefit of large field-of-view cameras for visual odometry. In Proceedings of the 2016 IEEE International Conference on Robotics and Automation (ICRA), Stockholm, Sweden, 16–21 May 2016; IEEE: Piscataway, NJ, USA, 2016; pp. 801–808. [Google Scholar]
  57. Zou, D.; Tan, P. Coslam: Collaborative visual slam in dynamic environments. IEEE Trans. Pattern Anal. Mach. Intell. 2012, 35, 354–366. [Google Scholar] [CrossRef] [PubMed]
  58. Li, A.; Zou, D.; Yu, W. Robust initialization of multi-camera slam with limited view overlaps and inaccurate extrinsic calibration. In Proceedings of the 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Prague, Czech Republic, 27 September–1 October 2021; IEEE: Piscataway, NJ, USA, 2021; pp. 3361–3367. [Google Scholar]
  59. Wang, Y.; Ng, Y.; Sa, I.; Parra, A.; Rodriguez-Opazo, C.; Lin, T.; Li, H. MAVIS: Multi-Camera Augmented Visual-Inertial SLAM using SE2(3) Based Exact IMU Pre-integration. In Proceedings of the 2024 IEEE International Conference on Robotics and Automation (ICRA), Yokohama, Japan, 13–17 May 2024; IEEE: Piscataway, NJ, USA, 2024; pp. 1694–1700. [Google Scholar] [CrossRef]
  60. Wang, L.; Xu, Y.; Shen, S. VINS-Multi: A Robust Asynchronous Multi-camera-IMU State Estimator. arXiv 2024, arXiv:2405.14539. [Google Scholar] [CrossRef]
  61. Song, H.; Liu, C.; Dai, H. BundledSLAM: An Accurate Visual SLAM System Using Multiple Cameras. In Proceedings of the 2024 IEEE 7th Advanced Information Technology, Electronic and Automation Control Conference (IAEAC), Chongqing, China, 15–17 March 2024; Volume 7, pp. 106–111. [Google Scholar] [CrossRef]
  62. Yang, Y.; Pan, M.; Tang, D.; Wang, T.; Yue, Y.; Liu, T.; Fu, M. MCOV-SLAM: A Multicamera Omnidirectional Visual SLAM System. IEEE/ASME Trans. Mechatron. 2024, 29, 3556–3567. [Google Scholar]
  63. Sewtz, M.; Fanger, Y.; Luo, X.; Bodenmüller, T.; Triebel, R. IndoorMCD: A Benchmark for Low-Cost Multi-Camera SLAM in Indoor Environments. IEEE Robot. Autom. Lett. 2023, 8, 1707–1714. [Google Scholar] [CrossRef]
  64. Morra, L.; Biondo, A.; Poerio, N.; Lamberti, F. MIXO: Mixture of Experts-based Visual Odometry for Multicamera Autonomous Systems. IEEE Trans. Consum. Electron. 2023, 69, 261–270. [Google Scholar] [CrossRef]
  65. Evangelista, D.; Olivastri, E.; Allegro, D.; Menegatti, E.; Pretto, A. A Graph-based Optimization Framework for Hand-Eye Calibration for Multi-Camera Setups. arXiv 2023, arXiv:2303.04747. [Google Scholar]
  66. Kaveti, P.; Vaidyanathan, S.N.; Chelvan, A.T.; Singh, H. Design and Evaluation of a Generic Visual SLAM Framework for Multi Camera Systems. IEEE Robot. Autom. Lett. 2023, 8, 7368–7375. [Google Scholar] [CrossRef]
  67. Zhang, Z.; Zou, K. MMO-SLAM: A Versatile and Accurate Multi Monocular SLAM System. J. Intell. Robot. Syst. 2022, 105, 58. [Google Scholar] [CrossRef]
  68. Karpyshev, P.; Kruzhkov, E.; Yudin, E.; Savinykh, A.; Potapov, A.; Kurenkov, M.; Kolomeytsev, A.; Kalinov, I.; Tsetserukou, D. Mucaslam: Cnn-based frame quality assessment for mobile robot with omnidirectional visual slam. In Proceedings of the 2022 IEEE 18th International Conference on Automation Science and Engineering (CASE), Mexico City, Mexico, 20–24 August 2022; IEEE: Piscataway, NJ, USA, 2022; pp. 368–373. [Google Scholar]
  69. Yin, J.; Li, A.; Li, T.; Yu, W.; Zou, D. M2dgr: A multi-sensor and multi-scenario slam dataset for ground robots. IEEE Robot. Autom. Lett. 2021, 7, 2266–2273. [Google Scholar]
  70. Javed, Z.; Kim, G.W. PanoVILD: A challenging panoramic vision, inertial and LiDAR dataset for simultaneous localization and mapping. J. Supercomput. 2022, 78, 8247–8267. [Google Scholar] [CrossRef]
  71. Ochoa, E.; Gracias, N.; Istenič, K.; Bosch, J.; Cieślak, P.; García, R. Collision detection and avoidance for underwater vehicles using omnidirectional vision. Sensors 2022, 22, 5354. [Google Scholar] [CrossRef]
  72. Gao, L.; Liang, Y.; Yang, J.; Wu, S.; Wang, C.; Chen, J.; Kneip, L. Vector: A versatile event-centric benchmark for multi-sensor slam. IEEE Robot. Autom. Lett. 2022, 7, 8217–8224. [Google Scholar] [CrossRef]
  73. Yu, J.; Xiang, Z.; Su, J. Hierarchical multi-level information fusion for robust and consistent visual SLAM. IEEE Trans. Veh. Technol. 2021, 71, 250–259. [Google Scholar] [CrossRef]
  74. Zhang, L.; Wisth, D.; Camurri, M.; Fallon, M. Balancing the budget: Feature selection and tracking for multi-camera visual-inertial odometry. IEEE Robot. Autom. Lett. 2021, 7, 1182–1189. [Google Scholar] [CrossRef]
  75. Wang, F.; Zhang, C.; Zhang, G.; Liu, Y.; Xia, Y.; Yang, X. PLMCVIO: Point-Line based Multi-Camera Visual Inertial Odometry. In Proceedings of the 2021 IEEE 11th Annual International Conference on CYBER Technology in Automation, Control, and Intelligent Systems (CYBER), Jiaxing, China, 27–31 July 2021; IEEE: Piscataway, NJ, USA, 2021; pp. 867–872. [Google Scholar]
  76. Eckenhoff, K.; Geneva, P.; Huang, G. MIMC-VINS: A versatile and resilient multi-IMU multi-camera visual-inertial navigation system. IEEE Trans. Robot. 2021, 37, 1360–1380. [Google Scholar] [CrossRef]
  77. Diamanti, E.; Løvås, H.S.; Larsen, M.K.; Ødegård, Ø. A multi-camera system for the integrated documentation of Underwater Cultural Heritage of high structural complexity; The case study of M/S Helma wreck. IFAC-PapersOnLine 2021, 54, 422–429. [Google Scholar] [CrossRef]
  78. He, D.; Chuang, H.M.; Chen, J.; Li, J.; Namiki, A. Real-time visual feedback control of multi-camera UAV. J. Robot. Mechatron. 2021, 33, 263–273. [Google Scholar] [CrossRef]
  79. Ali, I.; Durmush, A.; Suominen, O.; Yli-Hietanen, J.; Peltonen, S.; Collin, J.; Gotchev, A. FinnForest dataset: A forest landscape for visual SLAM. Robot. Auton. Syst. 2020, 132, 103610. [Google Scholar] [CrossRef]
  80. Rofallski, R.; Tholen, C.; Helmholz, P.; Parnum, I.; Luhmann, T. Measuring artificial reefs using a multi-camera system for unmanned underwater vehicles. Int. Arch. Photogramm. Remote Sens. Spat. Inf. Sci. 2020, 43, 999–1008. [Google Scholar] [CrossRef]
  81. Toft, C.; Maddern, W.; Torii, A.; Hammarstrand, L.; Stenborg, E.; Safari, D.; Okutomi, M.; Pollefeys, M.; Sivic, J.; Pajdla, T.; et al. Long-term visual localization revisited. IEEE Trans. Pattern Anal. Mach. Intell. 2020, 44, 2074–2088. [Google Scholar] [CrossRef]
  82. Seok, H.; Lim, J. ROVINS: Robust omnidirectional visual inertial navigation system. IEEE Robot. Autom. Lett. 2020, 5, 6225–6232. [Google Scholar] [CrossRef]
  83. Yang, Y.; Tang, D.; Wang, D.; Song, W.; Wang, J.; Fu, M. Multi-camera visual SLAM for off-road navigation. Robot. Auton. Syst. 2020, 128, 103505. [Google Scholar] [CrossRef]
  84. Li, J.; Deng, G.; Zhang, W.; Zhang, C.; Wang, F.; Liu, Y. Realization of CUDA-based real-time multi-camera visual SLAM in embedded systems. J. Real Time Image Process. 2020, 17, 713–727. [Google Scholar] [CrossRef]
  85. Eckenhoff, K.; Geneva, P.; Bloecker, J.; Huang, G. Multi-camera visual-inertial navigation with online intrinsic and extrinsic calibration. In Proceedings of the 2019 International Conference on Robotics and Automation (ICRA), Montreal, QC, Canada, 20–24 May 2019; IEEE: Piscataway, NJ, USA, 2019; pp. 3158–3164. [Google Scholar]
  86. Meng, X.; Gao, W.; Hu, Z. Dense RGB-D SLAM with multiple cameras. Sensors 2018, 18, 2118. [Google Scholar] [CrossRef]
  87. Rufli, M.; Scaramuzza, D.; Siegwart, R. Automatic detection of checkerboards on blurred and distorted images. In Proceedings of the 2008 IEEE/RSJ International Conference on Intelligent Robots and Systems, Nice, France, 22–26 September 2008; IEEE: Piscataway, NJ, USA, 2008; pp. 3121–3126. [Google Scholar]
  88. Li, B.; Heng, L.; Koser, K.; Pollefeys, M. A multiple-camera system calibration toolbox using a feature descriptor-based calibration pattern. In Proceedings of the 2013 IEEE/RSJ International Conference on Intelligent Robots and Systems, Tokyo, Japan, 3–7 November 2013; IEEE: Piscataway, NJ, USA, 2013; pp. 1301–1307. [Google Scholar]
  89. Kumar, R.K.; Ilie, A.; Frahm, J.M.; Pollefeys, M. Simple calibration of non-overlapping cameras with a mirror. In Proceedings of the 2008 IEEE Conference on Computer Vision and Pattern Recognition, Anchorage, AK, USA, 23–28 June 2008; IEEE: Piscataway, NJ, USA, 2008; pp. 1–7. [Google Scholar]
  90. Pollok, T.; Monari, E. A visual SLAM-based approach for calibration of distributed camera networks. In Proceedings of the 2016 13th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS), Colorado Springs, CO, USA, 23–26 August 2016; IEEE: Piscataway, NJ, USA, 2016; pp. 429–437. [Google Scholar]
  91. Zhang, L.; Zhang, J.; Zhang, W.; Zhang, C.; Liu, Y. An Online Automatic Calibration Method Based on Feature Descriptor for Non-Overlapping Multi-Camera Systems. In Proceedings of the 2018 IEEE International Conference on Robotics and Biomimetics (ROBIO), Kuala Lumpur, Malaysia, 12–15 December 2018; IEEE: Piscataway, NJ, USA, 2018; pp. 138–143. [Google Scholar]
  92. Ouyang, Z.; Hu, L.; Lu, Y.; Wang, Z.; Peng, X.; Kneip, L. Online calibration of exterior orientations of a vehicle-mounted surround-view camera system. In Proceedings of the 2020 IEEE International Conference on Robotics and Automation (ICRA), Paris, France, 31 May–31 August 2020; IEEE: Piscataway, NJ, USA, 2020; pp. 4990–4996. [Google Scholar]
  93. Rehder, J.; Nikolic, J.; Schneider, T.; Hinzmann, T.; Siegwart, R. Extending kalibr: Calibrating the extrinsics of multiple IMUs and of individual axes. In Proceedings of the 2016 IEEE International Conference on Robotics and Automation (ICRA), Stockholm, Sweden, 16–21 May 2016; IEEE: Piscataway, NJ, USA, 2016; pp. 4304–4311. [Google Scholar]
  94. Xu, J.; Li, R.; Zhao, L.; Yu, W.; Liu, Z.; Zhang, B.; Li, Y. CamMap: Extrinsic Calibration of Non-Overlapping Cameras Based on SLAM Map Alignment. IEEE Robot. Autom. Lett. 2022, 7, 11879–11885. [Google Scholar] [CrossRef]
  95. Dexheimer, E.; Peluse, P.; Chen, J.; Pritts, J.; Kaess, M. Information-theoretic online multi-camera extrinsic calibration. IEEE Robot. Autom. Lett. 2022, 7, 4757–4764. [Google Scholar] [CrossRef]
  96. Su, P.C.; Shen, J.; Xu, W.; Cheung, S.C.S.; Luo, Y. A fast and robust extrinsic calibration for RGB-D camera networks. Sensors 2018, 18, 235. [Google Scholar] [CrossRef]
  97. Newcombe, R.A.; Izadi, S.; Hilliges, O.; Molyneaux, D.; Kim, D.; Davison, A.J.; Kohi, P.; Shotton, J.; Hodges, S.; Fitzgibbon, A. Kinectfusion: Real-time dense surface mapping and tracking. In Proceedings of the 2011 10th IEEE International Symposium on Mixed and Augmented Reality, Basel, Switzerland, 26–29 October 2011; IEEE: Piscataway, NJ, USA, 2011; pp. 127–136. [Google Scholar]
  98. Tao, L.; Xia, R.; Zhao, J.; Zhang, T.; Chen, Y.; Fu, S. A convenient and high-accuracy multicamera calibration method based on imperfect spherical objects. IEEE Trans. Instrum. Meas. 2021, 70, 1–15. [Google Scholar] [CrossRef]
  99. Das, A.; Waslander, S.L. Calibration of a dynamic camera cluster for multi-camera visual SLAM. In Proceedings of the 2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Daejeon, Republic of Korea, 9–14 October 2016; IEEE: Piscataway, NJ, USA, 2016; pp. 4637–4642. [Google Scholar]
  100. Rebello, J.; Das, A.; Waslander, S. Autonomous active calibration of a dynamic camera cluster using next-best-view. In Proceedings of the 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Vancouver, BC, Canada, 24–28 September 2017; IEEE: Piscataway, NJ, USA, 2017; pp. 1484–1489. [Google Scholar]
  101. Rebello, J.; Fung, A.; Waslander, S.L. AC/DCC: Accurate Calibration of Dynamic Camera Clusters for Visual SLAM. In Proceedings of the 2020 IEEE International Conference on Robotics and Automation (ICRA), Paris, France, 31 May–31 August 2020; IEEE: Piscataway, NJ, USA, 2020; pp. 6035–6041. [Google Scholar]
  102. Hartenberg, R. Kinematic Synthesis of Linkages; McGraw-Hill, 1964; Volume 2, pp. 198–202. [Google Scholar]
  103. Choi, C.L.; Rebello, J.; Koppel, L.; Ganti, P.; Das, A.; Waslander, S.L. Encoderless gimbal calibration of dynamic multi-camera clusters. In Proceedings of the 2018 IEEE International Conference on Robotics and Automation (ICRA), Brisbane, QLD, Australia, 21–25 May 2018; IEEE: Piscataway, NJ, USA, 2018; pp. 2126–2133. [Google Scholar]
  104. Heng, L.; Li, B.; Pollefeys, M. Camodocal: Automatic intrinsic and extrinsic calibration of a rig with multiple generic cameras and odometry. In Proceedings of the 2013 IEEE/RSJ International Conference on Intelligent Robots and Systems, Tokyo, Japan, 3–7 November 2013; IEEE: Piscataway, NJ, USA, 2013; pp. 1793–1800. [Google Scholar]
  105. Wang, Y.; Jiang, W.; Huang, K.; Schwertfeger, S.; Kneip, L. Accurate calibration of multi-perspective cameras from a generalization of the hand-eye constraint. In Proceedings of the 2022 International Conference on Robotics and Automation (ICRA), Philadelphia, PA, USA, 23–27 May 2022; IEEE: Piscataway, NJ, USA, 2022; pp. 1244–1250. [Google Scholar]
  106. Tsai, R.Y.; Lenz, R.K. Real time versatile robotics hand/eye calibration using 3D machine vision. In Proceedings of the 1988 IEEE International Conference on Robotics and Automation, Philadelphia, PA, USA, 24–29 April 1988; IEEE: Piscataway, NJ, USA, 1988; pp. 554–561. [Google Scholar]
  107. Ou, J.; Huang, P.; Zhou, J.; Zhao, Y.; Lin, L. Automatic Extrinsic Calibration of 3D LIDAR and Multi-Cameras Based on Graph Optimization. Sensors 2022, 22, 2221. [Google Scholar] [CrossRef]
  108. Kim, E.S.; Park, S.Y. Extrinsic calibration between camera and LiDAR sensors by matching multiple 3D planes. Sensors 2019, 20, 52. [Google Scholar] [CrossRef]
  109. Pusztai, Z.; Eichhardt, I.; Hajder, L. Accurate calibration of multi-lidar-multi-camera systems. Sensors 2018, 18, 2139. [Google Scholar] [CrossRef] [PubMed]
  110. Hassanein, M.; Moussa, A.; El-Sheimy, N. A new automatic system calibration of multi-cameras and lidar sensors. Int. Arch. Photogramm. Remote Sens. Spat. Inf. Sci. 2016, 41, 589–594. [Google Scholar] [CrossRef]
  111. Hartzer, J.; Saripalli, S. Online Multi Camera-IMU Calibration. In Proceedings of the 2022 IEEE International Symposium on Safety, Security, and Rescue Robotics (SSRR), Sevilla, Spain, 8–10 November 2022; IEEE: Piscataway, NJ, USA, 2022; pp. 360–365. [Google Scholar]
  112. Brink, K.; Soloviev, A. Filter-based calibration for an IMU and multi-camera system. In Proceedings of the 2012 IEEE/ION Position, Location and Navigation Symposium, Myrtle Beach, SC, USA, 23–26 April 2012; IEEE: Piscataway, NJ, USA, 2012; pp. 730–739. [Google Scholar]
  113. Yang, Z.; Liu, T.; Shen, S. Self-calibrating multi-camera visual-inertial fusion for autonomous MAVs. In Proceedings of the 2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Daejeon, Republic of Korea, 9–14 October 2016; IEEE: Piscataway, NJ, USA, 2016; pp. 4984–4991. [Google Scholar]
  114. Davison, A.J.; Murray, D.W. Mobile robot localisation using active vision. In Proceedings of the Computer Vision—ECCV’98: 5th European Conference on Computer Vision, Freiburg, Germany, 2–6 June 1998; Springer: Berlin/Heidelberg, Germany, 1998; Volume II, pp. 809–825. [Google Scholar]
  115. Miraldo, P.; Dias, T.; Ramalingam, S. A minimal closed-form solution for multi-perspective pose estimation using points and lines. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 474–490. [Google Scholar]
  116. Uehara, K.; Saito, H.; Hara, K. Line-based SLAM considering prior distribution of distance and angle of line features in an urban environment. In Proceedings of the Computer Vision, Imaging and Computer Graphics–Theory and Applications: 12th International Joint Conference, VISIGRAPP 2017, Porto, Portugal, 27 February–1 March 2017; Revised Selected Papers 12. Springer: Berlin/Heidelberg, Germany, 2019; pp. 105–127. [Google Scholar]
  117. Jiao, Y.; Wang, Y.; Ding, X.; Fu, B.; Huang, S.; Xiong, R. 2-Entity random sample consensus for robust visual localization: Framework, methods, and verifications. IEEE Trans. Ind. Electron. 2020, 68, 4519–4528. [Google Scholar] [CrossRef]
  118. Geppert, M.; Liu, P.; Cui, Z.; Pollefeys, M.; Sattler, T. Efficient 2D-3D matching for multi-camera visual localization. In Proceedings of the 2019 International Conference on Robotics and Automation (ICRA), Montreal, QC, Canada, 20–24 May 2019; IEEE: Piscataway, NJ, USA, 2019; pp. 5972–5978. [Google Scholar]
  119. Zhang, Y.; Huang, F. Panoramic visual slam technology for spherical images. Sensors 2021, 21, 705. [Google Scholar] [CrossRef]
  120. Rublee, E.; Rabaud, V.; Konolige, K.; Bradski, G. ORB: An efficient alternative to SIFT or SURF. In Proceedings of the 2011 International Conference on Computer Vision, Barcelona, Spain, 6–13 November 2011; IEEE: Piscataway, NJ, USA, 2011; pp. 2564–2571. [Google Scholar] [CrossRef]
  121. Lowe, D.G. Distinctive Image Features from Scale-Invariant Keypoints. Int. J. Comput. Vis. 2004, 60, 91–110. [Google Scholar] [CrossRef]
  122. Kang, Y.; Song, Y.; Ge, W.; Ling, T. Robust Multi-camera SLAM with Manhattan Constraint toward Automated Valet Parking. In Proceedings of the 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Prague, Czech Republic, 27 September–1 October 2021; IEEE: Piscataway, NJ, USA, 2021; pp. 7615–7622. [Google Scholar]
  123. Won, C.; Seok, H.; Cui, Z.; Pollefeys, M.; Lim, J. OmniSLAM: Omnidirectional localization and dense mapping for wide-baseline multi-camera systems. In Proceedings of the 2020 IEEE International Conference on Robotics and Automation (ICRA), Paris, France, 31 May–31 August 2020; IEEE: Piscataway, NJ, USA, 2020; pp. 559–566. [Google Scholar]
  124. Yang, F.; Liming, Z.; Yi, Z.; Hengyang, K. Multi-camera System Depth Estimation. In Proceedings of the 2022 IEEE 6th Information Technology and Mechatronics Engineering Conference (ITOEC), Chongqing, China, 4–6 March 2022; IEEE: Piscataway, NJ, USA, 2022; Volume 6, pp. 1202–1207. [Google Scholar]
  125. Ghosh, S.; Gallego, G. Multi-Event-Camera Depth Estimation and Outlier Rejection by Refocused Events Fusion. Adv. Intell. Syst. 2022, 4, 2200221. [Google Scholar] [CrossRef]
  126. Rebecq, H.; Horstschäfer, T.; Gallego, G.; Scaramuzza, D. Evo: A geometric approach to event-based 6-dof parallel tracking and mapping in real time. IEEE Robot. Autom. Lett. 2016, 2, 593–600. [Google Scholar] [CrossRef]
  127. Rebecq, H.; Gallego, G.; Mueggler, E.; Scaramuzza, D. EMVS: Event-based multi-view stereo—3D reconstruction with an event camera in real-time. Int. J. Comput. Vis. 2018, 126, 1394–1414. [Google Scholar] [CrossRef]
  128. Wei, Y.; Zhao, L.; Zheng, W.; Zhu, Z.; Rao, Y.; Huang, G.; Lu, J.; Zhou, J. SurroundDepth: Entangling Surrounding Views for Self-Supervised Multi-Camera Depth Estimation. In Proceedings of the 6th Conference on Robot Learning (CoRL 2022), Auckland, New Zealand, 14–18 December 2022; PMLR: Cambridge, MA, USA, 2022. [Google Scholar] [CrossRef]
  129. Kangni, F.; Laganiere, R. Orientation and pose recovery from spherical panoramas. In Proceedings of the 2007 IEEE 11th International Conference on Computer Vision, Rio de Janeiro, Brazil, 14–21 October 2007; IEEE: Piscataway, NJ, USA, 2007; pp. 1–8. [Google Scholar]
  130. Rituerto, A.; Puig, L.; Guerrero, J.J. Visual SLAM with an omnidirectional camera. In Proceedings of the 2010 20th International Conference on Pattern Recognition, Istanbul, Turkey, 23–26 August 2010; IEEE: Piscataway, NJ, USA, 2010; pp. 348–351. [Google Scholar]
  131. Valiente, D.; Jadidi, M.G.; Miró, J.V.; Gil, A.; Reinoso, O. Information-based view initialization in visual SLAM with a single omnidirectional camera. Robot. Auton. Syst. 2015, 72, 93–104. [Google Scholar] [CrossRef]
  132. Valiente, D.; Gil, A.; Payá, L.; Sebastián, J.M.; Reinoso, Ó. Robust visual localization with dynamic uncertainty management in omnidirectional SLAM. Appl. Sci. 2017, 7, 1294. [Google Scholar] [CrossRef]
  133. Valiente, D.; Gil, A.; Reinoso, Ó.; Juliá, M.; Holloway, M. Improved omnidirectional odometry for a view-based mapping approach. Sensors 2017, 17, 325. [Google Scholar] [CrossRef] [PubMed]
  134. Gamallo, C.; Mucientes, M.; Regueiro, C.V. Omnidirectional visual SLAM under severe occlusions. Robot. Auton. Syst. 2015, 65, 76–87. [Google Scholar] [CrossRef]
  135. Caruso, D.; Engel, J.; Cremers, D. Large-scale direct SLAM for omnidirectional cameras. In Proceedings of the 2015 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Hamburg, Germany, 28 September–2 October 2015; IEEE: Piscataway, NJ, USA, 2015; pp. 141–148. [Google Scholar]
  136. Liu, P.; Heng, L.; Sattler, T.; Geiger, A.; Pollefeys, M. Direct visual odometry for a fisheye-stereo camera. In Proceedings of the 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Vancouver, BC, Canada, 24–28 September 2017; IEEE: Piscataway, NJ, USA, 2017; pp. 1746–1752. [Google Scholar]
  137. Matsuki, H.; Von Stumberg, L.; Usenko, V.; Stückler, J.; Cremers, D. Omnidirectional DSO: Direct sparse odometry with fisheye cameras. IEEE Robot. Autom. Lett. 2018, 3, 3693–3700. [Google Scholar] [CrossRef]
  138. Engel, J.; Koltun, V.; Cremers, D. Direct Sparse Odometry. IEEE Trans. Pattern Anal. Mach. Intell. 2018, 40, 611–625. [Google Scholar] [CrossRef]
  139. Forster, C.; Zhang, Z.; Gassner, M.; Werlberger, M.; Scaramuzza, D. SVO: Semidirect visual odometry for monocular and multicamera systems. IEEE Trans. Robot. 2016, 33, 249–265. [Google Scholar] [CrossRef]
  140. Heng, L.; Choi, B. Semi-direct visual odometry for a fisheye-stereo camera. In Proceedings of the 2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Daejeon, Republic of Korea, 9–14 October 2016; IEEE: Piscataway, NJ, USA, 2016; pp. 4077–4084. [Google Scholar]
  141. Sumikura, S.; Shibuya, M.; Sakurada, K. OpenVSLAM: A versatile visual SLAM framework. In Proceedings of the 27th ACM International Conference on Multimedia, Nice, France, 21–25 October 2019; ACM: New York, NY, USA, 2019; pp. 2292–2295. [Google Scholar]
  142. Pless, R. Using many cameras as one. In Proceedings of the 2003 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Madison, WI, USA, 18–20 June 2003; IEEE: Piscataway, NJ, USA, 2003; Volume 2, p. II-587. [Google Scholar]
  143. Lee, G.H.; Li, B.; Pollefeys, M.; Fraundorfer, F. Minimal solutions for the multi-camera pose estimation problem. Int. J. Robot. Res. 2015, 34, 837–848. [Google Scholar] [CrossRef]
  144. Chen, C.S.; Chang, W.Y. Pose estimation for generalized imaging device via solving non-perspective n point problem. In Proceedings of the 2002 IEEE International Conference on Robotics and Automation (Cat. No. 02CH37292), Washington, DC, USA, 11–15 May 2002; IEEE: Piscataway, NJ, USA, 2002; Volume 3, pp. 2931–2937. [Google Scholar]
  145. Nistér, D.; Stewénius, H. A minimal solution to the generalised 3-point pose problem. J. Math. Imaging Vis. 2007, 27, 67–79. [Google Scholar] [CrossRef]
  146. Kneip, L.; Scaramuzza, D.; Siegwart, R. A novel parametrization of the perspective-three-point problem for a direct computation of absolute camera position and orientation. In Proceedings of the CVPR 2011, Colorado Springs, CO, USA, 20–25 June 2011; IEEE: Piscataway, NJ, USA, 2011; pp. 2969–2976. [Google Scholar]
  147. Hesch, J.A.; Roumeliotis, S.I. A direct least-squares (DLS) method for PnP. In Proceedings of the 2011 International Conference on Computer Vision, Barcelona, Spain, 6–13 November 2011; IEEE: Piscataway, NJ, USA, 2011; pp. 383–390. [Google Scholar]
  148. Kukelova, Z.; Bujnak, M.; Pajdla, T. Automatic generator of minimal problem solvers. In Proceedings of the Computer Vision–ECCV 2008: 10th European Conference on Computer Vision, Marseille, France, 12–18 October 2008; Part III 10. Springer: Berlin/Heidelberg, Germany, 2008; pp. 302–315. [Google Scholar]
  149. Larsson, V.; Oskarsson, M.; Astrom, K.; Wallis, A.; Kukelova, Z.; Pajdla, T. Beyond grobner bases: Basis selection for minimal solvers. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–22 June 2018; pp. 3945–3954. [Google Scholar]
  150. Fragoso, V.; DeGol, J.; Hua, G. gdls*: Generalized pose-and-scale estimation given scale and gravity priors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 14–19 June 2020; pp. 2210–2219. [Google Scholar]
  151. Ross, J.; Mendez, O.; Saha, A.; Johnson, M.; Bowden, R. BEV-SLAM: Building a Globally-Consistent World Map Using Monocular Vision. In Proceedings of the 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Kyoto, Japan, 23–27 October 2022; IEEE: Piscataway, NJ, USA, 2022; pp. 3830–3836. [Google Scholar]
  152. Ben Hazem, Z.; Guler, N.; Altaif, A.H. Model-free trajectory tracking control of a 5-DOF mitsubishi robotic arm using deep deterministic policy gradient algorithm. Discov. Robot. 2025, 1, 4. [Google Scholar] [CrossRef]
  153. Labbé, Y.; Carpentier, J.; Aubry, M.; Sivic, J. Cosypose: Consistent multi-view multi-object 6d pose estimation. In Proceedings of the Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, 23–28 August 2020; Part XVII 16. Springer: Berlin/Heidelberg, Germany, 2020; pp. 574–591. [Google Scholar]
  154. Kaskman, R.; Shugurov, I.; Zakharov, S.; Ilic, S. 6 dof pose estimation of textureless objects from multiple rgb frames. In Proceedings of the Computer Vision–ECCV 2020 Workshops, Glasgow, UK, 23–28 August 2020; Part II 16. Springer: Berlin/Heidelberg, Germany, 2020; pp. 612–630. [Google Scholar]
  155. Li, C.; Bai, J.; Hager, G.D. A unified framework for multi-view multi-class object pose estimation. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 254–269. [Google Scholar]
  156. Sock, J.; Hamidreza Kasaei, S.; Seabra Lopes, L.; Kim, T.K. Multi-view 6D object pose estimation and camera motion planning using RGBD images. In Proceedings of the IEEE International Conference on Computer Vision Workshops, Venice, Italy, 22–29 October 2017; pp. 2228–2235. [Google Scholar]
  157. Erkent, Ö.; Shukla, D.; Piater, J. Integration of probabilistic pose estimates from multiple views. In Proceedings of the Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, 11–14 October 2016; Part VII 14. Springer: Berlin/Heidelberg, Germany, 2016; pp. 154–170. [Google Scholar]
  158. Sarlin, P.E.; DeTone, D.; Malisiewicz, T.; Rabinovich, A. Superglue: Learning feature matching with graph neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 14–19 June 2020; pp. 4938–4947. [Google Scholar]
  159. Yang, S.; Scherer, S. Cubeslam: Monocular 3-d object slam. IEEE Trans. Robot. 2019, 35, 925–938. [Google Scholar] [CrossRef]
  160. Li, F.; Shugurov, I.; Busam, B.; Yang, S.; Ilic, S. Ws-ope: Weakly supervised 6-d object pose regression using relative multi-camera pose constraints. IEEE Robot. Autom. Lett. 2022, 7, 3703–3710. [Google Scholar] [CrossRef]
  161. Kaygusuz, N.; Mendez, O.; Bowden, R. Multi-camera sensor fusion for visual odometry using deep uncertainty estimation. In Proceedings of the 2021 IEEE International Intelligent Transportation Systems Conference (ITSC), Indianapolis, IN, USA, 19–22 September 2021; IEEE: Piscataway, NJ, USA, 2021; pp. 2944–2949. [Google Scholar]
  162. Oskiper, T.; Zhu, Z.; Samarasekera, S.; Kumar, R. Visual odometry system using multiple stereo cameras and inertial measurement unit. In Proceedings of the 2007 IEEE Conference on Computer Vision and Pattern Recognition, Minneapolis, MN, USA, 17–22 June 2007; IEEE: Piscataway, NJ, USA, 2007; pp. 1–8. [Google Scholar]
  163. Hee Lee, G.; Pollefeys, M.; Fraundorfer, F. Relative pose estimation for a multi-camera system with known vertical direction. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA, 23–28 June 2014; pp. 540–547. [Google Scholar]
  164. Jaekel, J.; Kaess, M. Robust multi-stereo visual-inertial odometry. In Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). Workshop on Visual-Inertial Navigation: Challenges and Applications, Macau, China, 8 November 2019. [Google Scholar]
  165. Ye, W.; Zheng, R.; Zhang, F.; Ouyang, Z.; Liu, Y. Robust and efficient vehicles motion estimation with low-cost multi-camera and odometer-gyroscope. In Proceedings of the 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Macau, China, 3–8 November 2019; IEEE: Piscataway, NJ, USA, 2019; pp. 4490–4496. [Google Scholar]
  166. Tschopp, F.; Riner, M.; Fehr, M.; Bernreiter, L.; Furrer, F.; Novkovic, T.; Pfrunder, A.; Cadena, C.; Siegwart, R.; Nieto, J. Versavis—An open versatile multi-camera visual-inertial sensor suite. Sensors 2020, 20, 1439. [Google Scholar] [CrossRef]
  167. Miiller, M.; Steidle, F.; Schuster, M.J.; Lutz, P.; Maier, M.; Stoneman, S.; Tomic, T.; Stürzl, W. Robust visual-inertial state estimation with multiple odometries and efficient mapping on an MAV with ultra-wide FOV stereo vision. In Proceedings of the 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Madrid, Spain, 1–5 October 2018; IEEE: Piscataway, NJ, USA, 2018; pp. 3701–3708. [Google Scholar]
  168. Mur-Artal, R.; Tardós, J.D. ORB-SLAM2: An Open-Source SLAM System for Monocular, Stereo and RGB-D Cameras. IEEE Trans. Robot. 2017, 33, 1255–1262. [Google Scholar] [CrossRef]
  169. Shen, B.; Chen, Y.; Han, F.; Dai, S.; Xiong, R.; Wang, Y. Emv-lio: An efficient multiple vision aided lidar-inertial odometry. arXiv 2023, arXiv:2302.00216. [Google Scholar]
  170. Kaess, M.; Dellaert, F. Probabilistic structure matching for visual SLAM with a multi-camera rig. Comput. Vis. Image Underst. 2010, 114, 286–296. [Google Scholar] [CrossRef]
  171. Harmat, A.; Sharf, I.; Trentini, M. Parallel tracking and mapping with multiple cameras on an unmanned aerial vehicle. In Proceedings of the Intelligent Robotics and Applications: 5th International Conference, ICIRA 2012, Montreal, QC, Canada, 3–5 October 2012; Part I 5. Springer: Berlin/Heidelberg, Germany, 2012; pp. 421–432. [Google Scholar]
  172. Harmat, A.; Trentini, M.; Sharf, I. Multi-camera tracking and mapping for unmanned aerial vehicles in unstructured environments. J. Intell. Robot. Syst. 2015, 78, 291–317. [Google Scholar] [CrossRef]
  173. Kuo, J.; Muglikar, M.; Zhang, Z.; Scaramuzza, D. Redesigning SLAM for arbitrary multi-camera systems. In Proceedings of the 2020 IEEE International Conference on Robotics and Automation (ICRA), Paris, France, 31 May–31 August 2020; IEEE: Piscataway, NJ, USA, 2020; pp. 2116–2122. [Google Scholar]
  174. Urban, S.; Hinz, S. Multicol-slam-a modular real-time multi-camera slam system. arXiv 2016, arXiv:1610.07336. [Google Scholar]
  175. Mur-Artal, R.; Montiel, J.M.M.; Tardós, J.D. ORB-SLAM: A Versatile and Accurate Monocular SLAM System. IEEE Trans. Robot. 2015, 31, 1147–1163. [Google Scholar] [CrossRef]
  176. Yang, A.J.; Cui, C.; Bârsan, I.A.; Urtasun, R.; Wang, S. Asynchronous multi-view SLAM. In Proceedings of the 2021 IEEE International Conference on Robotics and Automation (ICRA), Xi’an, China, 30 May–5 June 2021; IEEE: Piscataway, NJ, USA, 2021; pp. 5669–5676. [Google Scholar]
  177. Olfati-Saber, R. Distributed Kalman filtering for sensor networks. In Proceedings of the 2007 46th IEEE Conference on Decision and Control, New Orleans, LA, USA, 12–14 December 2007; IEEE: Piscataway, NJ, USA, 2007; pp. 5492–5498. [Google Scholar]
  178. Talebi, S.P.; Werner, S. Distributed Kalman filtering and control through embedded average consensus information fusion. IEEE Trans. Autom. Control 2019, 64, 4396–4403. [Google Scholar] [CrossRef]
  179. Talebi, S.P.; Mandic, D.P. On the dynamics of multiagent nonlinear filtering and learning. In Proceedings of the 2024 IEEE 34th International Workshop on Machine Learning for Signal Processing (MLSP), London, UK, 22–25 September 2024; IEEE: Piscataway, NJ, USA, 2024; pp. 1–6. [Google Scholar]
  180. Mahmoud, M.S.; Khalid, H.M. Distributed Kalman filtering: A bibliographic review. IET Control Theory Appl. 2013, 7, 483–501. [Google Scholar] [CrossRef]
  181. Smith, M.; Baldwin, I.; Churchill, W.; Paul, R.; Newman, P. The new college vision and laser data set. Int. J. Robot. Res. 2009, 28, 595–599. [Google Scholar] [CrossRef]
  182. Ceriani, S.; Fontana, G.; Giusti, A.; Marzorati, D.; Matteucci, M.; Migliore, D.; Rizzi, D.; Sorrenti, D.G.; Taddei, P. Rawseeds ground truth collection systems for indoor self-localization and mapping. Auton. Robot. 2009, 27, 353–371. [Google Scholar] [CrossRef]
  183. Geiger, A.; Lenz, P.; Urtasun, R. Are we ready for autonomous driving? the kitti vision benchmark suite. In Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA, 16–21 June 2012; IEEE: Piscataway, NJ, USA, 2012; pp. 3354–3361. [Google Scholar]
  184. Korrapati, H.; Courbon, J.; Alizon, S.; Marmoiton, F. The Institut Pascal data sets. Proc. J. Francoph. Jeunes Cherch. Vis. Ordinat. 2013. [Google Scholar]
  185. Fallon, M.; Johannsson, H.; Kaess, M.; Leonard, J.J. The mit stata center dataset. Int. J. Robot. Res. 2013, 32, 1695–1699. [Google Scholar] [CrossRef]
  186. Koschorrek, P.; Piccini, T.; Oberg, P.; Felsberg, M.; Nielsen, L.; Mester, R. A multi-sensor traffic scene dataset with omnidirectional video. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Portland, OR, USA, 23–28 June 2013; pp. 727–734. [Google Scholar]
  187. Carlevaris-Bianco, N.; Ushani, A.K.; Eustice, R.M. University of Michigan North Campus long-term vision and lidar dataset. Int. J. Robot. Res. 2016, 35, 1023–1035. [Google Scholar] [CrossRef]
  188. Wang, S.; Bai, M.; Mattyus, G.; Chu, H.; Luo, W.; Yang, B.; Liang, J.; Cheverie, J.; Fidler, S.; Urtasun, R. Torontocity: Seeing the world with a million eyes. arXiv 2016, arXiv:1612.00423. [Google Scholar] [CrossRef]
  189. Urban, S.; Jutzi, B. LaFiDa—A laserscanner multi-fisheye camera dataset. J. Imaging 2017, 3, 5. [Google Scholar] [CrossRef]
  190. Pfrommer, B.; Sanket, N.; Daniilidis, K.; Cleveland, J. Penncosyvio: A challenging visual inertial odometry benchmark. In Proceedings of the 2017 IEEE International Conference on Robotics and Automation (ICRA), Singapore, 29 May–3 June 2017; IEEE: Piscataway, NJ, USA, 2017; pp. 3847–3854. [Google Scholar]
  191. Maddern, W.; Pascoe, G.; Linegar, C.; Newman, P. 1 year, 1000 km: The oxford robotcar dataset. Int. J. Robot. Res. 2017, 36, 3–15. [Google Scholar] [CrossRef]
  192. Huang, X.; Cheng, X.; Geng, Q.; Cao, B.; Zhou, D.; Wang, P.; Lin, Y.; Yang, R. The apolloscape dataset for autonomous driving. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Salt Lake City, UT, USA, 18–22 June 2018; pp. 954–960. [Google Scholar]
  193. Zhu, A.Z.; Thakur, D.; Özaslan, T.; Pfrommer, B.; Kumar, V.; Daniilidis, K. The multivehicle stereo event camera dataset: An event camera dataset for 3D perception. IEEE Robot. Autom. Lett. 2018, 3, 2032–2039. [Google Scholar] [CrossRef]
  194. Delmerico, J.; Cieslewski, T.; Rebecq, H.; Faessler, M.; Scaramuzza, D. Are we ready for autonomous drone racing? the UZH-FPV drone racing dataset. In Proceedings of the 2019 International Conference on Robotics and Automation (ICRA), Montreal, QC, Canada, 20–24 May 2019; IEEE: Piscataway, NJ, USA, 2019; pp. 6713–6719. [Google Scholar]
  195. Yogamani, S.; Hughes, C.; Horgan, J.; Sistu, G.; Varley, P.; O’Dea, D.; Uricár, M.; Milz, S.; Simon, M.; Amende, K.; et al. Woodscape: A multi-task, multi-camera fisheye dataset for autonomous driving. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27–28 October 2019; pp. 9308–9318. [Google Scholar]
  196. Chang, M.F.; Lambert, J.; Sangkloy, P.; Singh, J.; Bak, S.; Hartnett, A.; Wang, D.; Carr, P.; Lucey, S.; Ramanan, D.; et al. Argoverse: 3d tracking and forecasting with rich maps. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019; pp. 8748–8757. [Google Scholar]
  197. Shi, X.; Li, D.; Zhao, P.; Tian, Q.; Tian, Y.; Long, Q.; Zhu, C.; Song, J.; Qiao, F.; Song, L.; et al. Are we ready for service robots? the openloris-scene datasets for lifelong slam. In Proceedings of the 2020 IEEE International Conference on Robotics and Automation (ICRA), Paris, France, 31 May–31 August 2020; IEEE: Piscataway, NJ, USA, 2020; pp. 3139–3145. [Google Scholar]
  198. Wen, W.; Zhou, Y.; Zhang, G.; Fahandezh-Saadi, S.; Bai, X.; Zhan, W.; Tomizuka, M.; Hsu, L.T. UrbanLoco: A full sensor suite dataset for mapping and localization in urban scenes. In Proceedings of the 2020 IEEE International Conference on Robotics and Automation (ICRA), Paris, France, 31 May–31 August 2020; IEEE: Piscataway, NJ, USA, 2020; pp. 2310–2316. [Google Scholar]
  199. Ligocki, A.; Jelinek, A.; Zalud, L. Brno urban dataset-the new data for self-driving agents and mapping tasks. In Proceedings of the 2020 IEEE International Conference on Robotics and Automation (ICRA), Paris, France, 31 May–31 August 2020; IEEE: Piscataway, NJ, USA, 2020; pp. 3284–3290. [Google Scholar]
  200. Zuñiga-Noël, D.; Jaenal, A.; Gomez-Ojeda, R.; Gonzalez-Jimenez, J. The UMA-VI dataset: Visual–inertial odometry in low-textured and dynamic illumination environments. Int. J. Robot. Res. 2020, 39, 1052–1060. [Google Scholar] [CrossRef]
  201. Caesar, H.; Bankiti, V.; Lang, A.H.; Vora, S.; Liong, V.E.; Xu, Q.; Krishnan, A.; Pan, Y.; Baldan, G.; Beijbom, O. nuscenes: A multimodal dataset for autonomous driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 14–19 June 2020; pp. 11621–11631. [Google Scholar]
  202. Benseddik, H.E.; Morbidi, F.; Caron, G. PanoraMIS: An ultra-wide field of view image dataset for vision-based robot-motion estimation. Int. J. Robot. Res. 2020, 39, 1037–1051. [Google Scholar] [CrossRef]
  203. Pitropov, M.; Garcia, D.E.; Rebello, J.; Smart, M.; Wang, C.; Czarnecki, K.; Waslander, S. Canadian adverse driving conditions dataset. Int. J. Robot. Res. 2021, 40, 681–690. [Google Scholar] [CrossRef]
  204. Sun, P.; Kretzschmar, H.; Dotiwalla, X.; Chouard, A.; Patnaik, V.; Tsui, P.; Guo, J.; Zhou, Y.; Chai, Y.; Caine, B.; et al. Scalability in perception for autonomous driving: Waymo open dataset. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 14–19 June 2020; pp. 2446–2454. [Google Scholar]
  205. Geyer, J.; Kassahun, Y.; Mahmudi, M.; Ricou, X.; Durgesh, R.; Chung, A.S.; Hauswald, L.; Pham, V.H.; Mühlegg, M.; Dorn, S.; et al. A2d2: Audi autonomous driving dataset. arXiv 2020, arXiv:2004.06320. [Google Scholar] [CrossRef]
  206. Yan, Z.; Sun, L.; Krajník, T.; Ruichek, Y. EU long-term dataset with multiple sensors for autonomous driving. In Proceedings of the 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Las Vegas, NV, USA, 24 October–24 January 2021; IEEE: Piscataway, NJ, USA, 2020; pp. 10697–10704. [Google Scholar]
  207. Zhang, L.; Camurri, M.; Wisth, D.; Fallon, M. Multi-camera lidar inertial extension to the newer college dataset. arXiv 2021, arXiv:2112.08854. [Google Scholar]
  208. Gehrig, M.; Aarents, W.; Gehrig, D.; Scaramuzza, D. Dsec: A stereo event camera dataset for driving scenarios. IEEE Robot. Autom. Lett. 2021, 6, 4947–4954. [Google Scholar] [CrossRef]
  209. Klenk, S.; Chui, J.; Demmel, N.; Cremers, D. Tum-vie: The tum stereo visual-inertial event dataset. In Proceedings of the 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Prague, Czech Republic, 27 September–1 October 2021; IEEE: Piscataway, NJ, USA, 2021; pp. 8601–8608. [Google Scholar]
  210. Pandey, G.; McBride, J.R.; Eustice, R.M. Ford campus vision and lidar data set. Int. J. Robot. Res. 2011, 30, 1543–1552. [Google Scholar] [CrossRef]
  211. Helmberger, M.; Morin, K.; Berner, B.; Kumar, N.; Cioffi, G.; Scaramuzza, D. The hilti slam challenge dataset. IEEE Robot. Autom. Lett. 2022, 7, 7518–7525. [Google Scholar] [CrossRef]
  212. Wang, Z.; Liu, Y.; Duan, Y.; Li, X.; Zhang, X.; Ji, J.; Dong, E.; Zhang, Y. USTC FLICAR: A sensors fusion dataset of LiDAR-inertial-camera for heavy-duty autonomous aerial work robots. Int. J. Robot. Res. 2023, 42, 1015–1047. [Google Scholar] [CrossRef]
  213. Hess, W.; Kohler, D.; Rapp, H.; Andor, D. Real-time loop closure in 2D LIDAR SLAM. In Proceedings of the 2016 IEEE International Conference on Robotics and Automation (ICRA), Stockholm, Sweden, 16–21 May 2016; IEEE: Piscataway, NJ, USA, 2016; pp. 1271–1278. [Google Scholar]
  214. Zhang, J.; Singh, S. LOAM: Lidar odometry and mapping in real-time. In Proceedings of the Robotics: Science and Systems, Berkeley, CA, USA, 12–16 July 2014; Volume 2, pp. 1–9. [Google Scholar]
  215. Zujevs, A.; Pudzs, M.; Osadcuks, V.; Ardavs, A.; Galauskis, M.; Grundspenkis, J. An event-based vision dataset for visual navigation tasks in agricultural environments. In Proceedings of the 2021 IEEE International Conference on Robotics and Automation (ICRA), Xi’an, China, 30 May–5 June 2021; IEEE: Piscataway, NJ, USA, 2021; pp. 13769–13775. [Google Scholar]
Figure 1. The proportion of different types of SLAM surveys since 2015.
Figure 1. The proportion of different types of SLAM surveys since 2015.
Electronics 15 00602 g001
Figure 2. A comparison of visual sensors and environment on MCSLAM since 2018.
Figure 2. A comparison of visual sensors and environment on MCSLAM since 2018.
Electronics 15 00602 g002
Figure 3. Fundamental flowchart of visual SLAM.
Figure 3. Fundamental flowchart of visual SLAM.
Electronics 15 00602 g003
Figure 5. High-level schematic of the generalized MCSLAM pipeline, showing multi-camera input, feature extraction, data association, bundle adjustment, map update, and visualization modules.
Figure 5. High-level schematic of the generalized MCSLAM pipeline, showing multi-camera input, feature extraction, data association, bundle adjustment, map update, and visualization modules.
Electronics 15 00602 g005
Table 2. The Parameters of MCSs Used for SLAM (P: Pinhole Camera, F: Fisheye Camera, C: ACES, A: Arbitrary Configuration.) (Env.: Environment, I: Indoor, O: Outdoor).
Table 2. The Parameters of MCSs Used for SLAM (P: Pinhole Camera, F: Fisheye Camera, C: ACES, A: Arbitrary Configuration.) (Env.: Environment, I: Indoor, O: Outdoor).
Vehicle PlatformMCS TypeCamera QuantityParametersEnv.Year
Fixed bracket [59]P&A5 × cameras720 × 540@30HzI&O2024
Drone [60]P&A3 × RGB-D cameras
1 × stereo camera
1024 × 768@30Hz
1920 × 1080@30Hz
I2024
Fixed bracket [61]P&A2 × cameras752 × 480@20HzI&O2024
Car [62]F&A6 × fisheye cameras752 × 480@15HzO2024
Handheld ground robot [63]P&A3 × RGB-D cameras640 × 480@15HzI2023
Car [64]P&A6 × cameras in car1600 × 1200@12HzO2023
Fixed bracket [65]P&A3 × Kinect v21920 × 1080@130HzI2023
Fixed bracket [66]P&A7 × cameras128 × 2048@10HzO2023
Car [67]F&A4 × cameras in car1920 × 1208@20HzO2022
HermesBot [68]P&A1 × RGB-D camera
1 × binocular fisheye camera
1 × rolling-shutter camera
640 × 480@20HzI2022
Ground robot [69]P&A1 × camera
1 × infrared camera
1 × event camera
1280 × 1024@15Hz
640 × 512@25Hz
640 × 480@15Hz
I2022
Car [70]F&C6 × Ladybug31600 × 1200@5HzO2022
Small Submarine [71]F&C4 × GoPro Hero 4 Black3840 × 2160@30HzO2022
Head-mounted [72]P&A2 × Prophesee Gen3 CD with Kowa LM5JCM
2 × FLIR Grasshopper3 with Kowa LM6JCM
1 × Kinect
1 × Ouster OS0-128
640 × 480@N/A
1224 × 1024@30Hz
640 × 576@30Hz
128 × 2048@10Hz
O2022
Car [73]F&A4 × cameras in car1920 × 1208@25HzI&O2021
Handheld [74]F&A4 × grayscale fisheye cameras720 × 540@30HzO2021
Fixed bracket [75]P&A6 × mvBlueFOX-MLC 200 w752 × 480@20HzI&O2021
Simulation [76]P&A3 × binocular cameras installed on a semicircle640 × 480@30HzI&O2021
Underwater robot [77]P&F&A1 × Gopro
3 × Fish Eyes Gopro
2 × binocular cameras
1 × UHI camera
2704 × 1520@60Hz
2704 × 2028@30Hz
1920 × 1080@30Hz
648 × 486@20Hz
O2021
Drone [78]P&A1 × fast camera (XIMEA MQ 003 CG-CM)
1 × slow camera (Sony IU 233 N2-Z)
640 × 400@120Hz
648 × 488@500Hz
I2021
Car [79]P&A4 × Basler acA 1920 - 50 gc GigE1920 × 1200@50HzO2020
Underwater robot [80]P&A3 × Basler Ace acA 1920 - 48 gc1920 × 1200@50HzO2020
Truck [81]F&A4 × cameras1600 × 1532@20HzO2020
Assembled small robot [82]F&A4 × fisheye cameras1600 × 1532@10HzI2020
Car [83]P&C5 × cameras752 × 480@20HzO2020
Fixed bracket [84]P&A3 × mvBlueFox-MLC 200 wG752 × 480@93HzI2020
Drone [85]P&A3 × Pointgrey Blackfly1280 × 960@20HzI2019
Fixed bracket [86]P&A3 × Kinect320 × 240@30Hz (Depth)
640 × 480@30Hz (RGB)
I2018
Table 3. The Dataset for MCSLAM since 2009 (Env.: Environment, I: Indoor, O: Outdoor, U: Urban, S: Static, D: Dynamic, Dur.: Duration, L: Long-term, Sh: Short-term).
Table 3. The Dataset for MCSLAM since 2009 (Env.: Environment, I: Indoor, O: Outdoor, U: Urban, S: Static, D: Dynamic, Dur.: Duration, L: Long-term, Sh: Short-term).
DatasetEnv.PlatformCamerasOther SensersScene
Mode
Ground
Truth
Dur.Years
New
college [181]
IWheel
robot
1 Panoramic RGB
 5 × 384 × 512@3Hz
1 Stereo
 2 × 512 × 384@20Hz
2 LiDARs
 LMS 291-S14@75Hz
1 GPS@5Hz
1 IMU
S&DLiDARL2009
Rawseeds [182]I&OGround
robot
1 Trinocular gray
 3 × 640 × 480@15Hz
1 RGB
 640 × 480@30Hz
1 Fisheye RGB
 640 × 640@15Hz
1 IMU
 1 accel/gyro@128Hz
1 GPS
 RTK@5Hz
4 LiDARs
 2 Hokuyo@10Hz
 2 SICK@75Hz
1 Sonar belt (Indoor)
S&DGPS
LiDAR
Sh2009
KITTI [183]OCar1 Stereo
 2 × 1392 × 512@10Hz
1 Stereo RGB
 21,392 × 512@10Hz
1 LIDAR
 Velodyne HDL-64E 3D@10Hz
1 IMU
1 RTK
S&DGNSS
INS
L2012
IPDS [184]OElectric
cart
1 Classical camera
3 Synchronized cameras
2 Large FOV cameras@7.5Hz
1 Webcam
1 GPS@10Hz
1 IMU
 acc./gyro@50Hz
S&DGPSL2013
MIT stata
dataset [185]
IWheel
robot
1 Stereo RGB
1 Infrared RGB-D
2 LiDARs
 Hokuyo UTM-30LX
1 IMU
 Microstrain 3DM-GX2
SLiDARL2013
AMUSE [186]OVehicle1 Omnidirectional RGB
 6 × 1616 × 616@30Hz
1 IMU
 XSens MTi AHRS3
1 GPS
 uBlox AEK-4P
1 VF@250Hz
S&DGPS
INS
L2013
NCLT [187]I&OSegway6 RGB
 1600 × 1200@5Hz
1 IMU
  3DM-GX3-45
  3-axis acc./gyro@100Hz
3 LiDARs
  1 HDL-32E@10Hz
  1 UTM-30LX
  1 URG-04LX
1 FOG
  KVH DSP-3000
2 GPS
  1 Garmin 18x@5Hz
  1 RTK@1HZ
DGPS
IMU
LiDAR
L2016
TorontoCity [188]OVehicle1 GoProHero
4 RGB camera
1 PointGray Bumblebee
3 Stereo Camera
1 LIDAR
 Velodyne HDL-64E
S&DHigh-precision
maps
L2016
LaFiDa [189]I&OHelmet3 Fisheye cameras
 754 × 480@70Hz
3 LiDARs
 Hokuyo UTM-30LX-EW
S&DIMU
LiDAR
Sh2017
PennCOSYVIO [190]I&OHandheld4 Gopro hero 4
1 RGB (rolling shutter)
 1920 × 1080@30Hz
1 Stereo gray-scale
 2 × 752 × 480@20Hz
1 Fisheye gray-scale
 640 × 480@30Hz
3 IMUs
 1 ADIS16488
 3-axis acc./gyro@200Hz
 2 Tango
 3-axis acc.@128Hz
 3-axis gyro@100Hz
DINSSh2017
Oxford
Robot
car [191]
OutdoorVehicle1 Trinocular stereo
  3 × 1280 × 960@16Hz
3 Fisheye gray
  1024 × 1024@11.1Hz
3 LiDARs
 2 SICKLMS-151 2D@50Hz
 1 SICKLD-MRS 3D@12.5Hz
1 IMU
1 GPS
DLiDARL2017
ApolloScape [192]OCar6 video cameras
 3384 × 2710
2 LiDARs
 VUX-1HA
1 GNSS
1 IMU
S&DGNSS
LiDAR
L2018
MVSEC [193]I&OHandheld
Hexacopter
Car
Motorcycle
1 Event stereo
 2 × 346 × 260
1 Stereo
 2 × 752 × 480
1 LiDAR
 Velodyne Puck LITE@20Hz
1 GPS
 UBLOX NEO-M8N
2 IMUs
 1MPU 6150
 1 ADIS16488@200Hz
S&DGNSS
LiDAR
IMU
L2018
Table 4. The Dataset for MCSLAM since 2009 (Env.: Environment, I: Indoor, O: Outdoor, U: Urban, S: Static, D: Dynamic, Dur.: Duration, L: Long-term, Sh: Short-term).
Table 4. The Dataset for MCSLAM since 2009 (Env.: Environment, I: Indoor, O: Outdoor, U: Urban, S: Static, D: Dynamic, Dur.: Duration, L: Long-term, Sh: Short-term).
DatasetEnv.PlatformCamerasOther SensersScene
Mode
Ground
Truth
Dur.Years
UZH-FPV [194]I&OMAV1 Stereo gray-scale
 2 × 640 × 480@30Hz,
1 Event camera
 346 × 260@50Hz
1 Laser
 1 Leica Nova MS60 TotalStation@20HZ
2 IMUs
 3-axis acc./gyro/ magn.@500Hz
 3-axis acc./gyro@1000Hz
SLaserSh2019
WoodScape [195]OCar4 Fisheye
 1MPx RGB with 190 horizontal FOV
1 LiDAR
 Velodyne HDL-64E@20Hz
1 GNSS/IMU
 NovAtel Propak6 & SPAN-IGM-A1
 1 GNSS Positioning with SPS
S&DLiDARSh2019
Argoverse [196]OVehicle7 Ring cameras
 1920 × 1200@30Hz
2 Stereo cameras
 4 × 2056 × 2464@5Hz
2 LiDARs
 VLP-32
1 GPS
S&DGPSL2019
OpenLORIS [197]IGround
robot
1 RGB-D(rolling shutter)
 848 × 480@30Hz
1 Stereo fisheye RGB
 2 × 848 × 480@30Hz
2 IMUs
 2 BMI055
 3-axis acc.@250Hz
 3-axis gyro@400Hz
1 LiDAR
 UTM-30LX 30m
S&DLaser
Tracker
Pose
@40Hz
Sh2020
UrbanLoco [198]UCar6 Cameras (California)
 2048 × 1536@10Hz
1 LiDAR
 RS-LIDAR-32@10Hz (California)
1 IMU
 Xsens Mti@10100Hz
1 GNSS
 Ublox M8T GPS@1Hz
S&DGNSSL2020
Brno Urban [199]UCar4 RGB cemara
 1920 × 1200@10Hz
1 Thermal camera
 640 × 512@30Hz
2 LiDARs
 Velodyne HDL-32e@10Hz
1 GNSS
 Trimble BX982@20Hz
1 IMU
 Xsens MTi-G-710
S&DGNSSL2020
UMA-VI [200]I&OHandheld1 Stereo RGB camera
 2 × 1024 × 768@12.5Hz
1 Stereo camera
 2 × 752 × 480@25Hz
1 IMU
 XSens MTi-28A53G35 3D@250Hz
DPseudo-
ground
truth
poses
L2020
FinnForest
dataset [79]
OCar4 RGB
 1920 × 1200
1 IMU
 Fiber optic gyro@200Hz
1 GPS
 OEM7 GNSS@20Hz
S&DGPSL2020
nuScenes [201]OCar6 Cameras
 1600 × 900@12Hz
1 LiDAR
 32 beams@20Hz
1 GPS
 20 mm RTK
1 IMU
5 RADARs@13Hz
S&DGPSSh2020
PanoraMIS [202]I&OWheel
Aerial industrial
1 Catadioptric
 VStone objective hyperbolic mirror
1 Twin-Fisheye
 2 × 1280 × 720@30Hz
2 IMUs
2 GPS
S&DGPS
IMU
L2020
CADCD [203]OVehicle8 Cameras
 1280 × 1024@10Hz
1 LiDAR
 Velodyne VLP-32C@10Hz
1 GNSS
 Triple-Frequency
3 IMUs
 1 STIM300 MEMS@100Hz
 2 Xsens@200Hz
S&DGNSSL2020
Waymo Open [204]OVehicle5 Cameras
 3 × 1920 × 1280
 2 × 1920 × 1040
5 LiDARsS&DLiDARL2020
A2D2 [205]OVehicle6 Cameras
 5 Sekonix SF3324-100
 1928 × 1208@30Hz
 1 Sekonix SF3325-100
 1928 × 1208@30Hz
5 LiDARs
 Velodyne VLP-16@10Hz
Bus gateway
S&DLiDARL2020
EU [206]OVehicle2 Stereo cameras
 1 Bumblebee XB3
 1 Bumblebee2
2 Fisheye cameras
 Pixelink PL-B742F
4 LiDARs
 2 Velodyne HDL-32E
 1 ibeo LUX 4L
 1 SICK LMS100-10000
1 GPS
 RTK Magellan ProFlex500
1 IMU
 Xsens MTi-28A53G25
1 RADAR
 Continental ARS 308
S&DGPS
LiDAR
L2020
Table 5. The Dataset for MCSLAM since 2009 (Env.: Environment, I: Indoor, O: Outdoor, U: Urban, S: Static, D: Dynamic, Dur.: Duration, L: Long-term, Sh: Short-term).
Table 5. The Dataset for MCSLAM since 2009 (Env.: Environment, I: Indoor, O: Outdoor, U: Urban, S: Static, D: Dynamic, Dur.: Duration, L: Long-term, Sh: Short-term).
DatasetEnv.PlatformCamerasOther SensersScene
Mode
Ground
Truth
Dur.Years
The
Newer
College
Dataset [207]
I&OHandheld1 Stereo fisheye grayscale
 2 × 720 × 540@30Hz
2 Grayscale fisheye
 720 × 540@30Hz
1 LiDAR
 Ouster, OS0-128@10Hz
2 IMUs
 1 ICM-20948@100Hz
 1 Bosch BMI085@200Hz
DLiDARSh2021
M2DGR [69]I&OGround
robot
6 Fisheye RGB
 1280 × 1024@15Hz
1 Infrared camera
 640 × 512@25Hz
1 Event camera
 640 × 480@15Hz
1 RGB-D (rolling shutter)
 640 × 480@15Hz
2 IMUs
1 Handsfree A9
 3-axis acc./gyro/ magn.@150Hz
1 BMI055
 3-axis acc./gyro@200Hz
1 LiDAR
 Velodyne VLP-32C@10Hz
1 GNSS
 Ublox M8T@100Hz
DGNSS
LiDAR
L2021
AMV-Bench [176]OVehicle5 Wide-angle RGB
 1920 × 1200@10Hz
1 Stereo RGB
1 LiDAR
1 IMU
1 GPS
S&DGNSS
LiDAR
L2021
DSEC [208]OCar1 Event stereo
 2 × 640 × 480
1 Stereo RGB
 2 × 1440 × 1080@20Hz
1 LiDAR
 Velodyne VLP-16@10Hz
1 GNSS
 RTK@10Hz
S&DGNSS
LiDAR
Sh2021
TUM-VIE [209]I&OHandheld,
Head-
mounted
1 Event stereo
 2 × 1280 × 720
1 Stereo
 2 × 1024 × 1024@20Hz
1 IMU
 3-axis acc./gyro@200Hz
S&DIMUSh2021
Ford
campus [210]
OVehicle6 Cameras
 Pointgrey 2009
 1600 × 1200
2 LiDARs
 1 Velodyne 2007@15Hz
 1 RIEGL 2010
1 GPS
 Applanix 2010
1 IMU
 Xsens 2010
S&DGNSS
LiDAR
IMU
L2011
The Hilti
SLAM
Challenge
Dataset [211]
I&OHandheld5 Cameras1 Ouster OS0-64
1 Livox MID70
3 IMUs
 1 Bosch IMU@200Hz
 1 ADIS16445@800Hz
 1 Ivensense@100Hz
DIMUSh2022
PanoVILD [70]OVehicle6 Cameras
 2-Meg-apixel Ladybug3
 1600 × 1200@5Hz
1 LiDAR
 Ouster os1-64
1 GPS
 RTK NovAtel@100Hz
1 IMU
 3-axis acc./gyro@100Hz
S&DGPSL2022
Vector [72]IHandheld,
Helmet,
Wheeled
tripod
1 Event stereo camera
 2 × 640 × 480
1 Stereo camera
 2 × 1224 × 1024@30Hz
1 RGB-D Depth
 640 × 576@30Hz
1 RGB-D Color
 128 × 2048@10Hz
1 LiDAR
 FARO 128-channel
1 IMU
 XSens MTi-30 AHRS
 3-axis acc./gyro/ magn.@200Hz
S&DLiDAR
IMU
Sh2022
IndoorMCD [63]IHandheld,
Ground
robot
3 RGB-D (rolling shutter)
 640 × 480@15Hz
3 IMUs
 3 BMI055
 3-axis acc.@250Hz
 3-axis gyro@400Hz
S&DIMUSh2023
MultiCamSLAM [66]I&OGround
robot
7 Cameras
 FLIR BlackFly S 1.3 MP color
 720 × 540
1 IMU
 Vectornav@200Hz
1 GPS
S&DGPSSh2023
USTC
FLICAR [212]
OBucket
Truck
2 Cameras
 1 × 1440 × 1080@20Hz
 1 × 3072 × 2048@20Hz
2 Stereo cameras
 3 × 1280 × 960
 2 × 1024 × 768
1 IMU
 9 axis Xsens MTi-G-710@400Hz
4 LiDARs
 1 3D Velodyne HDL-32E
 1 3D Velodyne VLP-32C
 1 3D LiVOX Avia
 1 3D Ouster OS0-12
S&D3D Laser
Tracker
Sh2023
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Wang, G.; Wang, L.; He, J.; Jiang, Y.; Qi, Q.; Zhou, Y. Multi-Camera Simultaneous Localization and Mapping for Unmanned Systems: A Survey. Electronics 2026, 15, 602. https://doi.org/10.3390/electronics15030602

AMA Style

Wang G, Wang L, He J, Jiang Y, Qi Q, Zhou Y. Multi-Camera Simultaneous Localization and Mapping for Unmanned Systems: A Survey. Electronics. 2026; 15(3):602. https://doi.org/10.3390/electronics15030602

Chicago/Turabian Style

Wang, Guoyan, Likun Wang, Jun He, Yanwen Jiang, Qiming Qi, and Yueshang Zhou. 2026. "Multi-Camera Simultaneous Localization and Mapping for Unmanned Systems: A Survey" Electronics 15, no. 3: 602. https://doi.org/10.3390/electronics15030602

APA Style

Wang, G., Wang, L., He, J., Jiang, Y., Qi, Q., & Zhou, Y. (2026). Multi-Camera Simultaneous Localization and Mapping for Unmanned Systems: A Survey. Electronics, 15(3), 602. https://doi.org/10.3390/electronics15030602

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop