MDR–SLAM: Robust 3D Mapping in Low-Texture Scenes with a Decoupled Approach and Temporal Filtering
Abstract
1. Introduction
- A novel modular system architecture: We proposed the MDR–SLAM framework, whose decoupling design aims to accommodate computing-intensive mapping algorithms to cope with challenging scenarios that traditional monolithic systems find difficult to handle.
- A robust temporal density depth estimation algorithm: We have designed a temporal filtering algorithm driven by keyframes, which is specially used to enhance the density of point clouds and reconstruction integrity in low-texture environments.
- A dynamic fusion strategy for long-term mapping: We proposed an incremental fusion strategy containing a confidence model. It dynamically maintains the global map through a closed-loop process of “fusion-update-cleaning” to ensure the long-term consistency of large-scale scene mapping.
2. Related Work
2.1. Pose Estimation
2.2. Dense 3D Reconstruction in Low-Texture Scenes
2.3. System Architectures in SLAM
3. Method
3.1. System Overview
- Data input: The system uses a stereo camera as its main sensor, which releases left and right image streams at a fixed frequency, and comes with corresponding camera calibration information.
- Parallel processing: the position estimation module (node 1) and the depth estimation module (node 2) subscribe to these image streams in parallel. Node 1 focuses on calculating the six-degree-of-freedom (6-DoF) position of the camera in real time, while node 2 is responsible for generating a dense point cloud for each stereo image pair in the current camera coordinate system.
- Data fusion: As the map fusion module (node 3) at the back end of the system, subscribe to the output of the first two nodes. It adopts a key time synchronization mechanism to ensure that each point cloud is associated with its best-matched camera position in time.
- Map generation: After performing the coordinate transformation, node 3 merges the local point cloud in the world coordinate system into a global map. In order to effectively manage the data volume, the system down samples the global map, and finally publishes a global dense point cloud for visualization or other downstream applications.
3.2. Pose Estimation Module
- Bundle Adjustment (BA): As the gold standard for the refinement process in visual SLAM, bundle adjustment is a large-scale nonlinear optimization problem, which minimizes the total weight projection error on a series of camera positions and three-dimensional points. Its cost function is expressed as follows:
- 2.
- Pose Graph Optimization: The goal of pose graph optimization is to adjust the posture of all nodes in the diagram to minimize the error of these constraints. Its cost function can be written as follows:
3.3. Dense Depth Estimation with Temporal Filtering
3.3.1. ELAS: Single-Frame Depth Estimation
- Prior Probability : Based on the a priori assumption that the surface in the scene is generally parallel to the camera imaging plane, the model sets a Gaussian a priori for the plane parameters. Its probability density function is directly proportional to the following:
- 2.
- Likelihood Function : The likelihood function measures how likely the observed support point disparity is to occur in the case of a given specific plane hypothesis:
3.3.2. Temporal Depth Filtering
3.4. Global Map Fusion
3.4.1. Coordinate Transformation
3.4.2. Incremental Map Fusion with Confidence Updating
4. Results
4.1. Experimental Setup
4.1.1. Datasets
- EuRoC MAV dataset: This is a visual inertial dataset collected by drones in the indoor environment, which contains high-precision ground true values. We use its three-dimensional image sequence to quantitatively evaluate the trajectory accuracy of the position estimation front-end (node 1).
- KITTI odometer dataset: As one of the most authoritative benchmarks in the field of autonomous driving, the dataset contains a sequence of stereo images collected from the on-board platform in a large outdoor environment. We use it to evaluate the quality and long-term stability of the three-dimensional dense mapping of the whole system in such scenarios.
- In order to further verify the adaptability of the framework under different sensors and scenarios, we used the Intel RealSense D435i camera (Intel Corporation, Santa Clara, CA, USA) to collect a series of challenging indoor scene data. The dataset is used to qualitatively display the actual reconstruction performance of the system.
- For the core quantitative map evaluation, we utilize the recently published NUFR-M3F (Neufield Robotics Multi-Modal Multi-Floor) Dataset [35]. This dataset is specifically chosen because it provides synchronized stereo visual data (from a ZED camera (Stereolabs, San Francisco, CA, USA)) and high-fidelity 3D map ground truth (generated from a LiDAR sweep), which is essential for calculating our RMSE and Completeness metrics. We utilize the 2nd_floor.bag sequence, which simulates a challenging indoor office environment with prevalent low-texture walls, long corridors, and structural symmetry. The dataset is publicly available at https://github.com/neufieldrobotics/NUFR-M3F?tab=readme-ov-file (accessed on 7 December 2025).
4.1.2. Evaluation Metrics
- Absolute trajectory error (ATE, m): We use ATE to measure the global consistency of position estimation. After aligning the estimated trajectory with the true value of the ground, the indicator calculates the root mean square error (RMSE) between all corresponding positions.
- Map Accuracy (RMSE, m): This metric measures how closely the estimated points align with the true surface. It is defined as the Root Mean Square Error (RMSE) of the distances from every point to its nearest neighbor in .
- 3.
- Map Completeness (%): This measures the percentage of the ground truth map that has been successfully reconstructed. It is calculated as the proportion of points that are within a distance threshold of their nearest neighbor in .
4.1.3. Camera Calibration
4.2. Accuracy Analysis of Pose Estimation
4.3. Ablation Study
4.3.1. Low-Texture Qualitative Evidence
- Uniform Low-Texture (Column a): In the scenario dominated by a large, uniform white wall, the raw ELAS output is plagued by high noise and erroneous depth estimates, resulting in a fractured and highly inconsistent point cloud. The MDR–SLAM output, however, successfully filters this noise, generating a much smoother and geometrically accurate representation of the wall and surrounding structures.
- Specular/Ambiguous Texture (Column b): In the area featuring strip—like white light bands, raw ELAS fails to establish correct correspondences, leading to significant depth ambiguity and large voids in the reconstruction. The MDR–SLAM system successfully aggregates geometric constraints across multiple frames, effectively filling these voids and reconstructing the environment with high fidelity.
- High-Motion Cornering (Column c): During a corridor cornering motion (high translation/rotation), the ELAS depth map suffers from motion blur and frame inconsistency. The MDR–SLAM output maintains structural integrity during this challenging maneuver, demonstrating the filter’s ability to maintain consistency even when input quality is momentarily degraded.
4.3.2. Map Quality Analysis
4.4. Reconstruction Showcase
4.4.1. Quantitative Comparison with Baselines
4.4.2. Indoor Reconstruction in Complex and Degraded Scenes
4.5. Performance and Resource Analysis
4.5.1. Decoupling Advantage and Resource Isolation
4.5.2. Platform Adaptability and Efficiency
4.5.3. Real-Time Throughput Validation
5. Discussion
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Han, J.; Ma, C.; Zou, D.; Jiao, S.; Chen, C.; Wang, J. Distributed Multi-Robot SLAM Algorithm with Lightweight Communication and Optimization. Electronics 2024, 13, 4129. [Google Scholar] [CrossRef] [Scilit]
- Rostum, H.; Vásárhelyi, J. Enhancing Machine Learning Techniques in VSLAM for Robust Autonomous Unmanned Aerial Vehicle Navigation. Electronics 2025, 14, 1440. [Google Scholar] [CrossRef] [Scilit]
- Cadena, C.; Carlone, L.; Carrillo, H.; Latif, Y.; Scaramuzza, D.; Neira, J.; Reid, I.; Leonard, J.J. Past, Present, and Future of Simultaneous Localization and Mapping: Towards the Robust-Perception Age. IEEE Trans. Robot. 2016, 32, 1309–1332. [Google Scholar] [CrossRef] [Scilit]
- Rodríguez-Lira, D.-C.; Córdova-Esparza, D.-M.; Terven, J.; Romero-González, J.-A.; Alvarez-Alvarado, J.M.; González-Barbosa, J.-J.; Ramírez-Pedraza, A. Recent Developments in Image-Based 3D Reconstruction Using Deep Learning: Methodologies and Applications. Electronics 2025, 14, 3032. [Google Scholar] [CrossRef] [Scilit]
- Kerbl, B.; Kopanas, G.; Leimkuehler, T.; Drettakis, G. 3D Gaussian Splatting for Real-Time Radiance Field Rendering. ACM Trans. Graph. 2023, 42, 1–14. [Google Scholar] [CrossRef] [Scilit]
- Mildenhall, B.; Srinivasan, P.P.; Tancik, M.; Barron, J.T.; Ramamoorthi, R.; Ng, R. NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis. Commun. ACM 2022, 65, 99–106. [Google Scholar] [CrossRef] [Scilit]
- Mur-Artal, R.; Tardos, J.D. ORB-SLAM2: An Open-Source SLAM System for Monocular, Stereo, and RGB-D Cameras. IEEE Trans. Robot. 2017, 33, 1255–1262. [Google Scholar] [CrossRef] [Scilit]
- Campos, C.; Elvira, R.; Rodríguez, J.J.G.; Montiel, J.M.M.; Tardós, J.D. ORB-SLAM3: An Accurate Open-Source Library for Visual, Visual-Inertial and Multi-Map SLAM. IEEE Trans. Robot. 2021, 37, 1874–1890. [Google Scholar] [CrossRef] [Scilit]
- Qin, T.; Li, P.; Shen, S. VINS-Mono: A Robust and Versatile Monocular Visual-Inertial State Estimator. IEEE Trans. Robot. 2018, 34, 1004–1020. [Google Scholar] [CrossRef] [Scilit]
- Zhu, Z.; Peng, S.; Larsson, V.; Xu, W.; Bao, H.; Cui, Z.; Oswald, M.R.; Pollefeys, M. NICE-SLAM: Neural Implicit Scalable Encoding for SLAM. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; IEEE: New Orleans, LA, USA, 2022; pp. 12776–12786. [Google Scholar]
- Wang, Y.; Zhang, Y.; Hu, L.; Ge, G.; Wang, W.; Tan, S. Improved Feature Point Extraction Method of VSLAM in Low-Light Dynamic Environment. Electronics 2024, 13, 2936. [Google Scholar] [CrossRef] [Scilit]
- Yang, X.; Jiang, G. A Practical 3D Reconstruction Method for Weak Texture Scenes. Remote Sens. 2021, 13, 3103. [Google Scholar] [CrossRef] [Scilit]
- Ma, Y.; Lv, J.; Wei, J. High-Precision Visual SLAM for Dynamic Scenes Using Semantic–Geometric Feature Filtering and NeRF Maps. Electronics 2025, 14, 3657. [Google Scholar] [CrossRef] [Scilit]
- Hirschmuller, H. Stereo Processing by Semiglobal Matching and Mutual Information. IEEE Trans. Pattern Anal. Mach. Intell. 2008, 30, 328–341. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Barnes, C.; Shechtman, E.; Finkelstein, A.; Goldman, D.B. PatchMatch: A Randomized Correspondence Algorithm for Structural Image Editing. ACM Trans. Graph. 2009, 28, 1–11. [Google Scholar] [CrossRef] [Scilit]
- Engel, J.; Schöps, T.; Cremers, D. LSD-SLAM: Large-Scale Direct Monocular SLAM. In Computer Vision–ECCV 2014; Fleet, D., Pajdla, T., Schiele, B., Tuytelaars, T., Eds.; Lecture Notes in Computer Science; Springer International Publishing: Cham, Switzerland, 2014; Volume 8690, pp. 834–849. ISBN 978-3-319-10604-5. [Google Scholar]
- Engel, J.; Koltun, V.; Cremers, D. Direct Sparse Odometry. IEEE Trans. Pattern Anal. Mach. Intell. 2018, 40, 611–625. [Google Scholar] [CrossRef] [Scilit]
- Crocetti, F.; Brilli, R.; Dionigi, A.; Fravolini, M.L.; Costante, G.; Valigi, P. Comparison of DSO and ORB-SLAM3 in Low-Light Environments with Auxiliary Lighting and Deep Learning Based Image Enhancing. J. Field Robot. 2025, 42, 3748–3771. [Google Scholar] [CrossRef] [Scilit]
- Geiger, A.; Roser, M.; Urtasun, R. Efficient Large-Scale Stereo Matching. In Computer Vision–ACCV 2010; Kimmel, R., Klette, R., Sugimoto, A., Eds.; Lecture Notes in Computer Science; Springer: Berlin/Heidelberg, Germany, 2011; Volume 6492, pp. 25–38. ISBN 978-3-642-19314-9. [Google Scholar]
- Forster, C.; Pizzoli, M.; Scaramuzza, D. SVO: Fast Semi-Direct Monocular Visual Odometry. In Proceedings of the 2014 IEEE International Conference on Robotics and Automation (ICRA), Hong Kong, China, 31 May–7 June 2014; IEEE: Hong Kong, China, 2014; pp. 15–22. [Google Scholar]
- Newcombe, R.A.; Lovegrove, S.J.; Davison, A.J. DTAM: Dense Tracking and Mapping in Real-Time. In Proceedings of the 2011 International Conference on Computer Vision, Barcelona, Spain, 6–13 November 2011; IEEE: Barcelona, Spain, 2011; pp. 2320–2327. [Google Scholar]
- Luo, H.; Zhang, J.; Liu, X.; Zhang, L.; Liu, J. Large-Scale 3D Reconstruction from Multi-View Imagery: A Comprehensive Review. Remote Sens. 2024, 16, 773. [Google Scholar] [CrossRef] [Scilit]
- Teed, Z.; Deng, J. DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D Cameras. Adv. Neural Inf. Process. Syst. 2021, 34, 16558–16569. [Google Scholar]
- Romanoni, A.; Matteucci, M. TAPA-MVS: Textureless-Aware PAtchMatch Multi-View Stereo. In Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 10412–10421. [Google Scholar] [CrossRef] [Scilit]
- Asafa, G.F.; Ren, S.; Mamun, S.S.; Gobena, K.A. DepthCloud2Point: Depth Maps and Initial Point for 3D Point Cloud Reconstruction from a Single Image. Electronics 2025, 14, 1119. [Google Scholar] [CrossRef] [Scilit]
- Furukawa, Y.; Ponce, J. Accurate, Dense, and Robust Multiview Stereopsis. IEEE Trans. Pattern Anal. Mach. Intell. 2010, 32, 1362–1376. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Xu, Q.; Kong, W.; Tao, W.; Pollefeys, M. Multi-Scale Geometric Consistency Guided and Planar Prior Assisted Multi-View Stereo. IEEE Trans. PATTERN Anal. Mach. Intell. 2023, 45, 4945–4963. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yao, Y.; Luo, Z.; Li, S.; Shen, T.; Fang, T.; Quan, L. Recurrent MVSNet for High-Resolution Multi-View Stereo Depth Inference. In Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019; pp. 5520–5529. [Google Scholar] [CrossRef] [Scilit]
- Gu, X.; Fan, Z.; Zhu, S.; Dai, Z.; Tan, F.; Tan, P. Cascade Cost Volume for High-Resolution Multi-View Stereo and Stereo Matching. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 2492–2501. [Google Scholar] [CrossRef] [Scilit]
- Laga, H.; Jospin, L.V.; Boussaid, F.; Bennamoun, M. A Survey on Deep Learning Techniques for Stereo-Based Depth Estimation. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 44, 1738–1764. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Rosinol, A.; Abate, M.; Chang, Y.; Carlone, L. Kimera: An Open-Source Library for Real-Time Metric-Semantic Localization and Mapping. In Proceedings of the 2020 IEEE International Conference on Robotics and Automation (ICRA), Paris, France, 31 May–31 August 2020; pp. 1689–1696. [Google Scholar] [CrossRef] [Scilit]
- Klein, G.; Murray, D. Parallel Tracking and Mapping for Small AR Workspaces. In Proceedings of the 2007 6th IEEE and ACM International Symposium on Mixed and Augmented Reality, Nara, Japan, 13–16 November 2007; IEEE: Nara, Japan, 2007; pp. 1–10. [Google Scholar]
- Ragot, N.; Khemmar, R.; Pokala, A.; Rossi, R.; Ertaud, J.-Y. Benchmark of Visual SLAM Algorithms: ORB-SLAM2 vs RTAB-Map. In Proceedings of the 2019 Eighth International Conference on Emerging Security Technologies (EST), Colchester, UK, 22–24 July 2019; IEEE: Colchester, UK, 2019; pp. 1–6. [Google Scholar]
- Qin, T.; Cao, S.Z.; Pan, J.; Shen, S.J. A General Optimisation-Based Framework for Global Pose Estimation with Multiple Sensors. IET Cyber-Syst. Robot. 2025, 7, e70023. [Google Scholar] [CrossRef] [Scilit]
- Kaveti, P.; Gupta, A.; Giaya, D.; Karp, M.; Keil, C.; Nir, J.; Zhang, Z.; Singh, H. Challenges of Indoor SLAM: A Multi-Modal Multi-Floor Dataset for SLAM Evaluation. In Proceedings of the 2023 IEEE 19th International Conference on Automation Science and Engineering (CASE), Auckland, New Zealand, 26–30 August 2023; IEEE: Auckland, New Zealand, 2023; pp. 1–8. [Google Scholar]








| Sequence | RMSE (m) | Mean (m) | Std (m) | Median (m) | Max (m) | Min (m) |
|---|---|---|---|---|---|---|
| MH_01_easy | 0.035 | 0.028 | 0.022 | 0.022 | 0.095 | 0.002 |
| MH_03_medium | 0.036 | 0.032 | 0.017 | 0.028 | 0.106 | 0.002 |
| MH_05_difficult | 0.050 | 0.044 | 0.024 | 0.038 | 0.156 | 0.010 |
| V1_01_easy | 0.087 | 0.081 | 0.030 | 0.072 | 0.164 | 0.030 |
| V1_02_medium | 0.060 | 0.058 | 0.017 | 0.059 | 0.107 | 0.010 |
| Method | Accuracy (RMSE, m) | Completeness (%) | Trajectory (RMSE, m) |
|---|---|---|---|
| ELAS-only | 0.072 | 39.33 | 0.736 |
| MDR–SLAM | 0.012 | 45.15 | 0.327 |
| System | Type | Accuracy (RMSE, m) | Completeness (%) | Throughput (FPS) |
|---|---|---|---|---|
| MDR–SLAM | Real-time SLAM | 0.099 m | 26.31% | 4.7 Hz |
| OpenMVS | Offline MVS | 0.227 m | 34.07% | <0.1 Hz |
| Module (Node) | Component | CPU Usage (%) | Memory (MB) | GPU Usage (%) | Output Freq. (Hz) |
|---|---|---|---|---|---|
| Node 1 | ORB-SLAM2 (Tracking) | 98.0% | ~730 MB | 0% | ~20 Hz (Input) |
| Node 2 | Temporal ELAS (Depth) | 84.8% | ~164 MB | 0% | 4.7 Hz |
| Node 3 | Confidence Map (Fusion) | 80.4% | ~463 MB | 0% | 4.7 Hz |
| System Total | End-to-End Throughput | ~263% (2.6 Cores) | ~1.3 GB | 0% | 4.7 H |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
Share and Cite
Zhang, K.; Zhou, L. MDR–SLAM: Robust 3D Mapping in Low-Texture Scenes with a Decoupled Approach and Temporal Filtering. Electronics 2025, 14, 4864. https://doi.org/10.3390/electronics14244864
Zhang K, Zhou L. MDR–SLAM: Robust 3D Mapping in Low-Texture Scenes with a Decoupled Approach and Temporal Filtering. Electronics. 2025; 14(24):4864. https://doi.org/10.3390/electronics14244864
Chicago/Turabian StyleZhang, Kailin, and Letao Zhou. 2025. "MDR–SLAM: Robust 3D Mapping in Low-Texture Scenes with a Decoupled Approach and Temporal Filtering" Electronics 14, no. 24: 4864. https://doi.org/10.3390/electronics14244864
APA StyleZhang, K., & Zhou, L. (2025). MDR–SLAM: Robust 3D Mapping in Low-Texture Scenes with a Decoupled Approach and Temporal Filtering. Electronics, 14(24), 4864. https://doi.org/10.3390/electronics14244864

