Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (17)

Search Parameters:
Keywords = monocular ego-motion

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
21 pages, 36892 KB  
Article
Self-Supervised Depth and Ego-Motion Learning from Multi-Frame Thermal Images with Motion Enhancement
by Rui Yu, Guoliang Ma, Jian Guo and Lisong Xu
Appl. Sci. 2025, 15(22), 11890; https://doi.org/10.3390/app152211890 - 8 Nov 2025
Viewed by 1262
Abstract
Thermal cameras are known for their ability to overcome lighting constraints and provide reliable thermal radiation images. This capability facilitates methods for depth and ego-motion estimation, enabling efficient learning of poses and scene structures under all-day conditions. However, the existing studies on depth [...] Read more.
Thermal cameras are known for their ability to overcome lighting constraints and provide reliable thermal radiation images. This capability facilitates methods for depth and ego-motion estimation, enabling efficient learning of poses and scene structures under all-day conditions. However, the existing studies on depth prediction for thermal images are limited. In practical applications, thermal cameras capture sequential frames. Unfortunately, the potential of this multi-frame aspect is underutilized by the previous methods, resulting in limitations on the depth prediction accuracy of thermal videos. To leverage the multi-frame advantages of thermal videos and to improve the accuracy of monocular depth estimation from thermal images, we propose a framework for self-supervised depth and ego-motion learning from multi-frame thermal images. We construct a multi-view stereo (MVS) cost volume from temporally adjacent thermal frames. The construction process is adjusted based on the estimated pose, which serves as a motion hint. To stabilize the motion hint and improve pose estimation accuracy, we design a motion enhancement module that utilizes self-generated poses for additional supervisory signals. Additionally, we introduce RGB images in the training phase to form a multi-spectral loss, thereby augmenting the performance of the thermal model. The experimental results, conducted on a public dataset, demonstrate the proposed method’s accurate estimation of depth and ego-motion across varying light conditions, surpassing the performance of the self-supervised baseline. Full article
(This article belongs to the Special Issue Application of Artificial Intelligence in Image Processing)
Show Figures

Figure 1

28 pages, 10678 KB  
Article
Deep-DSO: Improving Mapping of Direct Sparse Odometry Using CNN-Based Single-Image Depth Estimation
by Erick P. Herrera-Granda, Juan C. Torres-Cantero, Israel D. Herrera-Granda, José F. Lucio-Naranjo, Andrés Rosales, Javier Revelo-Fuelagán and Diego H. Peluffo-Ordóñez
Mathematics 2025, 13(20), 3330; https://doi.org/10.3390/math13203330 - 19 Oct 2025
Cited by 2 | Viewed by 2754
Abstract
In recent years, SLAM, visual odometry, and structure-from-motion approaches have widely addressed the problems of 3D reconstruction and ego-motion estimation. Of the many input modalities that can be used to solve these ill-posed problems, the pure visual alternative using a single monocular RGB [...] Read more.
In recent years, SLAM, visual odometry, and structure-from-motion approaches have widely addressed the problems of 3D reconstruction and ego-motion estimation. Of the many input modalities that can be used to solve these ill-posed problems, the pure visual alternative using a single monocular RGB camera has attracted the attention of multiple researchers due to its low cost and widespread availability in handheld devices. One of the best proposals currently available is the Direct Sparse Odometry (DSO) system, which has demonstrated the ability to accurately recover trajectories and depth maps using monocular sequences as the only source of information. Given the impressive advances in single-image depth estimation using neural networks, this work proposes an extension of the DSO system, named DeepDSO. DeepDSO effectively integrates the state-of-the-art NeW CRF neural network as a depth estimation module, providing depth prior information for each candidate point. This reduces the point search interval over the epipolar line. This integration improves the DSO algorithm’s depth point initialization and allows each proposed point to converge faster to its true depth. Experimentation carried out in the TUM-Mono dataset demonstrated that adding the neural network depth estimation module to the DSO pipeline significantly reduced rotation, translation, scale, start-segment alignment, end-segment alignment, and RMSE errors. Full article
(This article belongs to the Section E1: Mathematics and Computer Science)
Show Figures

Figure 1

10 pages, 2952 KB  
Article
Weakly Supervised Monocular Fisheye Camera Distance Estimation with Segmentation Constraints
by Zhihao Zhang and Xuejun Yang
Electronics 2025, 14(17), 3429; https://doi.org/10.3390/electronics14173429 - 28 Aug 2025
Cited by 1 | Viewed by 1259
Abstract
Monocular fisheye camera distance estimation is a crucial visual perception task for autonomous driving. Due to the practical challenges of acquiring precise depth annotations, existing self-supervised methods usually consist of a monocular distance model and an ego-motion predictor with the goal of minimizing [...] Read more.
Monocular fisheye camera distance estimation is a crucial visual perception task for autonomous driving. Due to the practical challenges of acquiring precise depth annotations, existing self-supervised methods usually consist of a monocular distance model and an ego-motion predictor with the goal of minimizing a reconstruction matching loss. However, they suffer from inaccurate distance estimation in low-texture regions, especially road surfaces. In this paper, we introduce a weakly supervised learning strategy that incorporates semantic segmentation, instance segmentation, and optical flow as additional sources of supervision. In addition to the self-supervised reconstruction loss, we introduce a road surface flatness loss, an instance smoothness loss, and an optical flow loss to enhance the accuracy of distance estimation. We evaluate the proposed method on the WoodScape and SynWoodScape datasets, and it outperforms the self-supervised monocular baseline, FisheyeDistanceNet. Full article
Show Figures

Figure 1

32 pages, 14809 KB  
Article
A Comparison of Monocular Visual SLAM and Visual Odometry Methods Applied to 3D Reconstruction
by Erick P. Herrera-Granda, Juan C. Torres-Cantero, Andrés Rosales and Diego H. Peluffo-Ordóñez
Appl. Sci. 2023, 13(15), 8837; https://doi.org/10.3390/app13158837 - 31 Jul 2023
Cited by 11 | Viewed by 12174
Abstract
Pure monocular 3D reconstruction is a complex problem that has attracted the research community’s interest due to the affordability and availability of RGB sensors. SLAM, VO, and SFM are disciplines formulated to solve the 3D reconstruction problem and estimate the camera’s ego-motion; so, [...] Read more.
Pure monocular 3D reconstruction is a complex problem that has attracted the research community’s interest due to the affordability and availability of RGB sensors. SLAM, VO, and SFM are disciplines formulated to solve the 3D reconstruction problem and estimate the camera’s ego-motion; so, many methods have been proposed. However, most of these methods have not been evaluated on large datasets and under various motion patterns, have not been tested under the same metrics, and most of them have not been evaluated following a taxonomy, making their comparison and selection difficult. In this research, we performed a comparison of ten publicly available SLAM and VO methods following a taxonomy, including one method for each category of the primary taxonomy, three machine-learning-based methods, and two updates of the best methods to identify the advantages and limitations of each category of the taxonomy and test whether the addition of machine learning or updates on those methods improved them significantly. Thus, we evaluated each algorithm using the TUM-Mono dataset and benchmark, and we performed an inferential statistical analysis to identify the significant differences through its metrics. The results determined that the sparse-direct methods significantly outperformed the rest of the taxonomy, and fusing them with machine learning techniques significantly enhanced the geometric-based methods’ performance from different perspectives. Full article
(This article belongs to the Topic Computer Vision and Image Processing)
Show Figures

Figure 1

11 pages, 2973 KB  
Article
StereoVO: Learning Stereo Visual Odometry Approach Based on Optical Flow and Depth Information
by Chao Duan, Steffen Junginger, Kerstin Thurow and Hui Liu
Appl. Sci. 2023, 13(10), 5842; https://doi.org/10.3390/app13105842 - 9 May 2023
Cited by 4 | Viewed by 5246
Abstract
We present a novel stereo visual odometry (VO) model that utilizes both optical flow and depth information. While some existing monocular VO methods demonstrate superior performance, they require extra frames or information to initialize the model in order to obtain absolute scale, and [...] Read more.
We present a novel stereo visual odometry (VO) model that utilizes both optical flow and depth information. While some existing monocular VO methods demonstrate superior performance, they require extra frames or information to initialize the model in order to obtain absolute scale, and they do not take into account moving objects. To address these issues, we have combined optical flow and depth information to estimate ego-motion and proposed a framework for stereo VO using deep neural networks. The model simultaneously generates optical flow and depth information outputs from sequential stereo RGB image pairs, which are then fed into the pose estimation network to achieve final motion estimation. Our experiments have demonstrated that our combination of optical flow and depth information improves the accuracy of camera pose estimation. Our method outperforms existing learning-based and monocular geometry-based methods on the KITTI odometry dataset. Furthermore, we have achieved real-time performance, making our method both effective and efficient. Full article
(This article belongs to the Special Issue Robotics and Industrial Automation: From Methods to Applications)
Show Figures

Figure 1

14 pages, 5182 KB  
Article
Unsupervised Learning of Monocular Depth and Ego-Motion with Optical Flow Features and Multiple Constraints
by Baigan Zhao, Yingping Huang, Wenyan Ci and Xing Hu
Sensors 2022, 22(4), 1383; https://doi.org/10.3390/s22041383 - 11 Feb 2022
Cited by 9 | Viewed by 3839
Abstract
This paper proposes a novel unsupervised learning framework for depth recovery and camera ego-motion estimation from monocular video. The framework exploits the optical flow (OF) property to jointly train the depth and the ego-motion models. Unlike the existing unsupervised methods, our method extracts [...] Read more.
This paper proposes a novel unsupervised learning framework for depth recovery and camera ego-motion estimation from monocular video. The framework exploits the optical flow (OF) property to jointly train the depth and the ego-motion models. Unlike the existing unsupervised methods, our method extracts the features from the optical flow rather than from the raw RGB images, thereby enhancing unsupervised learning. In addition, we exploit the forward-backward consistency check of the optical flow to generate a mask of the invalid region in the image, and accordingly, eliminate the outlier regions such as occlusion regions and moving objects for the learning. Furthermore, in addition to using view synthesis as a supervised signal, we impose additional loss functions, including optical flow consistency loss and depth consistency loss, as additional supervision signals on the valid image region to further enhance the training of the models. Substantial experiments on multiple benchmark datasets demonstrate that our method outperforms other unsupervised methods. Full article
(This article belongs to the Section Sensing and Imaging)
Show Figures

Figure 1

16 pages, 7604 KB  
Article
Uncertainty Estimation of Dense Optical Flow for Robust Visual Navigation
by Yonhon Ng, Hongdong Li and Jonghyuk Kim
Sensors 2021, 21(22), 7603; https://doi.org/10.3390/s21227603 - 16 Nov 2021
Cited by 5 | Viewed by 3963
Abstract
This paper presents a novel dense optical-flow algorithm to solve the monocular simultaneous localisation and mapping (SLAM) problem for ground or aerial robots. Dense optical flow can effectively provide the ego-motion of the vehicle while enabling collision avoidance with the potential obstacles. Existing [...] Read more.
This paper presents a novel dense optical-flow algorithm to solve the monocular simultaneous localisation and mapping (SLAM) problem for ground or aerial robots. Dense optical flow can effectively provide the ego-motion of the vehicle while enabling collision avoidance with the potential obstacles. Existing research has not fully utilised the uncertainty of the optical flow—at most, an isotropic Gaussian density model has been used. We estimate the full uncertainty of the optical flow and propose a new eight-point algorithm based on the statistical Mahalanobis distance. Combined with the pose-graph optimisation, the proposed method demonstrates enhanced robustness and accuracy for the public autonomous car dataset (KITTI) and aerial monocular dataset. Full article
(This article belongs to the Special Issue Multi-Radio and/or Multi-Sensor Integrated Navigation System)
Show Figures

Figure 1

13 pages, 1945 KB  
Communication
Leveraging Deep Learning for Visual Odometry Using Optical Flow
by Tejas Pandey, Dexmont Pena, Jonathan Byrne and David Moloney
Sensors 2021, 21(4), 1313; https://doi.org/10.3390/s21041313 - 12 Feb 2021
Cited by 30 | Viewed by 8728
Abstract
In this paper, we study deep learning approaches for monocular visual odometry (VO). Deep learning solutions have shown to be effective in VO applications, replacing the need for highly engineered steps, such as feature extraction and outlier rejection in a traditional pipeline. We [...] Read more.
In this paper, we study deep learning approaches for monocular visual odometry (VO). Deep learning solutions have shown to be effective in VO applications, replacing the need for highly engineered steps, such as feature extraction and outlier rejection in a traditional pipeline. We propose a new architecture combining ego-motion estimation and sequence-based learning using deep neural networks. We estimate camera motion from optical flow using Convolutional Neural Networks (CNNs) and model the motion dynamics using Recurrent Neural Networks (RNNs). The network outputs the relative 6-DOF camera poses for a sequence, and implicitly learns the absolute scale without the need for camera intrinsics. The entire trajectory is then integrated without any post-calibration. We evaluate the proposed method on the KITTI dataset and compare it with traditional and other deep learning approaches in the literature. Full article
(This article belongs to the Section Optical Sensors)
Show Figures

Figure 1

15 pages, 4427 KB  
Article
Unsupervised Learning of Depth and Camera Pose with Feature Map Warping
by Ente Guo, Zhifeng Chen, Yanlin Zhou and Dapeng Oliver Wu
Sensors 2021, 21(3), 923; https://doi.org/10.3390/s21030923 - 30 Jan 2021
Cited by 4 | Viewed by 4987
Abstract
Estimating the depth of image and egomotion of agent are important for autonomous and robot in understanding the surrounding environment and avoiding collision. Most existing unsupervised methods estimate depth and camera egomotion by minimizing photometric error between adjacent frames. However, the photometric consistency [...] Read more.
Estimating the depth of image and egomotion of agent are important for autonomous and robot in understanding the surrounding environment and avoiding collision. Most existing unsupervised methods estimate depth and camera egomotion by minimizing photometric error between adjacent frames. However, the photometric consistency sometimes does not meet the real situation, such as brightness change, moving objects and occlusion. To reduce the influence of brightness change, we propose a feature pyramid matching loss (FPML) which captures the trainable feature error between a current and the adjacent frames and therefore it is more robust than photometric error. In addition, we propose the occlusion-aware mask (OAM) network which can indicate occlusion according to change of masks to improve estimation accuracy of depth and camera pose. The experimental results verify that the proposed unsupervised approach is highly competitive against the state-of-the-art methods, both qualitatively and quantitatively. Specifically, our method reduces absolute relative error (Abs Rel) by 0.017–0.088. Full article
(This article belongs to the Special Issue Neural Networks and Deep Learning in Image Sensing)
Show Figures

Figure 1

15 pages, 3756 KB  
Article
Ego-Motion Estimation Using Recurrent Convolutional Neural Networks through Optical Flow Learning
by Baigan Zhao, Yingping Huang, Hongjian Wei and Xing Hu
Electronics 2021, 10(3), 222; https://doi.org/10.3390/electronics10030222 - 20 Jan 2021
Cited by 19 | Viewed by 8230
Abstract
Visual odometry (VO) refers to incremental estimation of the motion state of an agent (e.g., vehicle and robot) by using image information, and is a key component of modern localization and navigation systems. Addressing the monocular VO problem, this paper presents a novel [...] Read more.
Visual odometry (VO) refers to incremental estimation of the motion state of an agent (e.g., vehicle and robot) by using image information, and is a key component of modern localization and navigation systems. Addressing the monocular VO problem, this paper presents a novel end-to-end network for estimation of camera ego-motion. The network learns the latent subspace of optical flow (OF) and models sequential dynamics so that the motion estimation is constrained by the relations between sequential images. We compute the OF field of consecutive images and extract the latent OF representation in a self-encoding manner. A Recurrent Neural Network is then followed to examine the OF changes, i.e., to conduct sequential learning. The extracted sequential OF subspace is used to compute the regression of the 6-dimensional pose vector. We derive three models with different network structures and different training schemes: LS-CNN-VO, LS-AE-VO, and LS-RCNN-VO. Particularly, we separately train the encoder in an unsupervised manner. By this means, we avoid non-convergence during the training of the whole network and allow more generalized and effective feature representation. Substantial experiments have been conducted on KITTI and Malaga datasets, and the results demonstrate that our LS-RCNN-VO outperforms the existing learning-based VO approaches. Full article
(This article belongs to the Special Issue Autonomous Vehicles Technology)
Show Figures

Figure 1

23 pages, 4246 KB  
Article
Joint Unsupervised Learning of Depth, Pose, Ground Normal Vector and Ground Segmentation by a Monocular Camera Sensor
by Lu Xiong, Yongkun Wen, Yuyao Huang, Junqiao Zhao and Wei Tian
Sensors 2020, 20(13), 3737; https://doi.org/10.3390/s20133737 - 3 Jul 2020
Cited by 6 | Viewed by 4816
Abstract
We propose a completely unsupervised approach to simultaneously estimate scene depth, ego-pose, ground segmentation and ground normal vector from only monocular RGB video sequences. In our approach, estimation for different scene structures can mutually benefit each other by the joint optimization. Specifically, we [...] Read more.
We propose a completely unsupervised approach to simultaneously estimate scene depth, ego-pose, ground segmentation and ground normal vector from only monocular RGB video sequences. In our approach, estimation for different scene structures can mutually benefit each other by the joint optimization. Specifically, we use the mutual information loss to pre-train the ground segmentation network and before adding the corresponding self-learning label obtained by a geometric method. By using the static nature of the ground and its normal vector, the scene depth and ego-motion can be efficiently learned by the self-supervised learning procedure. Extensive experimental results on both Cityscapes and KITTI benchmark demonstrate the significant improvement on the estimation accuracy for both scene depth and ego-pose by our approach. We also achieve an average error of about 3 for estimated ground normal vectors. By deploying our proposed geometric constraints, the IOU accuracy of unsupervised ground segmentation is increased by 35% on the Cityscapes dataset. Full article
(This article belongs to the Special Issue Camera as a Smart-Sensor (CaaSS))
Show Figures

Figure 1

18 pages, 22280 KB  
Article
DM-SLAM: A Feature-Based SLAM System for Rigid Dynamic Scenes
by Junhao Cheng, Zhi Wang, Hongyan Zhou, Li Li and Jian Yao
ISPRS Int. J. Geo-Inf. 2020, 9(4), 202; https://doi.org/10.3390/ijgi9040202 - 27 Mar 2020
Cited by 85 | Viewed by 7951
Abstract
Most Simultaneous Localization and Mapping (SLAM) methods assume that environments are static. Such a strong assumption limits the application of most visual SLAM systems. The dynamic objects will cause many wrong data associations during the SLAM process. To address this problem, a novel [...] Read more.
Most Simultaneous Localization and Mapping (SLAM) methods assume that environments are static. Such a strong assumption limits the application of most visual SLAM systems. The dynamic objects will cause many wrong data associations during the SLAM process. To address this problem, a novel visual SLAM method that follows the pipeline of feature-based methods called DM-SLAM is proposed in this paper. DM-SLAM combines an instance segmentation network with optical flow information to improve the location accuracy in dynamic environments, which supports monocular, stereo, and RGB-D sensors. It consists of four modules: semantic segmentation, ego-motion estimation, dynamic point detection and a feature-based SLAM framework. The semantic segmentation module obtains pixel-wise segmentation results of potentially dynamic objects, and the ego-motion estimation module calculates the initial pose. In the third module, two different strategies are presented to detect dynamic feature points for RGB-D/stereo and monocular cases. In the first case, the feature points with depth information are reprojected to the current frame. The reprojection offset vectors are used to distinguish the dynamic points. In the other case, we utilize the epipolar constraint to accomplish this task. Furthermore, the static feature points left are fed into the fourth module. The experimental results on the public TUM and KITTI datasets demonstrate that DM-SLAM outperforms the standard visual SLAM baselines in terms of accuracy in highly dynamic environments. Full article
(This article belongs to the Special Issue 3D Indoor Mapping and Modelling)
Show Figures

Graphical abstract

24 pages, 9593 KB  
Article
Forward and Backward Visual Fusion Approach to Motion Estimation with High Robustness and Low Cost
by Ke Wang, Xin Huang, JunLan Chen, Chuan Cao, Zhoubing Xiong and Long Chen
Remote Sens. 2019, 11(18), 2139; https://doi.org/10.3390/rs11182139 - 13 Sep 2019
Cited by 11 | Viewed by 5503
Abstract
We present a novel low-cost visual odometry method of estimating the ego-motion (self-motion) for ground vehicles by detecting the changes that motion induces on the images. Different from traditional localization methods that use differential global positioning system (GPS), precise inertial measurement unit (IMU) [...] Read more.
We present a novel low-cost visual odometry method of estimating the ego-motion (self-motion) for ground vehicles by detecting the changes that motion induces on the images. Different from traditional localization methods that use differential global positioning system (GPS), precise inertial measurement unit (IMU) or 3D Lidar, the proposed method only leverage data from inexpensive visual sensors of forward and backward onboard cameras. Starting with the spatial-temporal synchronization, the scale factor of backward monocular visual odometry was estimated based on the MSE optimization method in a sliding window. Then, in trajectory estimation, an improved two-layers Kalman filter was proposed including orientation fusion and position fusion. Where, in the orientation fusion step, we utilized the trajectory error space represented by unit quaternion as the state of the filter. The resulting system enables high-accuracy, low-cost ego-pose estimation, along with providing robustness capability of handing camera module degradation by automatic reduce the confidence of failed sensor in the fusion pipeline. Therefore, it can operate in the presence of complex and highly dynamic motion such as enter-in-and-out tunnel entrance, texture-less, illumination change environments, bumpy road and even one of the cameras fails. The experiments carried out in this paper have proved that our algorithm can achieve the best performance on evaluation indexes of average in distance (AED), average in X direction (AEX), average in Y direction (AEY), and root mean square error (RMSE) compared to other state-of-the-art algorithms, which indicates that the output results of our approach is superior to other methods. Full article
(This article belongs to the Special Issue Mobile Mapping Technologies)
Show Figures

Graphical abstract

20 pages, 3878 KB  
Article
A 3D Relative-Motion Context Constraint-Based MAP Solution for Multiple-Object Tracking Problems
by Zhongli Wang, Litong Fan and Baigen Cai
Sensors 2018, 18(7), 2363; https://doi.org/10.3390/s18072363 - 20 Jul 2018
Cited by 2 | Viewed by 5379
Abstract
Multi-object tracking (MOT), especially by using a moving monocular camera, is a very challenging task in the field of visual object tracking. To tackle this problem, the traditional tracking-by-detection-based method is heavily dependent on detection results. Occlusion and mis-detections will often lead to [...] Read more.
Multi-object tracking (MOT), especially by using a moving monocular camera, is a very challenging task in the field of visual object tracking. To tackle this problem, the traditional tracking-by-detection-based method is heavily dependent on detection results. Occlusion and mis-detections will often lead to tracklets or drifting. In this paper, the tasks of MOT and camera motion estimation are formulated as finding a maximum a posteriori (MAP) solution of joint probability and synchronously solved in a unified framework. To improve performance, we incorporate the three-dimensional (3D) relative-motion model into a sequential Bayesian framework to track multiple objects and the camera’s ego-motion estimation. A 3D relative-motion model that describes spatial relations among objects is exploited for predicting object states robustly and recovering objects when occlusion and mis-detections occur. Reversible jump Markov chain Monte Carlo (RJMCMC) particle filtering is applied to solve the posteriori estimation problem. Both quantitative and qualitative experiments with benchmark datasets and video collected on campus were conducted, which confirms that the proposed method is outperformed in many evaluation metrics. Full article
(This article belongs to the Special Issue Sensors Signal Processing and Visual Computing)
Show Figures

Figure 1

25 pages, 8746 KB  
Article
Adaptive Absolute Ego-Motion Estimation Using Wearable Visual-Inertial Sensors for Indoor Positioning
by Ya Tian, Zhe Chen, Shouyin Lu and Jindong Tan
Micromachines 2018, 9(3), 113; https://doi.org/10.3390/mi9030113 - 6 Mar 2018
Cited by 4 | Viewed by 4877
Abstract
This paper proposes an adaptive absolute ego-motion estimation method using wearable visual-inertial sensors for indoor positioning. We introduce a wearable visual-inertial device to estimate not only the camera ego-motion, but also the 3D motion of the moving object in dynamic environments. Firstly, a [...] Read more.
This paper proposes an adaptive absolute ego-motion estimation method using wearable visual-inertial sensors for indoor positioning. We introduce a wearable visual-inertial device to estimate not only the camera ego-motion, but also the 3D motion of the moving object in dynamic environments. Firstly, a novel method dynamic scene segmentation is proposed using two visual geometry constraints with the help of inertial sensors. Moreover, this paper introduces a concept of “virtual camera” to consider the motion area related to each moving object as if a static object were viewed by a “virtual camera”. We therefore derive the 3D moving object’s motion from the motions for the real and virtual camera because the virtual camera’s motion is actually the combined motion of both the real camera and the moving object. In addition, a multi-rate linear Kalman-filter (MR-LKF) as our previous work was selected to solve both the problem of scale ambiguity in monocular camera tracking and the different sampling frequencies of visual and inertial sensors. The performance of the proposed method is evaluated by simulation studies and practical experiments performed in both static and dynamic environments. The results show the method’s robustness and effectiveness compared with the results from a Pioneer robot as the ground truth. Full article
(This article belongs to the Special Issue MEMS Technology for Biomedical Imaging Applications)
Show Figures

Figure 1

Back to TopTop