Feature-PLPD-Aided Visual–Inertial Odometry for Low-Cost Embedded Systems
Abstract
1. Introduction
- The integration and extension of our previous hardware-aware visual processing developments into complete Feature-PLPD VIO pipelines, addressing the absence of inertial fusion and full-system deployment in the earlier studies.
- A resource-aware ESEKF-based loosely coupled fusion strategy, in which inertial prediction is performed at the camera rate to improve VO robustness while limiting the computational and memory overhead of high-frequency IMU processing.
- Two complementary end-to-end embedded implementations targeting different resource constraints: a high-performance GPU-based stereo VIO and a low-cost FPGA-based RGB-D-assisted monocular VIO with depth scale correction.
- A progressive experimental analysis from visual-only odometry and previously developed processing blocks to complete VIO implementations, together with dataset-based evaluation and on-the-fly robotic deployment of the stereo system.
- A joint quantitative evaluation of trajectory accuracy, throughput, and energy consumption to characterize the trade-offs involved in scaling the framework toward low-cost embedded motion estimation.
2. Related Work
2.1. Visual Odometry and Feature-Based Front-Ends
2.2. Visual–Inertial Odometry: Tightly and Loosely Coupled Approaches
2.3. Embedded Implementations of VO and VIO
3. Method
3.1. Feature-PLPD Tracking Overview
- The Feature-PLPD Extraction thread begins by using our Feature-Point (corners) and Line Point (segments) Detector [21] for initialization, ensuring enough long-line segments and corners. If features are insufficient, re-extraction occurs. The thread employs the Multi-level Edge Detector (MED) to create a pre-refined feature map of potential corners and edglets, then identifies line segments (FLS) by using a refined random anchor and orientation to grow the segment. Simultaneously, a corner refinement (CR) process identifies the true corners.
- The Feature-PLPD Tracker thread utilizes the sparse KLT optical flow algorithm (OF) [48] to estimate feature locations between frames using the Taylor equation. Due to its sensitivity to abrupt movement and noise, which can lead to point loss, we implemented a flow correction method (FC) [22] based on a line-fitting model that stabilizes line points and removes outliers through anchor and line-segment properties. Unlike line segments, corners lack a defined model, so we use a corner outlier rejection mechanism (COR) based on circular matching [30] or geometric modeling via RANSAC [49], depending on whether the camera is monocular or stereo, as detailed in Section 3.3 and Section 3.4.
3.2. Proposed Feature-PLPD VIO Back-End
3.2.1. ESEKF Kinematic Model in Discrete Time
- is the position of the mobile platform in the global frame;
- is the velocity in the global frame;
- is the orientation quaternion from the body frame to the global frame;
- is the accelerometer bias;
- is the gyroscope bias.
- and the different equations of the nominal state kinematics can be written as
3.2.2. ESEKF Algorithm
Prediction Step
- Input: Previous corrected estimate , ; IMU measurement .
- Output: Predicted estimate , .
- Nominal State Propagation:
- Covariance Propagation:
Correction Step (On Measurement)
- Input: Predicted estimate , ; position measurement .
- Output: Corrected estimate , .
- Kalman Gain:
- Error-State Correction:
- Inject Error into Nominal State:
- Covariance Update:
3.3. End-to-End Embedded Feature-PLPD Stereo VIO on GPU-Based Heterogeneous Architecture
- Extraction: The process begins with the extraction of a set of features (corners and line-segment points ) from the previous left camera frame.
- Tracking: The forward tracking of line-segment points is conducted using the proposed OF-based Flow correction method and ensures their stereo correspondence in the prior right camera. Simultaneously, corners are propagated in a circular manner across the four frames (previous left camera → previous right camera → current right camera → current left camera, and back to the previous left camera) to facilitate outlier rejection.
- Triangulation: The aforementioned step results in feature correspondences between the previous left and right frames, which enables 3D projection through triangulation based on intrinsic camera parameters.
- PnP solver: The feature tracking from Step 2, alongside the triangulation results from Step 3, provides the necessary input for the PnP solver to minimize the 3D-2D projection error within the current left camera. This involves comparing the current 2D features with their re-projected counterparts (3D to 2D). The PnP solver, supported by a complementary RANSAC method for outlier rejection, effectively resolves the optimization problem and estimates the relative pose between previous and current frames (). Then, the current position measurement can be computed by Equation (10).
- IMU propagation: Concurrently, the IMU data is propagated through the adopted ESEKF-based observer in the back-end across the sequence frames to refine the pose estimation. This approach corrects accumulated drift by updating the error states related to position, orientation, and velocity, thereby ensuring the trajectory remains consistent over time.
- Repeat from step 1: The tracking process is sustained as long as the number of extracted line-segment points remains sufficient. Should the number of features diminish, a re-extraction process is initiated, preserving the previously robust tracked features and repeating the process from step 1.
3.4. Low-Cost Embedded Feature-PLPD Monocular VIO with Depth Scale Correction on FPGA-Based Heterogeneous Architecture
- Extraction: We implement a stringent extraction and selection process for a set of features (corners and line-segment points) using the proposed Feature-PLPD from the previous camera frame. This solution primarily focuses on line-segment points rather than corners, as we have already fitted a model for controlling the tracking of these points in the next step.
- Tracking: The forward feature tracking step is crucial for system accuracy. It is important to note that tracking line segments using the proposed OF-based flow correction method effectively maintains the alignment of line-segment points. In contrast, corner tracking is discouraged since corners do not have a fitted model as segments do and are only used in the absence of segments in certain scenes. Therefore, we adopt a strict corner outlier rejection process based on RANSAC, utilizing the geometric information obtained in the next step.
- Essential matrix computation: We normalize the tracked features to compute the essential matrix between two frames.
- SVD: We decompose the essential matrix using SVD to recover the unscaled estimated relative pose, denoted as (refer to the orange parts in Figure 5).
- Depth scale correction and update: To address the scale ambiguity inherent in monocular VO, we apply metric scale correction using depth measurements from an RGB-D sensor, similarly to the approaches described in [25,51]. The scale estimation relies on the feature correspondences retained after the tracking and geometric outlier-rejection steps described above. Before computing the scale, invalid or unreliable depth measurements are discarded, with the upper depth bound limited to 3 m, consistently with the reliable operating range considered for the RealSense D435i [52]. The scale factor is then estimated from the remaining valid correspondences by comparing their 3D norms, following [53,54]:Here, N denotes the number of valid correspondences retained after the tracking, geometric outlier-rejection, and depth-validity checks, is the measured 3D position of feature in the current frame using its depth , and is the corresponding 3D position predicted from the previous frame using the unscaled monocular motion .To reduce frame-to-frame fluctuations, the estimated scale factor is temporally smoothed using the four previous scale estimates through median filtering and interpolation. The resulting scale factor is then applied multiplicatively to the unscaled translation, while the relative rotation remains unchanged:Although metric depth is available, it is used here only to recover the translation scale, thereby preserving the original Feature-PLPD motion estimation based on 2D correspondences and limiting the dependence on depth measurements. Accordingly, this configuration is referred to as RGB-D-assisted monocular VIO rather than conventional monocular VIO, since depth provides an external metric-scale reference instead of relying exclusively on visual–inertial information for scale recovery. This assistance requires additional depth access and validity checks, while preserving the underlying monocular Feature-PLPD motion-estimation structure. The reported results should therefore be interpreted within this RGB-D-assisted setting when compared with conventional monocular VIO systems. In practice, this formulation also provided more stable relative-rotation estimates than direct RGB-D PnP, which exhibited orientation ambiguities for some feature configurations.The resulting scaled relative pose is then used to compute the current visual measurement according to Equation (10). Although the scale factor could in principle be estimated from inertial measurements, the present implementation favors depth-based visual information. Depth observations provide a direct metric reference for the tracked features in the controlled indoor environment, whereas scale recovery from the low-rate IMU measurements would be more sensitive to noise, bias, and accumulated drift.
- IMU propagation: Concurrently, the IMU data is propagated through the ESEKF-based observer in the back-end across various frames to update the state vector, which includes position, orientation, and velocity.
- Repeat from step 1: The tracking process continues as long as there are sufficient extracted line-segment points. If the number of features diminishes, a re-extraction process is initiated, preserving previously robust tracked features and restarting from step 1.
4. Experimental Results
4.1. Experimental Setup
4.1.1. Platform Description
4.1.2. Datasets
- Outdoor KITTI stereo inertial dataset: This dataset provides stereo image sequences with a resolution of and IMU measurements extracted from the corresponding raw KITTI data. For each image timestamp, the temporally nearest IMU measurement is selected from the raw sequence, without averaging or interpolation. The selected IMU measurements are transformed into the camera coordinate frame using the corresponding extrinsic calibration provided with the KITTI dataset, ensuring consistency with the visual measurements used by the VIO framework. This nearest-neighbor association discards the intermediate high-rate inertial measurements and may therefore reduce the representation of fast motion dynamics; this limitation is consistent with the camera-rate fusion trade-off discussed in Section 3.2. A controlled comparison with full-rate IMU propagation or pre-integration requires a temporally consistent high-rate inertial stream aligned with the evaluated odometry sequences. For the considered KITTI sequences, reconstructing such a stream from the corresponding raw inertial data is subject to data-continuity and synchronization constraints, which prevent the effect of the propagation rate from being isolated reliably. Consequently, the present evaluation is restricted to the synchronized camera-rate configuration, and no equivalence with full-rate inertial propagation or pre-integration is claimed.The OXTS data are processed using the official KITTI development tools and calibration parameters. Although the inertial measurements used by the VIO estimator and the ground-truth trajectory originate from the same OXTS unit, only the inertial measurements are provided as inputs to the proposed VIO framework, while the GPS-derived trajectory is used exclusively as ground truth for evaluation and is not involved in the state estimation. Ground truth is available for both translation and rotation, enabling a complete evaluation of the estimated trajectory in outdoor conditions. It should be noted that the OXTS RT3003 used in KITTI is a higher-grade navigation system and is therefore not representative of the low-cost IMUs primarily targeted in this work. This limitation is complemented by the indoor experiments using the integrated IMU of the RealSense D435i, which provide an evaluation with a more representative low-cost sensing configuration. The rectified sequences and their corresponding raw sequences are:
- Sequence 00: 2011_10_03_drive_0027;
- Sequence 05: 2011_09_30_drive_0018;
- Sequence 06: 2011_09_30_drive_0020;
- Sequence 10: 2011_09_30_drive_0034.
These sequences were selected to limit the confounding influence of highly dynamic scenes and to provide consistent visual, inertial, and ground-truth data for evaluating the effect of visual–inertial fusion under identical conditions. This controlled selection enables the analysis to focus on the proposed embedded VO/VIO framework and its accuracy–runtime–energy trade-offs, rather than on dynamic-object handling, which is beyond the scope of this work. Consequently, the present KITTI evaluation should be interpreted as a controlled assessment of the proposed embedded VIO architecture rather than a comprehensive evaluation under highly dynamic or visually challenging conditions. - Indoor VICON stereo inertial dataset: Captured using the RealSense D435i (Intel Corporation, Santa Clara, CA, USA) in stereo mode and mounted on the Scout Mini R&D Kit, the dataset provides synchronized stereo images and inertial measurements for real-time robotic evaluation. Ground-truth translational motion is captured with high precision using the VICON system. The Scout Mini platform, illustrated in Figure 7, integrates an NVIDIA Jetson Xavier NX and enables on-the-fly processing of stereo image sequences at a resolution of .
- Indoor VICON RGB-D dataset: Also captured using the RealSense D435i, this dataset uses the RGB-D configuration at a resolution of . It is used for the offline evaluation of the RGB-D-assisted monocular VIO implementation. Ground truth is obtained from the VICON system, with only translational data available.
4.1.3. Evaluation Metrics
4.2. End-to-End Embedded Feature-PLPD Stereo VIO
4.2.1. Runtime Analysis
4.2.2. Performance Evaluation
4.3. Low-Cost Embedded Feature-PLPD RGB-D-Assisted Monocular VIO
4.3.1. Runtime Analysis
4.3.2. Performance Evaluation
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| A3 | Algorithm Architecture Adequacy |
| ATE | Absolute Trajectory Error |
| BRIEF | Binary Robust Independent Elementary Features |
| COR | Corner Outlier Rejection |
| CR | Corner Refinement |
| CUDA | Compute Unified Device Architecture |
| DSO | Direct Sparse Odometry |
| EDLines | EDge Line-segment detector |
| ESEKF | Error-State Extended Kalman Filter |
| FAST | Features from Accelerated Segment Test |
| FC | Flow Correction |
| FLS | Find Line Segments |
| FPGA | Field Programmable Gate Array |
| FPS | Frame Per Second |
| GFT | GoodFeatureToTrack detector |
| GPU | Graphics Processing Unit |
| HOOFR | Hessian ORB - Overlapped FREAK |
| IMU | Inertial Measurement Unit |
| LBD | Line Band Descriptor |
| LiDAR | Light Detection and Ranging |
| LSD | Line-Segment Detector |
| MED | Multi-level Edge Detector |
| OF | Optical Flow |
| OpenCL | Open Computing Language |
| ORB | Oriented FAST and Rotated BRIEF Oriented FAST and Rotated BRIEF |
| PLPD | Point and Line Points Detection |
| RANSAC | Random Sample Consensus |
| RPE | Relative Pose Error |
| SLAM | Simultaneous Localization and Mapping |
| VIO | Visual–Inertial Odometry |
| VO | Visual Odometry |
References
- Nistér, D.; Naroditsky, O.; Bergen, J. Visual odometry. In Proceedings of the 2004 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2004. CVPR 2004; IEEE: New York, NY, USA, 2004; Volume 1. [Google Scholar]
- Thrun, S. Simultaneous localization and mapping. In Robotics and Cognitive Approaches to Spatial Mapping; Springer: Berlin/Heidelberg, Germany, 2008; pp. 13–41. [Google Scholar]
- Bao, Y.; Ding, T.; Huo, J.; Liu, Y.; Li, Y.; Li, W.; Gao, Y.; Luo, J. 3D Gaussian Splatting: Survey, Technologies, Challenges, and Opportunities. IEEE Trans. Circuits Syst. Video Technol. 2025, 35, 6832–6852. [Google Scholar] [CrossRef] [Scilit]
- Zhu, D.; Wang, Z.; Fan, X.; Chen, M.; Chen, J. Gaussian landmarks tracking-based real-time splatting reconstruction model. Image Vis. Comput. 2026, 166, 105869. [Google Scholar] [CrossRef] [Scilit]
- Wang, K.; Guo, J.; Chen, K.; Lu, J. An in-depth examination of SLAM methods: Challenges, advancements, and applications in complex scenes for autonomous driving. IEEE Trans. Intell. Transp. Syst. 2025, 26, 11066–11087. [Google Scholar] [CrossRef] [Scilit]
- Qin, T.; Li, P.; Shen, S. Vins-mono: A robust and versatile monocular visual-inertial state estimator. IEEE Trans. Robot. 2018, 34, 1004–1020. [Google Scholar] [CrossRef] [Scilit]
- Ozyoruk, K.B.; Gokceler, G.I.; Bobrow, T.L.; Coskun, G.; Incetan, K.; Almalioglu, Y.; Mahmood, F.; Curto, E.; Perdigoto, L.; Oliveira, M.; et al. EndoSLAM dataset and an unsupervised monocular visual odometry and depth estimation approach for endoscopic videos. Med. Image Anal. 2021, 71, 102058. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Cao, Q.; Deng, R.; Pan, Y.; Liu, R.; Chen, Y.; Gong, G.; Zou, J.; Yang, H.; Han, D. Robotic wireless capsule endoscopy: Recent advances and upcoming technologies. Nat. Commun. 2024, 15, 4597. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Campos, C.; Elvira, R.; Rodríguez, J.J.G.; Montiel, J.M.; Tardós, J.D. Orb-slam3: An accurate open-source library for visual, visual–inertial, and multimap slam. IEEE Trans. Robot. 2021, 37, 1874–1890. [Google Scholar] [CrossRef] [Scilit]
- Leutenegger, S.; Lynen, S.; Bosse, M.; Siegwart, R.; Furgale, P. Keyframe-based visual–inertial odometry using nonlinear optimization. Int. J. Robot. Res. 2015, 34, 314–334. [Google Scholar] [CrossRef] [Scilit]
- Pumarola, A.; Vakhitov, A.; Agudo, A.; Sanfeliu, A.; Moreno-Noguer, F. PL-SLAM: Real-time monocular visual SLAM with points and lines. In Proceedings of the 2017 IEEE International Conference on Robotics and Automation (ICRA); IEEE: New York, NY, USA, 2017; pp. 4503–4508. [Google Scholar]
- Sorel, Y. Massively parallel computing systems with real time constraints: The “Algorithm Architecture Adequation” methodology. In Proceedings of the First International Conference on Massively Parallel Computing Systems (MPCS), The Challenges of General-Purpose and Special-Purpose Computing; IEEE: New York, NY, USA, 1994; pp. 44–53. [Google Scholar]
- Sola, J. Quaternion kinematics for the error-state Kalman filter. arXiv 2017, arXiv:1711.02508. [Google Scholar]
- Al-Tawil, B.; Hempel, T.; Abdelrahman, A.; Al-Hamadi, A. A review of visual SLAM for robotics: Evolution, properties, and future applications. Front. Robot. AI 2024, 11, 1347985. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Picard, Q.; Chevobbe, S.; Darouich, M.; Didier, J.Y. A survey on real-time 3D scene reconstruction with SLAM methods in embedded systems. arXiv 2023, arXiv:2309.05349. [Google Scholar]
- Ondrúška, P.; Kohli, P.; Izadi, S. Mobilefusion: Real-time volumetric surface reconstruction and dense tracking on mobile phones. IEEE Trans. Vis. Comput. Graph. 2015, 21, 1251–1258. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lin, C.H.; Wang, W.Y.; Liu, S.H.; Hsu, C.C.; Chien, C.H. Heterogeneous implementation of a novel indirect visual odometry system. IEEE Access 2019, 7, 34631–34644. [Google Scholar] [CrossRef] [Scilit]
- Nagy, B.; Foehn, P.; Scaramuzza, D. Faster than FAST: GPU-accelerated frontend for high-speed VIO. In Proceedings of the 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); IEEE: New York, NY, USA, 2020; pp. 4361–4368. [Google Scholar]
- Zhou, F.; Cao, Y.; Wang, X. Fast and Resource-Efficient Hardware Implementation of Modified Line Segment Detector. IEEE Trans. Circuits Syst. Video Technol. 2018, 28, 3262–3273. [Google Scholar] [CrossRef] [Scilit]
- Liu, M.; Delbruck, T. EDFLOW: Event Driven Optical Flow Camera With Keypoint Detection and Adaptive Block Matching. IEEE Trans. Circuits Syst. Video Technol. 2022, 32, 5776–5789. [Google Scholar] [CrossRef] [Scilit]
- Mamri, A.; Hadri, A.E.; Benallegue, A. Feature-PLPD: Feature-Point and Line Points Detection for Real-Time Embedded Visual Odometry-Based Systems. IEEE Signal Process. Lett. 2025, 32, 1890–1894. [Google Scholar] [CrossRef] [Scilit]
- Mamri, A.; El Hadri, A.; Benallegue, A. Embedded Feature-Line Points Tracker for real-time Visual Odometry-based system. In Proceedings of the 2025 33rd Mediterranean Conference on Control and Automation (MED), Tangier, Morocco, 10–13 June 2025; pp. 514–519. [Google Scholar] [CrossRef] [Scilit]
- Mamri, A.; Hadri, A.E.; Benallegue, A. Hardware/Software Co-Design of Multi-Level Edge Detector on Low-Cost FPGA-Based Embedded Heterogeneous Architecture. IEEE Embed. Syst. Lett. 2025, 18, 99–102. [Google Scholar] [CrossRef] [Scilit]
- Nistér, D. An efficient solution to the five-point relative pose problem. IEEE Trans. Pattern Anal. Mach. Intell. 2004, 26, 756–770. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Mur-Artal, R.; Tardós, J.D. Orb-slam2: An open-source slam system for monocular, stereo, and rgb-d cameras. IEEE Trans. Robot. 2017, 33, 1255–1262. [Google Scholar] [CrossRef] [Scilit]
- Engel, J.; Koltun, V.; Cremers, D. Direct sparse odometry. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 40, 611–625. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- He, Y.; Zhao, J.; Guo, Y.; He, W.; Yuan, K. PL-VIO: Tightly-coupled monocular visual–inertial odometry using point and line features. Sensors 2018, 18, 1159. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhu, Y.; Jin, R.; Lou, T.s.; Zhao, L. PLD-VINS: RGBD visual-inertial SLAM with point and line features. Aerosp. Sci. Technol. 2021, 119, 107185. [Google Scholar] [CrossRef] [Scilit]
- Zhang, C.; Fang, Z.; Luo, X.; Liu, W. Accurate and robust visual SLAM with a novel ray-to-ray line measurement model. Image Vis. Comput. 2023, 140, 104837. [Google Scholar] [CrossRef] [Scilit]
- Cvišić, I.; Petrović, I. Stereo odometry based on careful feature selection and tracking. In Proceedings of the 2015 European Conference on Mobile Robots (ECMR), Lincoln, UK, 2–4 September 2015; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
- Lupton, T.; Sukkarieh, S. Visual-inertial-aided navigation for high-dynamic motion in built environments without initial conditions. IEEE Trans. Robot. 2011, 28, 61–76. [Google Scholar] [CrossRef] [Scilit]
- Weiss, S.; Scaramuzza, D.; Siegwart, R. Monocular-SLAM–based navigation for autonomous micro helicopters in GPS-denied environments. J. Field Robot. 2011, 28, 854–874. [Google Scholar] [CrossRef] [Scilit]
- Kelly, J.; Sukhatme, G.S. Visual-inertial sensor fusion: Localization, mapping and sensor-to-sensor self-calibration. Int. J. Robot. Res. 2011, 30, 56–79, Correction in Int. J. Robot. Res. 2017, 36, 1619. [Google Scholar] [CrossRef] [Scilit]
- Bloesch, M.; Omari, S.; Hutter, M.; Siegwart, R. Robust visual inertial odometry using a direct EKF-based approach. In Proceedings of the 2015 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); IEEE: New York, NY, USA, 2015; pp. 298–304. [Google Scholar]
- Nguyen, D.D.; Elouardi, A.; Florez, S.A.R.; Bouaziz, S. HOOFR SLAM system: An embedded vision SLAM algorithm and its hardware-software mapping-based intelligent vehicles applications. IEEE Trans. Intell. Transp. Syst. 2018, 20, 4103–4118. [Google Scholar] [CrossRef] [Scilit]
- El Bouazzaoui, I.; Chghaf, M.; Rodriguez, S.; Nguyen, D.D.; El Ouardi, A. An Extended HOOFR SLAM Algorithm Using IR-D Sensor Data for Outdoor Autonomous Vehicle Localization. J. Intell. Robot. Syst. 2023, 109, 56. [Google Scholar] [CrossRef] [Scilit]
- Alahi, A.; Ortiz, R.; Vandergheynst, P. Freak: Fast retina keypoint. In Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2012; pp. 510–517. [Google Scholar]
- Aldegheri, S.; Bombieri, N.; Bloisi, D.D.; Farinelli, A. Data flow ORB-SLAM for real-time performance on embedded GPU boards. In Proceedings of the 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); IEEE: New York, NY, USA, 2019; pp. 5370–5375. [Google Scholar]
- Ma, T.; Bai, N.; Shi, W.; Wu, X.; Wang, L.; Wu, T.; Zhao, C. Research on the application of visual SLAM in embedded GPU. Wirel. Commun. Mob. Comput. 2021, 2021, 6691262. [Google Scholar] [CrossRef] [Scilit]
- Wan, Z.; Yu, B.; Li, T.Y.; Tang, J.; Zhu, Y.; Wang, Y.; Raychowdhury, A.; Liu, S. A survey of fpga-based robotic computing. IEEE Circuits Syst. Mag. 2021, 21, 48–74. [Google Scholar] [CrossRef] [Scilit]
- Weberruss, J.; Kleeman, L.; Boland, D.; Drummond, T. FPGA acceleration of multilevel ORB feature extraction for computer vision. In Proceedings of the 2017 27th International Conference on Field Programmable Logic and Applications (FPL), Ghent, Belgium, 4–8 September 2017; pp. 1–8. [Google Scholar] [CrossRef] [Scilit]
- Stumpp, D.C.; Akolkar, H.; George, A.D.; Benosman, R.B. hARMS: A Hardware Acceleration Architecture for Real-Time Event-Based Optical Flow. IEEE Access 2022, 10, 58181–58198. [Google Scholar] [CrossRef] [Scilit]
- Schulz, V.H.; Bombardelli, F.G.; Todt, E. A Harris corner detector implementation in SoC-FPGA for visual SLAM. In Proceedings of the Robotics: 12th Latin American Robotics Symposium and Third Brazilian Symposium on Robotics, LARS 2015/SBR 2015, Uberlândia, Brazil, 28 October–1 November 2015; Revised Selected Papers 12; Springer: Cham, Switzerland, 2016; pp. 57–71. [Google Scholar]
- Nguyen, D.D.; El Ouardi, A.; Rodriguez, S.; Bouaziz, S. FPGA implementation of HOOFR bucketing extractor-based real-time embedded SLAM applications. J. Real-Time Image Process. 2021, 18, 525–538. [Google Scholar] [CrossRef] [Scilit]
- Shi, X.; Cao, L.; Wang, D.; Liu, L.; You, G.; Liu, S.; Wang, C. HERO: Accelerating autonomous robotic tasks with FPGA. In Proceedings of the 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); IEEE: New York, NY, USA, 2018; pp. 7766–7772. [Google Scholar]
- Abouzahir, M.; Elouardi, A.; Latif, R.; Bouaziz, S.; Tajer, A. Embedding SLAM algorithms: Has it come of age? Robot. Auton. Syst. 2018, 100, 14–26. [Google Scholar] [CrossRef] [Scilit]
- Eisoldt, M.; Gaal, J.; Wiemann, T.; Flottmann, M.; Rothmann, M.; Tassemeier, M.; Porrmann, M. A fully integrated system for hardware-accelerated TSDF SLAM with LiDAR sensors (HATSDF SLAM). Robot. Auton. Syst. 2022, 156, 104205. [Google Scholar] [CrossRef] [Scilit]
- Bouguet, J.Y. Pyramidal implementation of the affine lucas kanade feature tracker description of the algorithm. Intel Corp. 2001, 5, 4. [Google Scholar]
- Fischler, M.A.; Bolles, R.C. Random sample consensus: A paradigm for model fitting with applications to image analysis and automated cartography. Commun. ACM 1981, 24, 381–395. [Google Scholar]
- Sirtkaya, S.; Seymen, B.; Alatan, A.A. Loosely coupled Kalman filtering for fusion of Visual Odometry and inertial navigation. In Proceedings of the 16th International Conference on Information Fusion; IEEE: New York, NY, USA, 2013; pp. 219–226. [Google Scholar]
- Kerl, C.; Sturm, J.; Cremers, D. Dense visual SLAM for RGB-D cameras. In Proceedings of the 2013 IEEE/RSJ International Conference on Intelligent Robots and Systems; IEEE: New York, NY, USA, 2013; pp. 2100–2106. [Google Scholar]
- Xiao, Y.; Li, B.; Xu, W.; Zhou, W.; Xu, B.; Zhang, H. Optimization of a Dense Mapping Algorithm with Enhanced Point-Line Features for Open-Pit Mining Environments. Appl. Sci. 2025, 15, 3579. [Google Scholar] [CrossRef] [Scilit]
- Engel, J.; Schöps, T.; Cremers, D. LSD-SLAM: Large-scale direct monocular SLAM. In Proceedings of the European Conference on Computer Vision; Springer: Cham, Switzerland, 2014; pp. 834–849. [Google Scholar]
- Zhou, T.; Brown, M.; Snavely, N.; Lowe, D.G. Unsupervised learning of depth and ego-motion from video. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 1851–1858. [Google Scholar]
- LIP6 – Sorbonne Université. Monolithe Platform. Available online: https://monolithe.proj.lip6.fr/ (accessed on 2 March 2026).










| Work | Algorithmic Contribution | Fusion | Embedded Implementation | Evaluation/Targeted Trade-Off |
|---|---|---|---|---|
| Feature-PLPD [21] | Hardware-aware joint corner and line-segment processing | N/A | ARM-GPU | VO accuracy–runtime trade-off |
| Enhanced line tracking [22] | Geometry-aware optical-flow correction for persistent line-segment tracking | N/A | ARM-GPU | Improved tracking accuracy and runtime |
| FPGA co-design [23] | Portability of the computationally intensive Feature-PLPD front-end | N/A | ARM-FPGA | Resource-efficient front-end implementation |
| Present work | Integration and extension toward complete point–line VIO | Resource-aware loosely coupled ESEKF | Stereo ARM-GPU VIO + RGB-D-assisted monocular ARM-FPGA VIO | Joint accuracy–runtime– energy trade-off under different embedded resource constraints |
| VO | Corners | Segments | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Front-End | tr (m) | rot (rad) | T (fps) | Detect | F1 (%) | T (fps) | Match | Detect | A (pix) | T (fps) | Match |
| SOFT-VO [30] | 12.56 | 0.17 | 23 | GFT | 39.2 | 82 | OF | N/A | N/A | N/A | N/A |
| PL-SLAM [11] | 12.02 | 0.22 | 2.2 | ORB | 13.6 | 44.2 | BRIEF | LSD | 20.8 | 5.5 | LBD |
| PL-VIO [27] | 32.65 | 0.14 | 2.4 | FAST | 31.2 | 177 | OF | LSD | 20.8 | 5.5 | LBD |
| PLD-VINS [28] | 18.04 | 0.19 | 18 | GFT | 39.2 | 82 | OF | EDlines | 30.1 | 15.6 | OF |
| Ray-To-Ray [29] | 33.15 | 0.21 | 12 | ORB | 13.6 | 44.2 | BRIEF | EDlines | 30.1 | 15.6 | OF |
| Feature-PLPD [21] | 9.7 | 0.17 | 22 | PLPD | 44.2 | 33 | OF | PLPD | 50.3 | 31 | OF |
| Method | Fusion | Estimator | Visual Information | Inertial Handling | Evaluation Dataset/ Context | Computing Platform | Embedded/ on-Board Validation | Energy Evaluation |
|---|---|---|---|---|---|---|---|---|
| OKVIS [10] | Tightly coupled | Nonlinear optimization | Keypoints | IMU error terms/high-rate | EuRoC/own experiments | General-purpose CPU | Real-time system | Not reported |
| ROVIO [34] | Tightly coupled | EKF | Image patches | State propagation | Own indoor/UAV experiments | Single CPU core | On-board UAV, 20 Hz | Not reported |
| VINS- Mono [6] | Tightly coupled | Sliding-window optimization | Point features | IMU pre-integration | EuRoC + real-world experiments | Desktop CPU/mobile platform | MAV and mobile-device demonstrations | Not reported |
| PL-VIO [27] | Tightly coupled | Sliding-window optimization | Points and lines | IMU pre-integration | EuRoC + PennCOSYVIO | Intel Core i7-6700HQ, 2.60 GHz, 16 GB | Dataset evaluation | Not reported |
| PLD-VINS [28] | Tightly coupled | Sliding-window optimization | Points, lines, and depth | IMU pre-integration | OpenLORIS + real-world experiments | Xeon E5645/Intel NUC i7-8700K | Dataset + robot evaluation | Not reported |
| Proposed | Loosely coupled | ESEKF | Points and lines | Camera-rate prediction | KITTI + VICON | Jetson Xavier NX/Jetson Nano/DE1-SoC | GPU robotic deployment + FPGA hardware execution | Reported |
| Category | Metrics | Unit |
|---|---|---|
| Accuracy | ATE: translation (tr), rotation (rot) | m, rad |
| RPE: translation (tr) | m | |
| Runtime | Average runtime | ms |
| Average throughput | fps | |
| Energy | Average power consumption | W |
| Energy per frame | mJ |
| KITTI | VICON | |||
|---|---|---|---|---|
| Front-end | Extraction | MED | 20.75 | 17.8 |
| FLS & CR | 6.26 | 8.5 | ||
| Tracker | OF | 14.4 | 12.8 | |
| FC | 4.1 | 3.86 | ||
| COR | 0.03 | 0.04 | ||
| SVO | Triangulation | 0.18 | 0.19 | |
| PnP solver | 4.63 | 3.82 | ||
| Back-end | ESEKF | Prediction | 0.021 | 0.023 |
| Update | 0.027 | 0.022 | ||
| Average throughput (fps) | 22 | 21 | ||
| Configuration | ATE | FPS | |
|---|---|---|---|
| tr. (m) | rot. (rad) | ||
| Feature-PLPD VO [21] | 9.70 | 0.17 | 22 |
| Enhanced line VO [22] | 7.87 | 0.14 | 50 |
| Combined Feature-PLPD VO | 8.90 | 0.16 | 22 |
| Proposed Feature-PLPD VIO | 5.40 | 0.13 | 22 |
| ATE | Energy | FPS | ||||||
|---|---|---|---|---|---|---|---|---|
| VO | VIO | |||||||
| Dataset/seq | tr | rot | tr | rot | W | mJ | ||
| KITTI | 00 | 8.9 | 0.16 | 5.4 | 0.13 | 8.1 | 368 | 22 |
| 05 | 7.57 | 0.11 | 6.1 | 0.09 | ||||
| 06 | 5.92 | 0.12 | 4.68 | 0.1 | ||||
| 10 | 5.53 | 0.17 | 4.8 | 0.14 | ||||
| VICON | 1.35 | N/A | 1.21 | N/A | 6.9 | 328 | 21 | |
| Resource (%) | MED | OF | MED & OF |
|---|---|---|---|
| Logic gate | 75 | 20 | 85 |
| Logic register | 33 | 14 | 42 |
| Memory block | 33 | 16 | 44 |
| DSP block | 71 | 24 | 82 |
| Jetson Nano | DE1-SoC | |||
|---|---|---|---|---|
| Front-end | Extraction | MED | 13 | 8 |
| FLS & CR | 7.8 | 10.3 | ||
| Tracker | OF | 10.8 | 2.5 | |
| FC | 7.1 | 15.3 | ||
| COR | 0.01 | 0.01 | ||
| VO | Essential | 8.4 | 11.8 | |
| SVD & scale | 2.05 | 2.87 | ||
| Back-end | ESEKF | Prediction | 0.031 | 0.05 |
| Update | 0.029 | 0.04 | ||
| Average throughput (fps) | 20.4 | 19.7 | ||
| Platform | Translation Error | Energy | FPS | ||
|---|---|---|---|---|---|
| ATE | RPE | W | mJ | ||
| Jetson Nano | 1.24 | 0.047 | 2.1 | 50 | 20.4 |
| DE1-SoC | 1.26 | 0.05 | 1.5 | 15.8 | 19.7 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Mamri, A.; El Hadri, A.; Benallegue, A.; Hachicha, K. Feature-PLPD-Aided Visual–Inertial Odometry for Low-Cost Embedded Systems. Sensors 2026, 26, 5940. https://doi.org/10.3390/s26185940
Mamri A, El Hadri A, Benallegue A, Hachicha K. Feature-PLPD-Aided Visual–Inertial Odometry for Low-Cost Embedded Systems. Sensors. 2026; 26(18):5940. https://doi.org/10.3390/s26185940
Chicago/Turabian StyleMamri, Ayoub, Abdelhafid El Hadri, Abdelaziz Benallegue, and Khalil Hachicha. 2026. "Feature-PLPD-Aided Visual–Inertial Odometry for Low-Cost Embedded Systems" Sensors 26, no. 18: 5940. https://doi.org/10.3390/s26185940
APA StyleMamri, A., El Hadri, A., Benallegue, A., & Hachicha, K. (2026). Feature-PLPD-Aided Visual–Inertial Odometry for Low-Cost Embedded Systems. Sensors, 26(18), 5940. https://doi.org/10.3390/s26185940

