Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (149)

Search Parameters:
Keywords = alignment of point clouds to images

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
21 pages, 1336 KB  
Article
Geometry-Guided Diffusion SAR Point Cloud Denoising
by Chengwei Zhang, Tao Jiang, Xinhao Xu, Wenjie Li, Fubo Zhang and Longyong Chen
Remote Sens. 2026, 18(15), 2458; https://doi.org/10.3390/rs18152458 - 26 Jul 2026
Viewed by 375
Abstract
Three-dimensional synthetic aperture radar (SAR) point clouds provide valuable geometric observations of urban scenes, but they often suffer from severe noise and layer-like artifacts caused by the low signal-to-noise ratio and tomographic imaging mechanism. These degradations make SAR point cloud denoising significantly more [...] Read more.
Three-dimensional synthetic aperture radar (SAR) point clouds provide valuable geometric observations of urban scenes, but they often suffer from severe noise and layer-like artifacts caused by the low signal-to-noise ratio and tomographic imaging mechanism. These degradations make SAR point cloud denoising significantly more challenging than conventional LiDAR point cloud denoising. In this paper, we propose a Geometry-guided Diffusion SAR Point Cloud Denoising (GDSD) framework to recover geometrically coherent building surfaces from noisy SAR point clouds.The key idea is to exploit relatively clean LiDAR point clouds as geometry priors while avoiding the need for paired SAR–LiDAR supervision or clean SAR ground truth. Specifically, we introduce a Forward Gaussian Noising Process to disrupt the intrinsic layer-like artifacts of SAR point clouds and reduce the input-level discrepancy between SAR and LiDAR domains. We further design a geometry prototype-based alignment module that projects SAR and LiDAR bottleneck features into a shared LiDAR-dominated latent space, enabling geometry-aware conditional reverse diffusion. A DiT-3D-based denoising network is then trained with LiDAR-domain diffusion supervision and applied to SAR point clouds using the aligned SAR geometry condition. To evaluate the proposed method, we construct a SAR point cloud denoising benchmark based on the MV3DSAR dataset with CAD-derived reference surfaces. Experimental results show that GDSD significantly improves the quality of noisy SAR point clouds and clearly outperforms the previous conventional LiDAR point cloud denoising baseline, producing more continuous and geometrically coherent SAR building point clouds. Full article
(This article belongs to the Section AI Remote Sensing)
Show Figures

Figure 1

23 pages, 4175 KB  
Article
BEV-Nexus: BEV Perception Algorithm Based on Depth Perception Enhancement and Dynamic Adaptive Fusion
by Xiaona Song, Haozhe Zhang, Zhengyi Huang, Jianlin Zhao and Lijun Wang
Sensors 2026, 26(15), 4720; https://doi.org/10.3390/s26154720 - 25 Jul 2026
Viewed by 340
Abstract
This paper proposes an improved multimodal fusion framework for 3D object detection, termed BEV-Nexus, which aims to address the issues of inaccurate depth estimation and inefficient fusion paradigms in existing image-point cloud fusion methods. We introduce a Point-Cloud-Guided Depth Prediction Network (PCGD-Net), which [...] Read more.
This paper proposes an improved multimodal fusion framework for 3D object detection, termed BEV-Nexus, which aims to address the issues of inaccurate depth estimation and inefficient fusion paradigms in existing image-point cloud fusion methods. We introduce a Point-Cloud-Guided Depth Prediction Network (PCGD-Net), which enhances the image branch’s depth prediction capability by embedding point cloud spatial prior, ground-truth loss constraint, and projected point cloud depth filling. Additionally, we design a Dynamic Self-adaptive Feature Fusion Module (DSF-Module), which computes multimodal feature similarity using window attention and performs weighted fusion based on self-adaptive weights, resolving alignment deviations in BEV features. Finally, we propose a Dilated Attention Enhancement Block (DAEB), which expands the receptive field through dilated convolution and integrates parameter-free attention mechanism (SimAM) for feature enhancement, ensuring efficiency while improving overall feature representation. Experimental results on nuScenes validation set show that BEV-Nexus outperforms it baseline (BEVFusion) by 1.8% mAP and 1.5% NDS. On the test set, BEV-Nexus improves mAP and NDS by 1.6% and 1.4%, respectively. Furthermore, the detection FPS remains nearly unchanged, demonstrating significant lightweight advantages. Full article
(This article belongs to the Special Issue AI-Powered Vision Sensing for Autonomous Driving)
Show Figures

Figure 1

24 pages, 24004 KB  
Article
Video Geospatial Mapping of Large-Scale Tower-Based Cameras Based on 3D GIS and Gradient Descent
by Xianguo Ling, Xingguo Zhang, Xin Li and Xiangfei Meng
ISPRS Int. J. Geo-Inf. 2026, 15(7), 316; https://doi.org/10.3390/ijgi15070316 - 12 Jul 2026
Viewed by 566
Abstract
To address the challenges of the large-scale georeferencing of tower-based cameras and the limited capability of video-based spatial analysis, we proposed a geospatial mapping method integrating 3D GIS and gradient descent optimization. Using a Digital Elevation Model (DEM), high-resolution remote sensing imagery, and [...] Read more.
To address the challenges of the large-scale georeferencing of tower-based cameras and the limited capability of video-based spatial analysis, we proposed a geospatial mapping method integrating 3D GIS and gradient descent optimization. Using a Digital Elevation Model (DEM), high-resolution remote sensing imagery, and tower-based video data as the primary data sources, the proposed method first estimates the intrinsic parameters of the tower-based camera by aligning a 3D GIS virtual camera with the video imagery. Subsequently, the initial camera extrinsic parameters are estimated using the PnP algorithm based on the previously estimated intrinsic matrix K and the corresponding control point pairs. Building upon these initial estimates, the camera intrinsic and extrinsic parameters are jointly optimized using a constrained L-BFGS-B framework that incorporates prior knowledge of the tower planar location, explicit box constraints, and a semi-constrained parameterization scheme with bounded parameter ranges. Furthermore, an outlier-removal and re-optimization strategy is employed to further improve the accuracy of parameter estimation. Finally, the optimized parameters are employed to transform image coordinates into three-dimensional world coordinates, and video geospatial mapping is achieved through the integration of colored point clouds with the 3D GIS scene. The results showed the following: (1) The 3D GIS scene constructed from publicly available DEM and high-resolution remote sensing imagery met the requirements for the initial estimation of intrinsic and extrinsic camera parameters. (2) Compared with PnP, RANSAC-PnP, SQPnP, and DLT, the proposed method achieves lower reprojection and 3D spatial errors. For the independent check points, the RMSE of the reprojection error is reduced by 66.4%, 73.6%, 68.0%, and 48.3%, respectively, while the RMSE of the 3D spatial error is reduced by 84.6%, 86.2%, 83.1%, and 69.4%, respectively. These results demonstrate that the proposed method provides reliable camera parameter estimates for video geospatial mapping. (3) Using the estimated camera parameters, image coordinates are transformed into 3D world coordinates to generate a georeferenced colored point cloud, which facilitates integrated analysis with existing geospatial datasets. The proposed method provides a feasible solution for tower-based camera georeferencing and three-dimensional visualization under conditions without field calibration. It offers a theoretical and technical basis for geospatial monitoring and related applications. Full article
Show Figures

Figure 1

25 pages, 1945 KB  
Article
Edge-Texture-Aware Semantic Dual-Query Fusion for Multimodal 3D Object Detection
by Yuehan Wu, Zheng Zheng, Kai Liu, Leyan Chen and Rihan Wu
Symmetry 2026, 18(7), 1133; https://doi.org/10.3390/sym18071133 - 2 Jul 2026
Viewed by 314
Abstract
Multimodal 3D object detection benefits from the complementary nature of camera images and LiDAR point clouds. However, existing voxel–pixel fusion methods typically rely on relatively coarse cross-modal interactions, which limit fine-grained structural modeling and degrade performance on small safety-critical objects. To address this [...] Read more.
Multimodal 3D object detection benefits from the complementary nature of camera images and LiDAR point clouds. However, existing voxel–pixel fusion methods typically rely on relatively coarse cross-modal interactions, which limit fine-grained structural modeling and degrade performance on small safety-critical objects. To address this issue, we propose ETA-SDQF, an edge-texture-aware semantic dual-query fusion framework designed to enhance 3D perception of vehicles, cyclists, and pedestrians. The proposed method first introduces an edge-texture-aware image backbone (ETAIB) based on the discrete wavelet transform (DWT), which improves the representation of multi-scale fine-grained image features. Then, we design a dual-query-guided attention fusion (DQGAF) module, which leverages deformable attention to adaptively aggregate voxel-aligned multi-scale image features under joint semantic and edge-texture guidance. Finally, we adopt a hybrid 3D feature learning strategy inspired by PV-RCNN, combining voxel-based feature learning with PointNet-style feature abstraction for processing fused features. This design improves the utilization of voxel features enriched with image semantics, thereby facilitating more reliable 3D object proposal generation. Experimental results on the KITTI dataset demonstrate that the proposed framework achieves better performance compared to existing baseline methods. It consistently improves pedestrian and cyclist detection, while maintaining competitive performance on car detection across different difficulty levels, showing potential benefits on challenging KITTI samples. Full article
(This article belongs to the Section A: Computer Science)
Show Figures

Figure 1

21 pages, 3211 KB  
Article
Object-Centric Seamless Pose Estimation in Multi-Object Scenes by Scale Alignment of Ray Diffusion and Iterative Closest Point
by YeonChang Jeong, Dong-Uk Seo, Kwanwoo Park and Soon-Yong Park
Appl. Sci. 2026, 16(13), 6624; https://doi.org/10.3390/app16136624 - 2 Jul 2026
Viewed by 364
Abstract
Robust estimation of camera trajectories from unconstrained image sequences remains a fundamental problem in computer vision and robotics. Recently, a diffusion-based camera tracking network has shown strong performance in sparse-view and single-object-centric settings, where a consistent object is observed across frames. However, when [...] Read more.
Robust estimation of camera trajectories from unconstrained image sequences remains a fundamental problem in computer vision and robotics. Recently, a diffusion-based camera tracking network has shown strong performance in sparse-view and single-object-centric settings, where a consistent object is observed across frames. However, when multiple objects appear sequentially in a video, the initially observed object may disappear as the sequence progresses, which prevents maintaining the “single-object-centric” paradigm across all frames and degrades pose estimation when the conventional method is applied to the multi-object sequence. In this work, we propose an object-centric camera pose estimation framework that handles such sequences by partitioning a video into object-level sub-scenes. As a baseline network, Ray Diffusion is applied to single-object sub-scenes, while frame-to-frame camera motion in multi-object sub-scenes is estimated using monocular video depth, object masks, and point cloud alignment using Iterative Closest Point (ICP). Since the domain of pose estimation from different sub-scenes is inconsistent in terms of pose scale, it requires seamless concatenation of pose estimation results through all sub-scenes. In this regard, we introduce a scale alignment strategy based on reprojection error minimization. This enables the pose estimates from individual sub-scenes to be integrated into a single and seamless camera trajectory. We evaluate the proposed method on a newly collected indoor dataset consisting of 40 multi-object video sequences. Experimental results compare our camera trajectory estimation with both the diffusion-based method and the state-of-the-art visual SLAM methods. Full article
(This article belongs to the Section Computing and Artificial Intelligence)
Show Figures

Figure 1

24 pages, 34784 KB  
Article
Occluder-Mask-Constrained 3D Reconstruction from Tower-Crane Construction Site Imagery
by Qirun He, Rong Zhang, Changjiang Yin, Qin Ye and Shaoming Zhang
Electronics 2026, 15(13), 2883; https://doi.org/10.3390/electronics15132883 - 1 Jul 2026
Viewed by 366
Abstract
3D reconstruction of construction scenes is an important enabling technology for digital and intelligent construction project management. Recurring foreground occluders and dynamic disturbances in tower-crane imagery can destabilize image registration and introduce spurious depth responses. This paper proposes an occluder-mask-constrained 3D reconstruction framework [...] Read more.
3D reconstruction of construction scenes is an important enabling technology for digital and intelligent construction project management. Recurring foreground occluders and dynamic disturbances in tower-crane imagery can destabilize image registration and introduce spurious depth responses. This paper proposes an occluder-mask-constrained 3D reconstruction framework driven by multi-view geometric anomalies. Adjacent-view geometric outliers are spatially aggregated to generate foreground prompt points, which are converted into occluder masks using Segment Anything Model 2 (SAM2). The masks are propagated as unified pixel-validity constraints through sparse feature filtering, Adaptive Patch Deformation Multi-View Stereo (APD-MVS) matching-cost evaluation, support-region selection, and depth-map fusion. Experiments on three real construction-site datasets show increased sparse-registration completeness in the tested sequences and fewer visually identifiable occluder-induced artifacts in dense point clouds. A representative 308-image sequence was further evaluated against no-mask reconstruction, You Only Look Once version 8 (YOLOv8) bounding-box removal, manually prompted Segment Anything Model 2.1 (SAM2.1), a Segment Anything Model 3 (SAM3) text-prompt baseline, and Visibility-Aware Multi-View Stereo Network (Vis-MVSNet). The evaluation combines sparse-reconstruction metrics, pixel-level mask-quality metrics from a manually annotated validation subset, module-wise runtime accounting, controlled ablations, and aligned dense-point-cloud visualization. These results show improved sparse-stage registration completeness and visible artifact suppression. Because high-precision 3D reference point clouds are unavailable, the dense results are interpreted as visual evidence of artifact suppression rather than as proof of improved absolute dense-reconstruction accuracy. Full article
(This article belongs to the Special Issue Advances in Object Tracking and Localization)
Show Figures

Figure 1

40 pages, 5967 KB  
Systematic Review
Radar-Camera Extrinsic Calibration for Roadside Infrastructure: A Systematic Review
by Zeynab Rokhi and Ali Emadi
Vehicles 2026, 8(6), 137; https://doi.org/10.3390/vehicles8060137 - 19 Jun 2026
Viewed by 464
Abstract
The growth of Intelligent Transportation Systems (ITS) has made high-quality perception data from multi-sensor setups essential. Pairing millimeter-wave (mmW) radar with a monocular camera is a common way to recover three-dimensional information about the environment, but aligning the two is difficult because sparse [...] Read more.
The growth of Intelligent Transportation Systems (ITS) has made high-quality perception data from multi-sensor setups essential. Pairing millimeter-wave (mmW) radar with a monocular camera is a common way to recover three-dimensional information about the environment, but aligning the two is difficult because sparse radar point clouds and dense camera images differ sharply in how they sense a scene. The problem grows more severe in roadside infrastructure, where the high mounting elevation introduces perspective distortion that vehicle-mounted systems rarely face. This paper presents a systematic review, conducted under the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines, of radar-camera extrinsic calibration for fixed roadside infrastructure, organizing existing work into a taxonomy that separates traditional two-stage pipelines from recent end-to-end learning frameworks. Because methods designed specifically for roadside units remain scarce, the review also covers vehicle- and robot-mounted methods whose static-sensor formulation carries over to fixed roadside deployment. For the two-stage pipeline, the analysis covers target-based and targetless correspondence registration along with the optimization techniques and algorithmic assumptions behind parameter estimation. The end-to-end learning literature shows a clear shift toward self-supervised and fusion-based models, some of which report real-time performance. The review also compares the metrics and procedures used to quantify calibration accuracy. Progress is evident, but robustness in cluttered urban environments remains an open challenge, and the paper closes by outlining future directions, arguing that standardized roadside benchmarks are needed before scalable, targetless calibration can mature. Full article
Show Figures

Figure 1

36 pages, 5240 KB  
Article
Single-View Scene Completion via Candidate Model Retrieval and Scale-Aware Registration
by Di Zhao, Yuxing Wang, Ziheng Shi and Junhan Shao
Appl. Sci. 2026, 16(12), 5778; https://doi.org/10.3390/app16125778 - 8 Jun 2026
Viewed by 262
Abstract
Single-view RGB-D observations are often affected by occlusion and restricted viewpoints, leading to incomplete object geometry and underestimated obstacle extents in indoor robot perception. This paper proposes a single-view scene completion framework that integrates candidate model retrieval and scale-aware registration. The framework first [...] Read more.
Single-view RGB-D observations are often affected by occlusion and restricted viewpoints, leading to incomplete object geometry and underestimated obstacle extents in indoor robot perception. This paper proposes a single-view scene completion framework that integrates candidate model retrieval and scale-aware registration. The framework first generates local RGB crops and partial point clouds through automatic instance segmentation; then retrieves complete candidate models by matching the local crops with multi-view rendered CAD images; and finally estimates candidate-to-observation rotation, translation, and scale to insert the selected aligned model into the original scene coordinate system. Experiments show that the retrieval module achieves Recall@1/Recall@5 of 80%/89%. The registration module reaches a success rate of 56.61%, outperforming the second-best method by 12.28 percentage points. More importantly, scene-level evaluation shows that the proposed method improves occupancy F1 from 0.445 to 0.523 and reduces boundary error from 0.202 m to 0.146 m compared with DiffCAD. These results indicate that the proposed framework improves navigation-oriented occupancy and obstacle-boundary recovery under CAD-library-based and segmentation-dependent single-view scene completion settings. Full article
(This article belongs to the Section Robotics and Automation)
Show Figures

Figure 1

22 pages, 6385 KB  
Article
Targetless Calibration of Wide-Baseline and Wide-Angle Surround-View Fisheye Cameras Using Cylindrical Projection Model
by Gee Hoon Lee and Soon-Yong Park
Sensors 2026, 26(12), 3622; https://doi.org/10.3390/s26123622 - 6 Jun 2026
Viewed by 511
Abstract
We propose a novel targetless extrinsic calibration method for wide-baseline and wide-angle fisheye cameras, which are mounted on a driving vehicle for surround view monitoring. Sequences of image frames from three fisheye cameras are obtained, and the object instance and depth around the [...] Read more.
We propose a novel targetless extrinsic calibration method for wide-baseline and wide-angle fisheye cameras, which are mounted on a driving vehicle for surround view monitoring. Sequences of image frames from three fisheye cameras are obtained, and the object instance and depth around the vehicle are used for calibration. Thus, the proposed method can be applied to online vehicle camera calibration. Fisheye images are first transformed into the cylindrical coordinate system by considering the panoramic formation of the cameras. Then, the state-of-the-art object detection and monocular depth estimation models are applied to the cylindrical images. Vehicle instances matched across different views are reconstructed into 3D point clouds, and their depths are scaled by employing the pose geometry of the front camera. The per-point depths and global scale are then jointly optimized to achieve accurate cross-view alignment and extrinsic calibration. Experiments on both real-world and synthetic video datasets show that the proposed method achieves higher accuracy than COLMAP and DUSt3R under challenging conditions such as wide baselines and low frame rates, without requiring an artificial calibration target. Full article
Show Figures

Figure 1

21 pages, 5157 KB  
Article
3D Quantitative Modeling for Stone Fruit Quality Assessment by LF-NMRI
by Kang Wang, Bing Li, Shan Zeng, Wei Tao, Ke Yang and Zhiguang Yang
Foods 2026, 15(11), 2012; https://doi.org/10.3390/foods15112012 - 4 Jun 2026
Viewed by 408
Abstract
The core volume ratio (CVR) is a key indicator for evaluating the proportion of edible fraction in stone fruits. Traditionally, CVR is determined through destructive sampling by separately measuring the masses of the core and entire fruit. Recently, low-field nuclear magnetic resonance imaging [...] Read more.
The core volume ratio (CVR) is a key indicator for evaluating the proportion of edible fraction in stone fruits. Traditionally, CVR is determined through destructive sampling by separately measuring the masses of the core and entire fruit. Recently, low-field nuclear magnetic resonance imaging (LF-NMRI) has been introduced as a non-destructive alternative, but its sparse sampling limits the ability to achieve accurate spatial and volumetric quantification of fruit quality. To address this limitation, we propose a novel method for high-precision three-dimensional (3D) modeling of stone fruits. The method acquires tomographic LF-NMRI sequences along three orthogonal axes. Each sequence is segmented into pulp and core regions using a SwinUNet deep learning model and converted into point clouds for each view. Point clouds from the three orthogonal views are registered via a genetic algorithm to align structural information from complementary perspectives and fused into a unified 3D model through Poisson surface reconstruction. Using prunes as a representative case, the method enables accurate quantification of core and entire fruit volumes, achieving a CVR estimation with a mean absolute error of 0.13% compared to manual measurements. The proposed three-view reconstruction strategy yields a volumetric error of only 0.73%, significantly outperforming single-view (4.57%) and dual-view (3.73%) approaches. This technology provides a robust and accurate non-destructive solution for 3D internal quality analysis of fruits. Full article
Show Figures

Figure 1

19 pages, 4328 KB  
Article
Safe Distance Monitoring for Substation Near-Current Operations via Image–LiDAR Cross-Modal Self-Registration
by Maonan Wang, Bo Wang, Xinming Fan, Tianrui Yin, Hengrui Ma and Peng Luo
Electronics 2026, 15(11), 2321; https://doi.org/10.3390/electronics15112321 - 27 May 2026
Cited by 1 | Viewed by 336
Abstract
Continuous monitoring of the minimum safety distance between construction machinery and energized bodies is essential during operations near energized equipment in substations. Existing methods mostly rely on fixed-view observation, online LiDAR, or rigid camera–LiDAR installation, leading to inflexible deployment, high extrinsic-maintenance cost, and [...] Read more.
Continuous monitoring of the minimum safety distance between construction machinery and energized bodies is essential during operations near energized equipment in substations. Existing methods mostly rely on fixed-view observation, online LiDAR, or rigid camera–LiDAR installation, leading to inflexible deployment, high extrinsic-maintenance cost, and insufficient metric consistency across viewpoints. To address these limitations, this paper proposes a safety-distance monitoring method based on cross-modal self-registration between monocular images and a pre-built LiDAR map. During online operation, only the current monocular image is used. Monocular depth estimation first generates a pseudo-point cloud, which is then registered with historical LiDAR point clouds to solve the camera pose and align the current observation with the map. Combined with target-boundary segmentation and prior energized-hazardous-region information, the method localizes key parts of construction machinery in 3D and computes the minimum safety distance to energized regions. Experiments show that the proposed method achieves a registration recall of 92.2%, a mean absolute error of 0.748 m, a maximum error of 0.822 m, and a single-frame latency of 180 ms. Notably, these results are achieved without real-time LiDAR input or on-site extrinsic recalibration. These results demonstrate the feasibility of the proposed framework in a representative substation scenario and indicate its potential for auxiliary online safety-distance monitoring. Full article
Show Figures

Figure 1

14 pages, 1804 KB  
Article
Air Target ISAR Recognition Based on Data Augmentation and Transfer Learning
by Moqian Wang, Zuzhen Huang, Jinjian Cai, Tao Wu and Youquan Lin
Sensors 2026, 26(11), 3323; https://doi.org/10.3390/s26113323 - 23 May 2026
Viewed by 679
Abstract
Aiming at the problems of extremely scarce measured samples and significant domain shift between simulated and measured data in automatic target recognition (ATR) of air targets for spaceborne radar, this paper proposes an inverse synthetic aperture radar (ISAR) image recognition method for air [...] Read more.
Aiming at the problems of extremely scarce measured samples and significant domain shift between simulated and measured data in automatic target recognition (ATR) of air targets for spaceborne radar, this paper proposes an inverse synthetic aperture radar (ISAR) image recognition method for air targets combining physics-driven data augmentation guided by detection prior information with domain adversarial transfer learning. First, the mapping relationship between scattering point projection and ISAR images is established by using the target 3D point cloud and radar observation geometric priors, and a 2D sinc kernel function is introduced for energy distribution rendering. Then, under the unsupervised transfer learning paradigm, aiming at the distribution inconsistency between augmented data (source domain) and unlabeled simulated data (target domain), this paper designs a cross-domain recognition task experiment including six types of typical aircraft targets, and compares the cross-domain recognition performance of three transfer learning methods (model fine-tuning, deep domain confusion (DDC) and domain-adversarial neural networks (DANN)) on the target domain. Meanwhile, t-distributed stochastic neighbor embedding (t-SNE) visualization is used to analyze the feature distribution alignment ability of the models. Simulation experiments show that the DANN model with a dynamic inversion coefficient introduced in the gradient reversal layer (GRL) achieves a recognition accuracy of 99.5% on the unlabeled target domain, which is significantly superior to the model fine-tuning and DDC methods. Moreover, it makes the feature distributions of source and target domain samples highly overlapping, and maintains a strong inter-class discriminability while eliminating the domain shift. The proposed scheme provides a physically interpretable and robust technical path for few-shot radar target image recognition. Full article
(This article belongs to the Section Radar Sensors)
Show Figures

Figure 1

23 pages, 11707 KB  
Technical Note
HyperCoreg: An Automated, Operational Pipeline for Co-Registering PRISMA and EnMAP Hyperspectral Imagery
by José Antonio Gámez García, Giacomo Lazzeri and Deodato Tapete
Geomatics 2026, 6(3), 47; https://doi.org/10.3390/geomatics6030047 - 11 May 2026
Viewed by 851
Abstract
HyperCoreg is an automated, end-to-end pipeline for geometric co-registration of spaceborne hyperspectral imagery (PRISMA L2D and EnMAP L2A) to Sentinel-2 Level-2A reference data. The workflow addresses scene-dependent geolocation errors that hinder reliable data fusion and multi-temporal analyses, particularly in cloud-affected acquisitions. HyperCoreg builds [...] Read more.
HyperCoreg is an automated, end-to-end pipeline for geometric co-registration of spaceborne hyperspectral imagery (PRISMA L2D and EnMAP L2A) to Sentinel-2 Level-2A reference data. The workflow addresses scene-dependent geolocation errors that hinder reliable data fusion and multi-temporal analyses, particularly in cloud-affected acquisitions. HyperCoreg builds on the AROSICS framework without replacing its image-matching engine and extends it at the workflow level through four operational functions: automated Sentinel-2 candidate selection, hyperspectral-to-multispectral band pairing, sequential alignment logic, and quality-controlled acceptance. The main output is a co-registered hyperspectral cube along with comprehensive metrics, per-scene reports, and optional diagnostic products that support accessible quality control. Performance is evaluated on a long time series of PRISMA images collected from 2019 to 2025 and an EnMAP test set acquired in 2025, over the Metropolitan City of Rome (Italy). The multi-sensor dataset encompasses heterogeneous acquisition conditions, including variable cloud cover, illumination, and seasonal variability. The results show systematic reductions in mean residual error compared with a controlled basic AROSICS-based pipeline configuration. The largest gains are achieved in challenging conditions where tie points are sparse or unevenly distributed. By improving geometric consistency, this pipeline facilitates spatial layering and integration of hyperspectral data with higher-resolution urban layers and supports a range of downstream applications where data integration and spatiotemporal consistency are cornerstones of further analysis. Full article
Show Figures

Graphical abstract

34 pages, 17465 KB  
Article
Backpack System Development and Image-LiDAR Integration for Improved Geospatial Data Alignment in Forest Mapping
by Raja Manish, Songlin Fei and Ayman Habib
Remote Sens. 2026, 18(9), 1443; https://doi.org/10.3390/rs18091443 - 6 May 2026
Viewed by 468
Abstract
Backpack mobile mapping systems (MMS) equipped with LiDAR and RGB cameras, as well as an optional GNSS/INS direct georeferencing unit, are increasingly utilized in forest inventory applications. In general, LiDAR point clouds provide detailed structural information, whereas imagery offers visual specifics of surface [...] Read more.
Backpack mobile mapping systems (MMS) equipped with LiDAR and RGB cameras, as well as an optional GNSS/INS direct georeferencing unit, are increasingly utilized in forest inventory applications. In general, LiDAR point clouds provide detailed structural information, whereas imagery offers visual specifics of surface features. However, cameras typically operate at lower acquisition rates compared to LiDAR. In proximal mapping, another challenge is the inconsistent reception of GNSS signals beneath forest canopies. Additionally, georeferencing accuracy may differ between LiDAR and imagery due to biases in the system calibration parameters and variations in post-processing approaches. To address these challenges, this study introduces a Backpack MMS that uses cameras configured at elevated frame rates to enhance image overlap. Concurrently, this study presents an algorithmic approach to addressing georeferencing issues by integrating imagery and LiDAR data, thereby enhancing system calibration and improving platform trajectory. The method is based on the hypothesis that forest environments are rich with geometrically well-defined features, such as tree trunks and ground patches. By identifying conjugate primitives in point clouds from both imagery and LiDAR, the procedure optimizes feature models while simultaneously minimizing calibration biases and/or trajectory errors. The proposed approach is validated using multiple field datasets collected in diverse forest environments. Quantitative results show that the procedure reduces image–LiDAR feature misalignment across all datasets from up to 1.1 m in the planimetric direction and 2 m in the vertical direction to within 5 cm in both. The feature fitting accuracy also improves from 2.9 cm to 0.85 cm for LiDAR point clouds and from 10 cm to 0.9 cm for image-based point clouds. However, the results indicate that despite increased data availability, imagery alone remains less reliable than LiDAR for extracting structural information. Nevertheless, the proposed image–LiDAR alignment strategy represents a crucial step toward developing a comprehensive tree inventory. Full article
Show Figures

Figure 1

24 pages, 4915 KB  
Article
Semantic-Guided Matching of Heterogeneous UAV Imagery and Mobile LiDAR Data Using Deep Learning and Graph Neural Networks
by Tee-Ann Teo, Hao Yu and Pei-Cheng Chen
Drones 2026, 10(3), 185; https://doi.org/10.3390/drones10030185 - 8 Mar 2026
Cited by 1 | Viewed by 807
Abstract
The integration of heterogeneous geospatial data, specifically low-cost unmanned aerial vehicle (UAV) imagery and mobile light detection and ranging (LiDAR) system point clouds, presents a significant challenge due to the significant radiometric and structural discrepancies between the two modalities. This study proposes a [...] Read more.
The integration of heterogeneous geospatial data, specifically low-cost unmanned aerial vehicle (UAV) imagery and mobile light detection and ranging (LiDAR) system point clouds, presents a significant challenge due to the significant radiometric and structural discrepancies between the two modalities. This study proposes a novel air-to-ground semantic feature matching framework to achieve precise geometric registration between these data sources by effectively incorporating semantic-constraint deep learning-based matching. The methodology transformed the cross-sensor alignment challenge into a robust two-dimensional image matching problem. This was achieved by first using YOLOv11 for semantic segmentation of common road markings in both the UAV orthoimage and the converted LiDAR intensity image to generate highly consistent feature references. Subsequently, the SuperPoint detector and a graph neural network matcher, SuperGlue, were applied to these semantic images to establish reliable geomatics information correspondence points. Experimental results confirmed that this semantic-guided strategy consistently outperformed traditional feature-based matching (i.e., scale-invariant feature transform + fast library for approximate nearest neighbors), particularly by converting the noisy LiDAR intensity image into a stabilized semantic representation. The explicit application of semantic constraints further proved effective in eliminating false matches between geometrically similar but semantically distinct objects. The final object-specific analysis demonstrated that features with clear, complex geometric structures (e.g., pedestrian crossings and directional arrows) provide the most robust matching control. In summary, the proposed framework successfully leverages semantic context to overcome cross-sensor heterogeneity, offering an automated and precise solution for the geometric alignment of mobile LiDAR data. Full article
Show Figures

Figure 1

Back to TopTop