Skip to Content
Remote SensingRemote Sensing
  • Article
  • Open Access

25 February 2026

A Coarse-to-Fine Optical-SAR Image Registration Algorithm for UAV-Based Multi-Sensor Systems Using Geographic Information Constraints and Cross-Modal Feature Consistency Mapping

,
,
,
,
,
and
1
College of Intelligence Science and Technology, National University of Defense Technology, Changsha 410073, China
2
School of Physical Education and Health, Hunan University of Technology and Business, Changsha 410205, China
*
Author to whom correspondence should be addressed.

Highlights

What are the main findings?
  • A coarse-to-fine registration framework integrating geographic information constraints with cross-modal feature consistency mapping is proposed, achieving automatic elimination of scale differences and rotation deviations between optical and SAR images through imaging geometry-based coordinate transformation.
  • On the integrated airborne optical/SAR dataset, the method achieves 2.00 (CPU)/1.97 (GPU) pixels in average RMSE, outperforming traditional and state-of-the-art deep learning methods while reducing computation time by 37.0%; its geographic-constrained coarse registration improves SuperGlue/LoFTR CMR by 167%/109%, and the hybrid GLS + LightGlue refinement yields the optimal 1.95-pixel RMSE with GPU acceleration.
What are the implications of the main findings?
  • The geographic information constraint approach circumvents radiometric differences that challenge both traditional intensity-based methods and data-driven deep learning approaches, providing a practical registration framework for UAV-based multi-sensor fusion systems in GNSS-aided scenarios.
  • The method demonstrates robust performance under challenging conditions (dense smoke, fire, illumination variations) and provides the best accuracy–efficiency trade-off for resource-constrained UAV platforms, enabling real-time applications in environmental monitoring, disaster assessment, target detection, and visual navigation tasks.
  • The demonstrated complementary benefits between geometric constraints and learned features suggest promising directions for future hybrid approaches combining physics-based geometric priors with data-driven feature representations for cross-modal image matching.

Abstract

Optical and synthetic aperture radar (SAR) image registration faces challenges from nonlinear radiometric distortions and geometric deformations caused by different imaging mechanisms. This paper proposes a coarse-to-fine registration algorithm integrating geographic information constraints with cross-modal feature consistency mapping. The coarse stage employs imaging geometry-based coordinate transformation with airborne navigation data to eliminate scale and rotation differences. The fine stage constructs a multi-scale phase congruency-based feature response aggregation model combined with rotation-invariant descriptors and global-to-local search for sub-pixel alignment. Experiments on integrated airborne optical/SAR datasets demonstrate superior performance with an average RMSE of 2.00 pixels, outperforming both traditional handcrafted methods (3MRS, OS-SIFT, POS-GIFT, GLS-MIFT) and state-of-the-art deep learning approaches (SuperGlue, LoFTR, ReDFeat, SAROptNet) while reducing execution time by 37.0% compared with the best-performing baseline. The proposed coarse registration also serves as an effective preprocessing module that improves SuperGlue’s matching rate by 167% and LoFTR’s by 109%, with a hybrid refinement strategy achieving 1.95 pixels RMSE. The method demonstrates robust performance under challenging conditions, enabling real-time UAV-based multi-sensor fusion applications.

1. Introduction

With the rapid advancement of remote sensing technology, the acquisition and comprehensive utilization of multimodal remote sensing imagery have become increasingly critical in various applications, including environmental monitoring, disaster assessment, target detection, and visual navigation [1,2,3]. Among different imaging modalities, synthetic aperture radar (SAR) and optical (visible light) imagery provide complementary information: SAR sensors operate independently of weather conditions and illumination, offering all-weather, day-and-night imaging capabilities with rich structural information, while optical sensors capture high-resolution spectral information consistent with human visual perception [4,5]. The fusion of SAR and optical imagery can significantly enhance the performance of remote sensing applications by leveraging the advantages of both modalities [6,7].
Image registration establishes spatial correspondence between images captured from different sensors, viewpoints, or times, serving as a fundamental prerequisite for multimodal image fusion [8,9]. However, registration of SAR and optical imagery remains a challenging task due to significant differences in imaging mechanisms [10,11]. These differences manifest in several aspects: (1) nonlinear radiometric distortions (NRD) caused by different spectral responses; (2) geometric deformations resulting from different viewing angles and acquisition conditions; (3) speckle noise inherent in SAR images that reduces feature detection accuracy; and (4) nonlinear intensity mappings of structural features, where edges, corners, and textures exhibit different appearances across modalities despite representing the same physical structures [12,13,14].
In unmanned aerial vehicle (UAV)-based detection systems, the integration of optical cameras and SAR sensors introduces unique challenges and opportunities [15,16]. UAV platforms provide flexible deployment and high-resolution data acquisition capabilities, but airborne sensors typically have limited accuracy in position and attitude measurements [17]. Moreover, UAV-based optical and SAR images often exhibit significant scale differences and rotation deviations due to different sensor mounting configurations and imaging geometries, which substantially increases the search space for feature matching, as shown in Figure 1. This limitation has motivated the development of advanced image registration algorithms that can achieve precise alignment without complete reliance on navigation data [18,19].
Figure 1. Illustration of UAV-based multi-sensor imaging system and cross-modal image characteristics. (a) Diagram of the UAV multi-sensor imaging system showing the visible light camera (nadir view) and SAR sensor (side-looking view) with their respective imaging geometries. (b) Optical image with high spectral resolution, capturing detailed color and texture information but dependent on weather and illumination conditions. (c) SAR image of the same geographic area with all-weather, day-and-night imaging capability but exhibiting speckle noise. The comparison highlights the complementary nature of optical and SAR imagery and the challenges in cross-modal registration due to different representations of the same ground features (water body, building, and vegetation).
Traditional feature-based registration methods, such as scale-invariant feature transform (SIFT) [20] and speeded-up robust features (SURF) [21], have achieved remarkable success in registering homogeneous images. However, these methods primarily rely on intensity gradient information, which exhibits significant variations between SAR and optical images due to the presence of NRD [22,23]. To address this challenge, researchers have proposed various approaches, including modified gradient definitions [24], structural feature descriptors [25], and phase congruency-based methods [26,27]. Recent years have witnessed the emergence of deep learning-based registration methods, such as SuperGlue, LoFTR, and their variants, which have achieved impressive performance in homogeneous image matching. However, these learning-based approaches typically require large-scale training datasets with accurate ground truth correspondences, which are difficult to obtain for SAR-optical image pairs. Furthermore, their computational demands often exceed the real-time processing requirements of UAV-based systems. Despite these advances, achieving efficient, robust, and accurate SAR-optical image registration remains an open problem.
Existing SAR-optical registration methods can be broadly categorized into two paradigms: (1) direct feature matching approaches that attempt to extract modality-invariant features from raw images and (2) image translation approaches that convert one modality to another before matching. The first paradigm faces the fundamental challenge of bridging the radiometric gap between modalities, often resulting in low matching rates. The second paradigm introduces additional complexity and potential artifacts from the translation process. In contrast, this paper proposes a third paradigm that leverages geographic information as a modality-invariant bridge. By transforming the registration problem from the image domain to the geographic coordinate domain, we circumvent the radiometric differences that challenge traditional intensity-based methods. This geographic constraint-based approach is particularly suitable for UAV platforms equipped with navigation sensors (GPS/IMU), which provide readily available pose information.
This paper proposes a novel coarse-to-fine image registration algorithm that integrates geographic information constraints with cross-modal feature consistency mapping for optical and SAR image alignment. Unlike existing methods that rely solely on image features, our approach exploits the complementary strengths of navigation data and image-based refinement. The main contributions are as follows:
(1)
A geographic information-constrained coarse registration strategy is proposed, which leverages imaging geometry-based coordinate transformation with UAV pose data to establish initial correspondence between images, effectively eliminating most scale differences and rotation deviations. This strategy reduces the matching search space by orders of magnitude compared with direct global matching.
(2)
A cross-modal feature consistency descriptor based on multi-scale Gaussian filtering and feature response aggregation is developed to overcome nonlinear radiometric differences between SAR and optical images. The descriptor exploits phase congruency theory to extract radiation-invariant structural features that exhibit consistent responses across modalities.
(3)
A global-to-local search (GLS) matching strategy is designed to improve matching efficiency while maintaining high accuracy. By constraining the search space using geometric predictions from coarse registration, the GLS strategy achieves 37% reduction in computational time compared with exhaustive search methods.
(4)
Comprehensive experimental validation demonstrates the effectiveness of the proposed algorithm, showing superior performance compared with state-of-the-art methods, including both traditional handcrafted feature methods (3MRS, OS-SIFT, POS-GIFT, GLS-MIFT) and deep learning-based approaches, with an average RMSE of 2.0042 pixels on integrated airborne optical/SAR datasets.
The remainder of this paper is organized as follows: Section 2 reviews related work on multimodal image registration; Section 3 details the proposed algorithm; Section 4 describes the experimental setup and results; and Section 5 concludes the paper.

3. Methodology

The optical-SAR image registration problem aims to find the optimal geometric transformation achieving best spatial alignment between the two modalities. This problem faces three core challenges: nonlinear radiometric mapping between different imaging mechanisms; complex geometric transformations involving scale, rotation, and translation; and an enormous non-convex search space. To address these challenges, we propose a hierarchical framework comprising coarse and fine stages, as shown in Figure 2. The coarse stage leverages airborne geographic information via imaging geometry-based coordinate transformation to eliminate scale differences and rotation deviations. The fine stage employs multi-scale phase congruency features with rotation-invariant descriptors and global-to-local search for sub-pixel alignment. The final transformation is obtained by composing the two-stage results.
Figure 2. Overview of the proposed coarse-to-fine optical-SAR image registration framework. The input section includes the optical image with UAV pose data (GPS, IMU) and the SAR image. Stage 1 (coarse registration) employs imaging geometry-based coordinate transformation with geographic information constraints, involving trisection point selection, pixel-to-geographic coordinate transformation, geographic coordinate matching, and affine transformation estimation to eliminate scale differences and rotation deviations. Stage 2 (fine registration) utilizes cross-modal feature consistency mapping, comprising multi-scale Gaussian pyramid construction, multi-direction gradient response computation, phase congruency feature aggregation, rotation-invariant descriptor construction, and global-to-local search (GLS) strategy for sub-pixel alignment. The final registration result is obtained by composing the coarse and fine transformation matrices.

3.1. Coarse Registration via Imaging Geometry-Based Coordinate Transformation with Geographic Information Constraints

Traditional registration methods directly seek correspondence in the image domain, facing high false matching rates due to cross-modal radiometric differences. To circumvent this limitation, we leverage geographic information from the airborne platform as a bridge, transforming the registration problem into point-pair matching in geographic space where coordinates are invariant to imaging modalities, as shown in Figure 3. The imaging geometry-based coordinate transformation model establishes pixel-to-geographic coordinate mapping through cascaded transformations across four coordinate systems: the pixel coordinate system F pixel , the camera coordinate system F cam , the UAV body coordinate system F uav , and the East–North–Up geographic coordinate system F geo .
Figure 3. Schematic diagram of the imaging geometry-based coordinate transformation model illustrating the cascaded coordinate transformations. The model establishes pixel-to-geographic coordinate mapping through four coordinate systems: the pixel coordinate system, the camera coordinate system, the UAV body coordinate system, and the East–North–Up geographic coordinate system. S1 denotes the transformation from pixel coordinates to camera coordinates via the pinhole camera model. S2 represents the transformation from camera to UAV body coordinates determined by gimbal attitude angles. S3 describes the transformation from UAV body to geographic coordinates using UAV pose parameters, including position (λ, φ, h) and attitude angles.
The translation vector t uav geo is determined by the UAV’s position in the geographic coordinate system. Let the UAV GPS position be ( λ , φ , h ) , where λ is longitude, φ is latitude, and h is altitude relative to the takeoff point. The translation vector can be expressed as t uav geo = [ t x , t y , h ] T , where t x and t y are the eastward and northward offsets, respectively, converted from longitude and latitude to the local ENU coordinate system. The imaging geometry-based coordinate transformation model establishes pixel-to-geographic coordinate mapping through cascaded transformations across four coordinate systems: the pixel coordinate system, the camera coordinate, the UAV body coordinate system (Fuav), and the East–North–Up geographic coordinate system.
Combining the above transformations, the complete mapping relationship from pixel coordinates to geographic coordinates is as follows:
p geo   = R uav geo R cam uav Z cam K 1 u v 1 + t uav geo
where p geo = X geo , Y geo , Z geo T is the 3D coordinate in the geographic coordinate system. To solve for the depth parameter Z cam , this paper adopts the ground plane assumption, assuming all ground points in the imaging area satisfy Z geo = 0 . Based on this constraint, we can derive the following:
Z cam = n T t geo uav n T R nav geo R cam uav K 1 u v 1
where n = [ 0 , 0 , 1 ] T is the ground plane normal vector.
Based on the established imaging geometry-based coordinate transformation model, the coarse registration algorithm first selects four uniformly distributed feature points in the optical reference image. To effectively constrain the six degrees of freedom of affine transformation, a trisection strategy is adopted. Let the optical image size be H v × W v , then the selected point coordinates are as follows:
p 1 v = W v 3 , H v 3
p 2 v = 2 W v 3 , H v 3
p 3 v = W v 3 , 2 H v 3
p 4 v = 2 W v 3 , 2 H v 3
For each pixel point p i v = ( u i , v i ) , Equations (7) and (8) are applied to compute its corresponding geographic coordinates g i = ( λ i , φ i ) . In the SAR image, corresponding points are searched by minimizing the geographic coordinate distance. The geographic coordinate distance is defined as follows:
ϵ = ( λ λ i ) 2 + ( φ φ i ) 2
where ( λ , φ ) are the geographic coordinates corresponding to pixels in the SAR image. The optimal matching point is found through a search and is as follows:
p i s = arg min ( u , v ) ϵ
where u ,   v are pixel coordinates in the SAR image. A search threshold ϵ th is set, and a match is considered valid only when ϵ < ϵ th , thereby establishing four groups of point-pair correspondences ( p i v , p i s ) i = 1 4 .
After obtaining point-pair correspondences, the affine transformation matrix is solved. The matrix form of 2D affine transformation is as follows:
u s v s 1 = a 11 a 12 t x a 21 a 22 t y 0 0 1 u v v v 1
where A = a 11 a 12 a 21 a 22 contains rotation, scale, and shear transformations, and t = [ t x , t y ] T is the translation vector. For each point pair ( p i v , p i s ) , where p i v = ( u i v , v i v ) and p i s = ( u i s , v i s ) , linear constraints can be written as follows:
u i s = a 11 u i v + a 12 v i v + t x
v i s = a 21 u i v + a 22 v i v + t y
Combining the constraints from four point pairs into matrix form b = M x , where x = [ a 11 , a 12 , t x , a 21 , a 22 , t y ] T is the parameter vector to be estimated, b is the observation vector, and M is the coefficient matrix. Solving through least squares we find the following:
x * = ( M T M ) 1 M T b
This yields the optimal affine transformation matrix H coarse . The registration quality of this matrix can be assessed through reprojection error, as follows:
E reproj = 1 N p i = 1 N p | H coarse p i v p i s | 2
where N p   =   4 is the number of point pairs. Although the reprojection error is typically on the order of several pixels due to GPS positioning errors and deviations from the ground plane assumption, coarse registration has effectively eliminated large-scale differences and large-angle rotations, narrowing the subsequent fine registration search range to a local neighborhood, as shown in Figure 4.
Figure 4. Schematic of the geographic information-constrained coarse registration process. Four trisection points are selected from the optical image, transformed to geographic coordinates via the imaging geometry-based coordinate transformation model, and matched with corresponding SAR pixels by minimizing geographic distance. The affine transformation matrix is then estimated through least squares optimization.

3.2. Multi-Scale Feature Response Aggregation Based on Phase Congruency

Although coarse registration effectively eliminates large-scale differences and rotation deviations, residual spatial deviations of several pixels remain due to simplifying assumptions and sensor measurement errors. Moreover, intrinsic radiometric differences between optical and SAR images make traditional intensity-based methods inapplicable. Phase congruency theory indicates that locations where different frequency components align in phase correspond to salient structures exhibiting good cross-modal consistency. Based on this principle, we propose a multi-scale feature response aggregation model to extract radiation-invariant feature points, as shown in Figure 5. Considering residual scale differences between images, a Gaussian scale-space pyramid is constructed with n o octave layers and n s scale sub-layers. The image at the i-th octave is obtained by downsampling the original image I ( x , y ) by a factor of 2 i :
I ( i ) ( x , y ) = I ( 2 i x , 2 i y ) , i = 0 , 1 , , n o 1
where I ( i ) is the image at the i-th octave, and the downsampling factor 2 i reduces the image size by half at each layer. Within each octave, multiple scale sub-layers are generated through Gaussian filtering with different scale parameters σ j :
L ( i ) ( x , y , σ j ) = G ( x , y , σ j ) I ( i ) ( x , y )
where the 2D Gaussian kernel is defined as follows:
G ( x , y , σ ) = 1 2 π σ 2 exp x 2 + y 2 2 σ 2
where σ is the standard deviation of the Gaussian kernel controlling the smoothing degree and denotes the convolution operation. The scale parameters increase geometrically:
σ j = σ 0 k j , j = 0 , 1 , , n s 1
where k = 2 1 / n s is the scale factor and σ 0 is the base scale.
Figure 5. Schematic of the multi-scale phase congruency-based feature response aggregation model. The process includes Gaussian pyramid construction, multi-direction gradient response computation, and feature aggregation through weighted summation and nonlinear transformation. The resulting feature response maps M(x, y) demonstrate cross-modal consistency between visible light and SAR images at structurally salient locations.
After constructing the multi-scale pyramid, gradient features are computed at each scale level. This paper employs Gaussian gradient operators to suppress speckle noise in SAR images. At scale σ , the Gaussian derivative kernels in the x and y directions are as follows:
G x ( x , y , σ ) = x σ 2 G ( x , y , σ )
G y ( x , y , σ ) = y σ 2 G ( x , y , σ )
Convolving with the image yields gradient responses, as follows:
I x ( i ) ( x , y , σ j ) = G x L ( i ) ( x , y , σ j )
I y ( i ) ( x , y , σ j ) = G y L ( i ) ( x , y , σ j )
where I x ( i ) and I y ( i ) are the gradient responses in the x and y directions at the i-th octave and scale σ j , respectively.
To extract multi-directional gradient information, n θ uniformly distributed orientation angles are defined, as follows:
θ k = 2 π k n θ , k = 0 , 1 , , n θ 1
The gradient response in orientation θ k is obtained through linear combination:
E θ k ( i ) ( x , y , σ j ) = cos ( θ k ) I x ( i ) ( x , y , σ j ) + sin ( θ k ) I y ( i ) ( x , y , σ j )
where E θ k ( i ) represents the gradient projection intensity in orientation θ k .
Based on the phase congruency principle, when multiple scales exhibit consistent high responses at the same location, that location likely corresponds to a salient image structure. To quantify this consistency, first compute the weighted sum of all scale responses at each location in orientation θ k , as follows:
F θ k ( x , y ) = j = 1 n s w j | E θ k ( i ) ( x , y , σ j ) |
where the weights w j adopt normalization based on response magnitude, as follows:
w j = | E θ k ( i ) ( x , y , σ j ) | j = 1 n s | E θ k ( i ) ( x , y , σ j ) |
where the numerator is the response magnitude at the current scale and the denominator is the sum of response magnitudes across all scales, ensuring that scales with strong responses receive higher weights. To enhance salient features and suppress noise, a nonlinear transformation is applied to the aggregation result, as follows:
F θ k ( x , y ) F θ k 2 ( x , y )
This squaring operation amplifies high response values while suppressing low-response noise. Integrating responses across all orientations yields the final feature response map, as follows:
M ( x , y ) = k = 1 n θ F θ k ( x , y )
where M ( x ,   y ) is the comprehensive feature response map, exhibiting high response values at image salient structures.
Based on the feature response map M ( x ,   y ) , local maxima are extracted as feature point candidates through non-maximum suppression (NMS). For each pixel ( x ,   y ) , if its response value is the maximum within an r × r neighborhood and exceeds a preset threshold, it is marked as a feature point, as follows:
K = ( x , y ) |   M ( x , y ) > τ feat M ( x , y ) = max ( x , y ) N r ( x , y ) M ( x , y )
where N r ( x , y ) denotes the x ,   y neighborhood centered at r × r , and τ feat is the response threshold.

3.3. Rotation-Invariant Feature Descriptor and Global-to-Local Hierarchical Search

The role of feature descriptors is to generate compact mathematical representations for each feature point, such that corresponding points from different images have similar descriptors, while different points have significantly different descriptors. Traditional SIFT descriptors based on gradient orientation histograms rely on the intensity consistency assumption and easily fail between optical and SAR images. This paper designs a rotation-invariant descriptor based on feature consistency mapping, which utilizes multi-scale multi-directional feature responses rather than raw intensity gradients, achieving robustness to radiometric variations, as shown in Figure 6. Meanwhile, to improve matching efficiency and accuracy, a global-to-local hierarchical search strategy is proposed that fully exploits geometric constraints provided by coarse registration.
Figure 6. Schematic of the rotation-invariant feature descriptor and global-to-local search (GLS) strategy. Left: Descriptor construction through support region definition, dominant orientation estimation, angular-radial spatial partitioning, and normalized orientation response vector generation. Right: GLS strategy comprising global search with multi-scale pyramid matching and RANSAC verification, followed by local optimization using predicted positions and bidirectional matching within constrained search windows.
To achieve rotation invariance, the dominant orientation needs to be determined for each feature point. With feature point ( x 0 ,   y 0 ) as the center, a circular support region with radius R is defined, as follows:
R = ( x , y ) | ( x x 0 ) 2 + ( y y 0 ) 2 R 2
where R is the support region radius, typically set to several times the feature point detection scale, e.g., R = 6 σ , where σ is the scale parameter corresponding to that feature point. The dominant orientation determination is based on the statistical distribution of gradient orientations within the support region. For each pixel within the support region, its gradient orientation on the feature response map is computed as follows:
θ ( x , y ) = arctan M y ( x , y ) M x ( x , y )
where M x and M y are the derivatives of the feature response map M in the x and y directions, respectively. All pixel gradient orientations within the support region are binned into an orientation histogram with n bin bins. Each pixel’s voting weight is the product of its gradient magnitude and a Gaussian weight, as follows:
w vote ( x , y ) = M x 2 ( x , y ) + M y 2 ( x , y ) exp ( x x 0 ) 2 + ( y y 0 ) 2 2 ( 1.5 σ ) 2
where the first term is the gradient magnitude and the second term is the Gaussian weight function with standard deviation 1.5 σ . The orientation corresponding to the bin with maximum response in the histogram is selected as the dominant orientation θ dom . After determining the dominant orientation, coordinate rotation transformation is applied to the support region, as follows:
x y = cos θ dom sin θ dom sin θ dom cos θ dom x x 0 y y 0
where x y are the rotated coordinates.
The rotated support region is partitioned spatially to capture local spatial structure information. A combination of angular and radial partitioning is adopted. Angular partitioning divides the circular region into 2 n o uniformly distributed sector regions along the orientation, each sector corresponding to an angular range of Δ θ = 2 π / ( 2 n o ) . Radial partitioning divides the circular region into n r concentric annular bands along the radial direction, with the inner and outer radii of the k-th band being as follows:
r inner ( k ) = ( k 1 ) R n r , r outer ( k ) = k R n r ,   k = 1 , 2 , , n r
The combination of the two partitioning schemes divides the support region into 2 n o × n r sub-regions, each denoted as R i , j .
For each sub-region R i , j , the feature responses in different orientations are statistically analyzed. With n θ uniformly distributed reference orientations defined, the orientation response vector for the sub-region is defined as h i , j = [ h i , j , 1 , h i , j , 2 , , h i , j , n θ ] T , where the k-th component is as follows:
h i , j , k = ( x , y ) R * i , j | E * θ k ( x , y ) | w g ( x , y )
where E θ k is the gradient response in orientation θ k , and the Gaussian weighting function is as follows:
w g ( x , y ) = exp ( x ) 2 + ( y ) 2 2 ( 0.5 R ) 2
Concatenating the orientation response vectors of all sub-regions in sequence yields the raw descriptor vector:
d raw = [ h 1 , 1 T , h 1 , 2 T , , h 2 n o , n r T ] T D
where the dimension is D = 2 n o × n r × n θ . To enhance the descriptor’s robustness to illumination variations, L2 normalization is first applied, as follows:
d norm = d raw | d raw | 2
Furthermore, the normalized vector is clipped, limiting components exceeding threshold t clip to that threshold, as follows:
d i min ( d i , t clip ) ,   i
where t clip is typically set to 0.2. L2 normalization is then applied again, yielding the final descriptor d D satisfying | d | 2 = 1 .
After extracting feature points from both images and constructing corresponding descriptors, a global-to-local hierarchical search strategy is employed to establish point-pair correspondences. In the global search stage, feature matching is performed independently at multiple levels of the multi-scale pyramid. For the ll l-th level of the pyramid, let the feature point sets of optical and SAR images be K v ( l ) and K s ( l ) , with corresponding descriptor sets d i v , ( l ) and d j s , ( l ) . For each optical feature point, compute the Euclidean distance between its descriptor and all SAR feature point descriptors, as follows:
d i j = | d i v , ( l ) d j s , ( l ) | 2
Select the point with minimum distance as the candidate match, as follows:
j * = arg min j d i j
Introducing Lowe’s ratio test, compute the ratio of nearest neighbor distance to second-nearest neighbor distance, as follows:
r i = d i , j * ( 1 ) d i , j * ( 2 )
where j * ( 1 ) and j * ( 2 ) correspond to the nearest and second-nearest neighbor indices, respectively. A match is considered reliable only when r i < r th , where r th is typically set to 0.8.
After independent matching at each level, RANSAC algorithm is used for geometric consistency verification. The inlier criterion is as follows:
| H ( l ) p i v , ( l ) p j s , ( l ) | 2 < ϵ inlier
where H ( l ) is the fitted affine matrix and ϵ inlier is the inlier threshold. Integrating information from multiple levels yields the global transformation matrix H global .
The local optimization stage utilizes geometric constraints obtained from global search to perform fine matching within local neighborhoods. For each feature point p i v , in the optical image, its corresponding position in the SAR image is predicted through the global transformation matrix, as follows:
p ^ i s = H global p i v
where p ^ i s is the predicted position. A square search window centered at p ^ i s with side length 2w × 2w is constructed, as follows:
W i = p | p p ^ i s | w
where | | denotes the infinity norm and w is the window half-width. Within the search window, a bidirectional matching strategy is adopted. Forward matching finds the nearest neighbor from optical to SAR, as follows:
j i * = arg min p j s W i | d i v d j s | 2
Reverse matching finds the nearest neighbor from SAR to optical, as follows:
i * = arg min i | d i v d j i * s | 2
Only when i * = i is the pair ( p i v , p j i * s ) considered a valid match. All valid matching point pairs are aggregated into set M local , and the RANSAC algorithm is applied again for geometric verification. The refined affine matrix is solved through weighted least squares, as follows:
H fine = arg min H ( p i v , p j s ) M final w i j | H p i v p j s | 2 2
where M final is the final inlier set and w i j are point pair weights. The final registration transformation matrix is obtained through composition of coarse and fine registration transformations, as follows:
H final = H fine H coarse
Thus, the proposed algorithm completes the entire process from coarse to fine registration. Large-scale differences and rotation deviations are eliminated through the imaging geometry-based coordinate transformation model with geographic information constraints, radiation-invariant features are extracted through multi-scale feature response aggregation based on phase congruency, and efficient and accurate point-pair matching is achieved through rotation-invariant feature consistency descriptors and a global-to-local hierarchical search strategy, ultimately realizing the high-precision cross-modal registration of optical and SAR images.

4. Experiments

4.1. Experimental Setup

4.1.1. Dataset

Data acquisition was conducted over the Qingshui Lake area near Hanshou County, Changde City, Hunan Province, China, using an integrated airborne optical/radar platform. The optical images were captured by a DJI Zenmuse P1 camera (45-megapixel full-frame sensor) at an average altitude of 340 m, providing a ground sampling distance (GSD) of approximately 0.15 m/pixel. The optical images have dimensions of 2784 × 2796 pixels, covering an approximately 418 m × 419 m ground area. The SAR images were acquired by a Ka-band miniaturized SAR system with 0.30 m/pixel slant range resolution and 1.5 m azimuth resolution, operating in stripmap mode with side-looking geometry (incidence angle: 45–55°). The optical image data include the UAV’s pose information at the time of image capture, comprising longitude, latitude, absolute altitude, relative altitude, and heading angle, as well as the gimbal’s roll, yaw, and pitch angles during image acquisition, as shown in Figure 7.
Figure 7. Example of optical image data format. (a) Optical image with associated UAV pose and gimbal attitude metadata, including GPS coordinates (longitude, latitude), absolute and relative altitudes, gimbal angles (roll, yaw, pitch), and UAV heading angle. (b) Demonstration of pixel-to-geographic coordinate mapping, showing the correspondence between pixel coordinates (1635, 1537) in the image plane and their computed geographic coordinates (longitude: 111.9417576°, latitude: 28.8110523°) for an image of size 2784 × 2796 pixels.
Because the SAR radar synchronously records the platform’s pose information during data acquisition, the SAR image data contain geographic location information for each pixel.

4.1.2. Implementation Details

The hardware environment consists of an NVIDIA GeForce RTX 3070 Laptop GPU, with MATLAB R2020a as the software platform running on Windows 11. For the coarse registration stage, the positions of the registration points are set at four trisection points: (640, 360), (1280, 360), (640, 720), and (1280, 720). In the fine registration stage, a 4-level multi-scale pyramid is employed, with the number of Gaussian filter orientations set to 6 and the feature descriptor region size set to 72 pixels. The maximum number of extracted feature points is set to 5000 for all algorithms, and the error threshold in the matching process is uniformly set to 3 pixels.

4.1.3. Evaluation Metrics

Three quantitative evaluation metrics are employed for objective assessment: Correct matching rate (CMR), root mean square error (RMSE), and algorithm execution time. CMR represents the ratio of correctly matched point pairs to total matched pairs in an image pair, as follows:
C M R = Correctly Matched Pairs Total Matched Pairs
RMSE measures the discrepancy between predicted matching point positions and their true locations, as follows:
R M S E = 1 N i = 1 N ( x i x ^ i ) 2
To obtain reliable ground truth correspondences for RMSE calculation, we employ a semi-automatic annotation process: (1) Expert annotators manually identify 20–30 distinctive landmarks visible in both optical and SAR images (e.g., building corners, road intersections, bridge endpoints); (2) initial pixel locations are refined to sub-pixel accuracy using template matching with normalized cross-correlation (NCC threshold > 0.7); and (3) only correspondences with mutual agreement from two independent annotators are retained. This procedure yields high-quality ground truth with estimated sub-pixel localization accuracy.

4.2. Coarse Registration Experiments

First, we analyze the effectiveness of the geographic information-constrained coarse registration algorithm for optical and SAR images. The geographic location information is computed for four registration points in each sample optical image, with the geographic coordinates recorded in Table 1. To validate the statistical significance of our method’s performance improvement, we conduct paired t-tests comparing our approach against the best-performing baseline methods across the three image groups. Table 1 presents the statistical analysis results.
Table 1. Statistical significance test results comparing the proposed method with state-of-the-art baselines.
RMSE differences are calculated as (Baseline RMSE − Ours RMSE) averaged across three image groups. Positive values indicate that our method achieves lower RMSE. The paired t-test confirms that our method’s performance improvement over all baseline methods is statistically significant at the 0.05 level (p < 0.05). The smallest improvement is against GLS-MIFT (0.04 pixels, p = 0.041), while the largest is against SuperGlue (1.29 pixels, p < 0.001), demonstrating that our geographic information-constrained approach significantly outperforms both traditional handcrafted methods and deep learning-based approaches.
The statistical significance of our method’s performance advantage (Table 1) provides strong evidence that the proposed coarse-to-fine framework with geographic information constraints achieves robust and reliable registration. Based on the geographic location information derived from pixel positions in the optical images, corresponding points with consistent latitude and longitude are located in the SAR images, thereby obtaining four groups of matching point pairs to achieve preliminary registration, as shown in Figure 8. The registration results demonstrate that the geographic information-constrained approach effectively establishes spatial correspondence between optical and SAR images by incorporating geographic location information as constraints, achieving coarse cross-modal alignment. However, due to differences in sensor imaging principles, projection distortions caused by terrain undulations, and localization errors inherent in the monocular visual positioning model, the registration results still exhibit certain nonlinear geometric deviations, manifested as local position shifts, edge misalignments, and micro-scale deformations.
Figure 8. Coarse registration results for three image groups. For each group: (a) shows the pre-coarse registration SAR image, (b) shows the optical (visible light) reference image, and (c) shows the coarsely registered SAR image. Red boxes highlight corresponding regions for comparison, demonstrating the effectiveness of geographic information-constrained coarse registration in eliminating large-scale differences and rotation deviations.

4.3. Fine Registration Experiments

In this section, the proposed cross-modal feature consistency mapping-based fine registration algorithm is compared with four state-of-the-art registration methods (3MRS, OS-SIFT, POS-GIFT, and GLS-MIFT) on three groups of coarsely registered images, as shown in Figure 9. All comparison algorithms employ unified parameter configurations.
Figure 9. Feature matching results of five registration algorithms on three image groups. Each row shows results from a different algorithm (from top to bottom: 3MRS, OS-SIFT, POS-GIFT, GLS-MIFT, and the proposed method), and each column corresponds to one image group. Yellow lines indicate matched feature point pairs between optical and SAR images.
The registration results reveal that the 3MRS algorithm achieves the fewest correct matching point pairs among the five algorithms, exhibiting the poorest registration performance, as shown in Figure 10. This deficiency primarily stems from 3MRS’s insufficient robustness to geometric transformations such as rotation, scaling, and perspective distortions in images. The OS-SIFT algorithm obtains a relatively larger number of correct matching point pairs, yet still falls short compared with POS-GIFT, GLS-MIFT, and the proposed method, as OS-SIFT fails to adequately incorporate geometric constraints during global search matching.
Figure 10. Checkerboard visualizations of registration results for five algorithms on three image groups. Each row corresponds to one image group, and columns from left to right show results of 3MRS, OS-SIFT, POS-GIFT, GLS-MIFT, and the proposed method. Red boxes highlight regions with noticeable misalignment for detailed comparison.
Among the deep learning-based methods, SuperGlue and LoFTR, which have achieved state-of-the-art performance on homogeneous image matching benchmarks, show significantly degraded performance on cross-modal SAR-optical pairs. SuperGlue achieves only 55/5000 average CMR with 3.29 pixels RMSE, while LoFTR obtains 76/5000 average CMR with 2.99 pixels RMSE, as shown in Table 2. This performance degradation can be attributed to the substantial domain gap between their training data (predominantly optical–optical pairs) and the target SAR-optical modality. The learned feature representations optimized for intensity-consistent image pairs fail to generalize to cross-modal scenarios with nonlinear radiometric distortions.
Table 2. Comparative experimental results of ten optical-SAR image registration algorithms including both handcrafted feature-based methods and deep learning-based methods. Correct matching rate (CMR) indicates the number of correctly matched point pairs out of 5000 maximum features; Root mean square error (RMSE) measures registration accuracy in pixels; time represents algorithm execution time in seconds. Bold indicates the best performance.
ReDFeat and SAROptNet, which are specifically designed for multimodal image matching, demonstrate improved performance over general-purpose deep learning methods. ReDFeat achieves 139/5000 average CMR with 2.45 pixels RMSE, while SAROptNet obtains 158/5000 average CMR with 2.28 pixels RMSE. However, both methods still underperform compared with the best handcrafted methods (POS-GIFT, GLS-MIFT) and our proposed approach. This observation suggests that current deep learning methods, even those tailored for multimodal matching, have not fully addressed the fundamental challenges of SAR-optical registration.
The checkerboard visualizations demonstrate that the 3MRS algorithm exhibits severe registration failures across all three image groups. Although OS-SIFT successfully achieves basic alignment for the first and third image groups, position offsets exceeding 10 pixels remain in the marked regions. POS-GIFT achieves accurate alignment for the first image group, with only slight misalignments at junction areas for the second and third groups. GLS-MIFT accurately registers the second and third image groups, while exhibiting minor deviations in the first group. The deep learning methods (SuperGlue, LoFTR, ReDFeat, SAROptNet) exhibit inconsistent alignment quality across different image groups, with visible misalignments, particularly in regions with strong radiometric differences such as water bodies and vegetation boundaries. In contrast, the proposed algorithm successfully achieves accurate alignment for all three image groups.
The experimental results indicate that the average CMR values for the five algorithms across three image pairs range from 30/5000 (3MRS) to 261/5000 (POS-GIFT) for handcrafted methods, and from 55/5000 (SuperGlue) to 158/5000 (SAROptNet) for deep learning methods, while our method achieves 179/5000 (CPU) and 198/5000 (GPU-accelerated). Although the proposed algorithm does not achieve the highest average CMR among handcrafted methods, the checkerboard visualizations demonstrate that the correctly matched points extracted by our algorithm are sufficient to meet the requirements for accurate image registration. Notably, our method significantly outperforms all deep learning-based approaches in terms of CMR, demonstrating the effectiveness of combining geographic information constraints with cross-modal feature consistency mapping.
In terms of average RMSE, with the inlier threshold set to 3 pixels, the results for handcrafted methods range from 2.69 (3MRS) to 2.04 pixels (GLS-MIFT), deep learning methods range from 3.29 (SuperGlue) to 2.28 pixels (SAROptNet), and our method achieves 2.00 pixels (CPU) and 1.97 pixels (GPU-accelerated), with our method achieving the best performance. The deep learning methods exhibit higher RMSE values compared with most handcrafted methods, further confirming that these approaches struggle with cross-modal radiometric variations. Even the best-performing deep learning method (SAROptNet, 2.28 pixels) shows 14.0% higher RMSE than our CPU version (2.00 pixels).
Regarding computational efficiency, the proposed algorithm’s average execution time is 35.69 s less than POS-GIFT (which ranks first in average CMR among handcrafted methods), representing a 37.0% reduction. While deep learning methods achieve faster inference times (2.34–5.87 s with GPU acceleration), they require dedicated GPU hardware, which may not be available on resource-constrained UAV platforms. Our GPU-accelerated version achieves comparable speed (1.89 s) while delivering significantly better registration accuracy. For CPU-only deployments, our method provides the best trade-off between accuracy and computational cost.
Considering the matching connection diagrams, checkerboard visualizations, and performance metrics comprehensively, the proposed algorithm significantly improves computational efficiency while maintaining high matching accuracy. This superior performance stems from two primary factors: First, the algorithm achieves cross-modal feature consistency through rotation and scale invariance, effectively handling complex geometric variations; second, the proposed feature consistency descriptor, with its relatively simple structure based on circular region partitioning centered at feature points and dominant orientation estimation, not only reduces computational complexity but also enhances descriptor stability; third, the geographic information-constrained coarse registration effectively reduces the search space and eliminates large-scale geometric differences, which is particularly advantageous over deep learning methods that attempt to match features globally without such geometric priors.
The comparison with deep learning methods reveals several important insights: (1) General-purpose learned matchers (SuperGlue, LoFTR) trained on optical image pairs do not generalize well to SAR-optical matching, highlighting the unique challenges of cross-modal registration; (2) even specialized multimodal methods (ReDFeat, SAROptNet) underperform compared with well-designed handcrafted approaches that explicitly address radiometric invariance; (3) the integration of geographic information constraints provides a modality-invariant bridge that effectively circumvents the radiometric differences challenging both traditional and learning-based methods. These findings suggest that, for UAV-based SAR-optical registration with available navigation data, leveraging geometric constraints combined with radiation-invariant feature descriptors remains more effective than purely data-driven approaches.

4.4. Ablation Studies

To validate the effectiveness of each component in the proposed algorithm, ablation experiments are conducted on three groups of optical and SAR image pairs. The experiments analyze two dimensions—a coarse registration stage and a fine registration stage—systematically evaluating the contributions of the geographic information constraint module, multi-scale feature pyramid, feature response aggregation model, and feature consistency descriptor to registration performance. Additionally, we investigate how the proposed coarse registration strategy benefits deep learning-based methods, providing insights into the complementary relationship between geometric constraints and learned feature matching.

4.4.1. Coarse Registration Stage Ablation

The core of the proposed coarse registration algorithm lies in introducing geographic information as constraints. To verify the effectiveness of this constraint, the following comparative experiments are designed: (1) Direct fine registration without geographic information constraints; (2) fine registration after coarse registration with geographic information constraints; (3) deep learning methods (SuperGlue, LoFTR) without coarse registration; (4) deep learning methods with our proposed coarse registration as preprocessing. This comprehensive comparison not only validates the necessity of our coarse registration stage but also demonstrates its generalizability as a preprocessing module for other matching methods.
The experimental results demonstrate that direct fine registration without geographic information constraints struggles to find correct matching point pairs globally due to significant scale differences and rotation deviations between optical and SAR images, resulting in an average CMR of only 12/5000 and an average RMSE as high as 4.88 pixels. With geographic information constraints introduced, the average CMR improves to 179/5000, the average RMSE decreases to 2.00 pixels, and the average execution time is reduced by 32.0%, as shown in Table 3. This fully validates the effectiveness of geographic information constraints in eliminating scale differences and rotation deviations between images.
Table 3. Ablation study results validating the necessity of geographic information constraints. “Fine Only” denotes direct fine registration without coarse alignment; “Coarse + Fine” denotes the proposed two-stage strategy. Deep learning methods are also included for comparison under the same coarse registration setting. Bold indicates the best performance in each metric.

4.4.2. Fine Registration Stage Ablation

For the fine registration stage, the following ablation experiments are designed to verify the contribution of each component: (1) Baseline method: single-scale feature extraction only; (2) adding multi-scale feature pyramid; (3) adding feature response aggregation model; and (4) adding feature consistency descriptor. All experiments are conducted on coarsely registered images. Additionally, we compare each ablation variant with deep learning methods (SuperGlue, LoFTR) under the same coarse registration setting to provide reference points for understanding the contribution of each component.
The experimental results indicate that each component contributes positively to registration performance, as shown in Table 4. Specifically, the multi-scale feature pyramid improves the average CMR from 85/5000 to 126/5000 and reduces the RMSE from 2.56 to 2.31 pixels by detecting feature points in different scale spaces. The feature response aggregation model further improves the average CMR to 158/5000 and reduces the RMSE to 2.18 pixels by fusing multi-directional, multi-scale gradient features. The feature consistency descriptor ultimately achieves an average RMSE of 2.00 pixels by constructing rotation-invariant feature representations, realizing sub-pixel registration accuracy.
Table 4. Ablation study evaluating the contribution of multi-scale pyramid, feature response aggregation, and feature consistency descriptor in the fine registration stage. All experiments are conducted on coarsely registered images.

4.4.3. Feature Descriptor Parameter Analysis

To further analyze the impact of design parameters of the feature consistency descriptor on registration performance, parameter sensitivity experiments are conducted on the number of angular partitions and radial subdivisions. To provide comprehensive context, we also compare our descriptor configurations with deep learning methods (SuperGlue, LoFTR) and the classical SIFT descriptor in terms of both performance and descriptor dimensionality, as shown in Table 5.
Table 5. Sensitivity analysis of descriptor parameters: effects of angular partitions and radial subdivisions on registration accuracy and descriptor dimensionality. A comparison with deep learning feature dimensions is also provided.
The experimental results show that registration accuracy gradually improves as the number of partitions and subdivisions increases, but the descriptor dimension also increases accordingly, leading to higher matching computational costs. Considering both registration accuracy and computational efficiency, we select 12 angular partitions (corresponding to n o = 6 ) and 3 radial subdivisions as the optimal parameter configuration. With this setting, the descriptor dimension is 216, maintaining high registration accuracy while achieving fast execution speed.
The parameter sensitivity analysis also reveals that even our smallest configuration (8 angular partitions, 2 radial subdivisions, 96 dimensions) achieves a competitive performance (142/5000 CMR, 2.23 pixels RMSE) that approaches SuperGlue (147/5000 CMR) while using only 37.5% of its descriptor dimension. This suggests that our phase congruency-based feature representation is highly efficient in encoding cross-modal structural information. For applications with strict computational constraints, smaller descriptor configurations can be employed with acceptable performance degradation, providing flexibility for deployment on various UAV platforms with different processing capabilities.

4.4.4. GLS Strategy Validation

To validate the effectiveness of the global-to-local search (GLS) strategy-based feature matcher, it is compared with traditional brute-force matching and FLANN matching methods. Furthermore, we include comparisons with state-of-the-art learning-based matchers (SuperGlue, LoFTR, LightGlue) and explore hybrid approaches that combine our GLS strategy with learning-based refinement to investigate potential complementary benefits, the results are shown in Table 6.
Table 6. Performance comparison of feature matching strategies on coarsely registered image pairs, including both traditional and learning-based matchers.
The experimental results demonstrate that the GLS strategy achieves higher CMR and lower RMSE than both brute-force and FLANN matching through its hierarchical matching mechanism. By obtaining the initial affine transformation matrix in the global search stage and significantly reducing the matching search space through constructing local search windows in the local optimization stage, the GLS strategy achieves superior performance.
The hybrid approaches demonstrate interesting complementary effects between geometric constraints and learned matching. Combining our coarse registration with LightGlue (Coarse + LightGlue) improves performance to 168/5000 CMR and 2.07 pixels RMSE, surpassing standalone LightGlue by 10.5% in CMR and 5.7% in RMSE. More notably, the GLS + LightGlue refinement strategy, which applies LightGlue as a refinement step after our GLS matching, achieves the best overall performance (185/5000 CMR, 1.9534 pixels RMSE) among all evaluated methods. This hybrid approach improves upon standalone GLS by 3.3% in CMR and 2.5% in RMSE, suggesting that learned features can provide complementary information for refining the matches obtained by our geometric constraint-based approach. However, this improvement comes at the cost of requiring GPU hardware for the LightGlue refinement step. For UAV applications without GPU support, our standalone GLS strategy remains the optimal choice, achieving the best performance among all CPU-only methods while maintaining practical computational efficiency.

5. Discussion

The comprehensive experimental evaluation reveals several important insights regarding SAR-optical image registration. First, the comparison with deep learning methods demonstrates that general-purpose learned matchers (SuperGlue, LoFTR) trained on optical image pairs exhibit significant performance degradation on cross-modal matching tasks, with average RMSE exceeding 2.99 pixels without coarse registration, highlighting the fundamental challenge of bridging the radiometric gap between modalities through purely data-driven approaches. Second, our geographic information-constrained coarse registration proves to be a generalizable preprocessing module that benefits not only our fine registration algorithm but also deep learning methods, improving SuperGlue’s CMR by 167% and LoFTR’s by 109% when applied as preprocessing. Third, the ablation studies confirm that each component of our framework contributes meaningfully to the final performance, with the multi-scale pyramid addressing residual scale variations, the feature response aggregation capturing phase-consistent structural information, and the rotation-invariant descriptor ensuring geometric robustness. Fourth, the parameter sensitivity analysis demonstrates that our descriptor achieves superior performance with moderate dimensionality (216-D) compared with deep learning features (256-D), indicating efficient utilization of descriptor capacity for encoding cross-modal invariant information. Finally, while deep learning methods offer faster inference with GPU acceleration, our CPU-based approach provides the best accuracy–efficiency trade-off for resource-constrained UAV platforms, and the hybrid GLS + LightGlue refinement strategy achieves the overall best performance (185/5000 CMR, 1.95 pixels RMSE) when GPU resources are available, suggesting promising directions for future research combining geometric constraints with learned representations.

6. Conclusions

This paper proposes a coarse-to-fine optical-SAR image registration algorithm integrating geographic information constraints with cross-modal feature consistency mapping to address nonlinear radiometric distortions and geometric deformations. The coarse stage leverages airborne navigation data to eliminate scale and rotation differences through imaging geometry-based coordinate transformation, while the fine stage employs a phase congruency-based feature response aggregation model combined with rotation-invariant descriptors and global-to-local hierarchical search for sub-pixel alignment. Comprehensive experiments on integrated airborne optical/SAR datasets demonstrate superior performance with an average RMSE of 2.00 pixels, outperforming both four traditional handcrafted methods (3MRS, OS-SIFT, POS-GIFT, GLS-MIFT) and four state-of-the-art deep learning approaches (SuperGlue, LoFTR, ReDFeat, SAROptNet), while reducing execution time by 37.0% compared with POS-GIFT. The ablation studies confirm the effectiveness of each component and demonstrate that the proposed coarse registration serves as a generalizable preprocessing module that significantly improves deep learning matchers’ performance. The hybrid GLS + LightGlue refinement strategy achieves the best overall performance (1.95 pixels RMSE), indicating complementary benefits between geometric constraints and learned features. The algorithm exhibits robust performance under challenging conditions including dense smoke and fire, validating its effectiveness for real-time UAV-based multi-sensor fusion applications. Future research directions include developing end-to-end trainable frameworks that jointly optimize geometric constraints and learned feature representations, exploring self-supervised learning strategies to reduce dependence on labeled cross-modal training data, extending the framework to complex urban scenarios with significant elevation variations and occlusions, and investigating lightweight network architectures suitable for onboard UAV deployment.

Author Contributions

Conceptualization, X.S., Z.Z. and X.G.; methodology, X.S., Z.Z. and X.G.; software, X.S.; validation, X.S.; formal analysis, X.S. and Z.Z.; investigation, X.S. and R.G.; resources, Z.Z., X.G. and X.S.; data curation, X.S. and R.G.; writing—original draft preparation, X.S.; writing—review and editing, Z.Z., X.G., X.L. and S.S.; visualization, X.S. and P.Z.; supervision, Z.Z., X.G. and S.S.; project administration, Z.Z. and X.S. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

The experimental datasets used in this study were collected using proprietary airborne optical/SAR platforms and are not publicly available due to confidentiality restrictions. Representative sample images and detailed data acquisition parameters can be obtained from the corresponding author upon reasonable request.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Sommervold, O.; Gazzea, M.; Arghandeh, R. A Survey on SAR and Optical Satellite Image Registration. Remote Sens. 2023, 15, 850. [Google Scholar] [CrossRef] [Scilit]
  2. Zhu, B.; Zhou, L.; Pu, S.; Fan, J.; Ye, Y. Advances and Challenges in Multimodal Remote Sensing Image Registration. IEEE J. Miniat. Air Space Syst. 2023, 4, 165–174. [Google Scholar] [CrossRef] [Scilit]
  3. Zhang, X.; Leng, C.; Hong, Y.; Pei, Z.; Cheng, I.; Basu, A. Multimodal Remote Sensing Image Registration Methods and Advancements: A Survey. Remote Sens. 2021, 13, 5128. [Google Scholar] [CrossRef] [Scilit]
  4. Kulkarni, S.C.; Rege, P.P. Pixel Level Fusion Techniques for SAR and Optical Images: A Review. Inf. Fusion 2020, 59, 13–29. [Google Scholar] [CrossRef] [Scilit]
  5. Wang, J.; Yang, H.; He, Y.; Zheng, F.; Liu, Z.; Chen, H. An Unpaired SAR-to-Optical Image Translation Method Based on Schrödinger Bridge Network and Multi-Scale Feature Fusion. Sci. Rep. 2024, 14, 27047. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  6. Vivone, G.; Deng, L.-J.; Deng, S.; Hong, D.; Jiang, M.; Li, C.; Li, W.; Shen, H.; Wu, X.; Xiao, J.-L.; et al. Deep Learning in Remote Sensing Image Fusion: Methods, Protocols, Data, and Future Perspectives. IEEE Geosci. Remote Sens. Mag. 2025, 13, 269–310. [Google Scholar] [CrossRef] [Scilit]
  7. Huang, M.; Xu, Y.; Qian, L.; Shi, W.; Zhang, Y.; Bao, W.; Wang, N.; Liu, X.; Xiang, X. A Bridge Neural Network-Based Optical-SAR Image Joint Intelligent Interpretation Framework. Space Sci. Technol. 2021, 2021, 9841456. [Google Scholar] [CrossRef] [Scilit]
  8. Brown, L.G. A Survey of Image Registration Techniques. ACM Comput. Surv. 1992, 24, 325–376. [Google Scholar] [CrossRef] [Scilit]
  9. Zitová, B.; Flusser, J. Image Registration Methods: A Survey. Image Vis. Comput. 2003, 21, 977–1000. [Google Scholar] [CrossRef] [Scilit]
  10. Chen, S.; Mei, L. Structure Similarity Virtual Map Generation Network for Optical and SAR Image Matching. Front. Phys. 2024, 12, 1287050. [Google Scholar] [CrossRef] [Scilit]
  11. Zhang, W. Robust Registration of SAR and Optical Images Based on Deep Learning and Improved Harris Algorithm. Sci. Rep. 2022, 12, 5901. [Google Scholar] [CrossRef] [Scilit]
  12. Yang, X.; Wang, Z.; Zhao, J.; Yang, D. FG-GAN: A Fine-Grained Generative Adversarial Network for Unsupervised SAR-to-Optical Image Translation. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5621211. [Google Scholar] [CrossRef] [Scilit]
  13. Li, H.; Gu, C.; Wu, D.; Cheng, G.; Guo, L.; Liu, H. Multiscale Generative Adversarial Network Based on Wavelet Feature Learning for SAR-to-Optical Image Translation. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5236115. [Google Scholar] [CrossRef] [Scilit]
  14. Zhang, K.; Yu, A.; Tong, W.; Dong, Z. A Robust SAR-Optical Heterologous Image Registration Method Based on Region-Adaptive Keypoint Selection. Remote Sens. 2024, 16, 3289. [Google Scholar] [CrossRef] [Scilit]
  15. Tan, S.; Duan, Z.; Pu, L. Multi-Scale Object Detection in UAV Images Based on Adaptive Feature Fusion. PLoS ONE 2024, 19, e0300120. [Google Scholar] [CrossRef] [Scilit]
  16. Samaras, S.; Diamantidou, E.; Ataloglou, D.; Sakellariou, N.; Vafeiadis, A.; Magoulianitis, V.; Lalas, A.; Dimou, A.; Zarpalas, D.; Votis, K.; et al. Deep Learning on Multi Sensor Data for Counter Uav Applications—A Systematic Review. Sensors 2019, 19, 4837. [Google Scholar] [CrossRef] [Scilit]
  17. Yao, H.; Qin, R.; Chen, X. Unmanned Aerial Vehicle for Remote Sensing Applications—A Review. Remote Sens. 2019, 11, 1443. [Google Scholar] [CrossRef] [Scilit]
  18. Goforth, H.; Lucey, S. GPS-Denied UAV Localization Using Pre-Existing Satellite Imagery. In Proceedings of the 2019 International Conference on Robotics and Automation, ICRA 2019, Montreal, QC, Canada, 20–24 May 2019; IEEE: New York, NY, USA, 2019; pp. 2974–2980. [Google Scholar]
  19. Chen, S.; Wu, X.; Mueller, M.W.; Sreenath, K. Real-Time Geo-Localization Using Satellite Imagery and Topography for Unmanned Aerial Vehicles. In Proceedings of the 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems, IROS 2021, Prague, Czech Republic, 27 September–1 October 2021; IEEE: New York, NY, USA, 2021; pp. 2275–2281. [Google Scholar]
  20. Lowe, D.G. Distinctive Image Features from Scale-Invariant Keypoints. Int. J. Comput. Vis. 2004, 60, 91–110. [Google Scholar] [CrossRef] [Scilit]
  21. Bay, H.; Ess, A.; Tuytelaars, T.; Van Gool, L. Speeded-Up Robust Features (SURF). Comput. Vis. Image Underst. 2008, 110, 346–359. [Google Scholar] [CrossRef] [Scilit]
  22. Sedaghat, A.; Mokhtarzade, M.; Ebadi, H. Uniform Robust Scale-Invariant Feature Matching for Optical Remote Sensing Images. IEEE Trans. Geosci. Remote Sens. 2011, 49, 4516–4527. [Google Scholar] [CrossRef] [Scilit]
  23. Mikolajczyk, K.; Schmid, C. A Performance Evaluation of Local Descriptors. IEEE Trans. Pattern Anal. Mach. Intell. 2005, 27, 1615–1630. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  24. Zhang, W.; Zhao, Y. An Improved SIFT Algorithm for Registration between SAR and Optical Images. Sci. Rep. 2023, 13, 6346. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  25. Ye, Y.; Shan, J.; Bruzzone, L.; Shen, L. Robust Registration of Multimodal Remote Sensing Images Based on Structural Similarity. IEEE Trans. Geosci. Remote Sens. 2021, 55, 2941–2958. [Google Scholar] [CrossRef] [Scilit]
  26. Kovesi, P. Image Features from Phase Congruency. J. Comput. Vis. Res. 1995, 1, 1–26. [Google Scholar]
  27. Ma, W.; Wu, Y.; Liu, S.; Su, Q.; Zhong, Y. Remote Sensing Image Registration Based on Phase Congruency Feature Detection and Spatial Constraint Matching. IEEE Access 2018, 6, 77554–77567. [Google Scholar] [CrossRef] [Scilit]
  28. Fan, B.; Wu, F.; Hu, Z. Rotationally Invariant Descriptors Using Intensity Order Pooling. IEEE Trans. Pattern Anal. Mach. Intell. 2012, 34, 2031–2045. [Google Scholar] [CrossRef] [Scilit]
  29. Ma, J.; Jiang, X.; Fan, A.; Jiang, J.; Yan, J. Image Matching From Handcrafted to Deep Features: A Survey. Int. J. Comput. Vis. 2021, 129, 23–79. [Google Scholar] [CrossRef] [Scilit]
  30. Viola, P.; Wells, W.M. Alignment by Maximization of Mutual Information. In Proceedings of the IEEE International Conference on Computer Vision; IEEE: New York, NY, USA, 1995; pp. 16–23. [Google Scholar]
  31. Cole-Rhodes, A.A.; Johnson, K.L.; LeMoigne, J.; Zavorin, L. Multiresolution Registration of Remote Sensing Imagery by Optimization of Mutual Information Using a Stochastic Gradient. IEEE Trans. Image Process. 2003, 12, 1495–1510. [Google Scholar] [CrossRef]
  32. Pluim, J.P.W.; Maintz, J.B.A.A.; Viergever, M.A. Mutual-Information-Based Registration of Medical Images: A Survey. IEEE Trans. Med. Imaging 2003, 22, 986–1004. [Google Scholar] [CrossRef] [Scilit]
  33. Maes, F.; Vandermeulen, D.; Suetens, P. Comparative Evaluation of Multiresolution Optimization Strategies for Multimodality Image Registration by Maximization of Mutual Information. Med. Image Anal. 1999, 3, 373–386. [Google Scholar] [CrossRef] [Scilit]
  34. Suri, S.; Reinartz, P. Mutual-Information-Based Registration of TerraSAR-X and Ikonos Imagery in Urban Areas. IEEE Trans. Geosci. Remote Sens. 2010, 48, 939–949. [Google Scholar] [CrossRef] [Scilit]
  35. Xiong, X.; Jin, G.; Xu, Q.; Zhang, H. Robust SAR Image Registration Using Rank-Based Ratio Self-Similarity. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2021, 14, 2358–2368. [Google Scholar] [CrossRef] [Scilit]
  36. Hel-Or, Y.; Hel-Or, H.; David, E. Matching by Tone Mapping: Photometric Invariant Template Matching. IEEE Trans. Pattern Anal. Mach. Intell. 2014, 36, 317–330. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  37. Pallotta, L.; Giunta, G.; Clemente, C. SAR Image Registration in the Presence of Rotation and Translation: A Constrained Least Squares Approach. IEEE Geosci. Remote Sens. Lett. 2021, 18, 1595–1599. [Google Scholar] [CrossRef] [Scilit]
  38. Harris, C.; Stephens, M. A Combined Corner and Edge Detector. In Proceedings of the Alvey Vision Conference 1988; Alvey Vision Club: Manchester, UK, 1988; pp. 23.1–23.6. [Google Scholar]
  39. Ke, Y.; Sukthankar, R. PCA-SIFT: A More Distinctive Representation for Local Image Descriptors. In Proceedings of the 2004 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2004. CVPR 2004, Washington, DC, USA, 27 June–2 July 2004; Volume 2, p. II. [Google Scholar]
  40. Mikolajczyk, K.; Tuytelaars, T.; Schmid, C.; Zisserman, A.; Matas, J.; Schaffalitzky, F.; Kadir, T.; Van Gool, L. A Comparison of Affine Region Detectors. Int. J. Comput. Vis. 2005, 65, 43–72. [Google Scholar] [CrossRef] [Scilit]
  41. Dellinger, F.; Delon, J.; Gousseau, Y.; Michel, J.; Tupin, F. SAR-SIFT: A SIFT-like Algorithm for SAR Images. IEEE Trans. Geosci. Remote Sens. 2015, 53, 453–466. [Google Scholar] [CrossRef] [Scilit]
  42. Xiang, Y.; Wang, F.; You, H. OS-SIFT: A Robust SIFT-Like Algorithm for High-Resolution Optical-to-SAR Image Registration in Suburban Areas. IEEE Trans. Geosci. Remote Sens. 2018, 56, 3078–3090. [Google Scholar] [CrossRef] [Scilit]
  43. Fan, J.; Wu, Y.; Wang, F.; Zhang, Q.; Liao, G.; Li, M. SAR Image Registration Using Phase Congruency and Nonlinear Diffusion-Based SIFT. IEEE Geosci. Remote Sens. Lett. 2015, 12, 562–566. [Google Scholar] [CrossRef] [Scilit]
  44. Kovesi, P. Phase Congruency Detects Corners and Edges. In Digital Image Computing: Techniques and Applications 2003; CSIRO: Sydney, Australia, 2003. [Google Scholar]
  45. Li, J.; Hu, Q.; Ai, M. RIFT: Multi-Modal Image Matching Based on Radiation-Variation Insensitive Feature Transform. IEEE Trans. Image Process. 2020, 29, 3296–3310. [Google Scholar] [CrossRef] [Scilit]
  46. Li, J.; Shi, P.; Hu, Q.; Zhang, Y. RIFT2: Speeding-up RIFT with A New Rotation-Invariance Technique. arXiv 2023, arXiv:2303.00319. [Google Scholar] [CrossRef] [Scilit]
  47. Ye, Y.; Shan, J.; Hao, S.; Bruzzone, L.; Qin, Y. A Local Phase Based Invariant Feature for Remote Sensing Image Matching. ISPRS J. Photogramm. Remote Sens. 2018, 142, 205–221. [Google Scholar] [CrossRef] [Scilit]
  48. Aguilera, C.A.; Aguilera, F.J.; Sappa, A.D.; Toledo, R. Learning Cross-Spectral Similarity Measures with Deep Convolutional Neural Networks. In Proceedings of the 29th IEEE Conference on Computer Vision and Pattern Recognition Workshops, CVPRW 2016, Las Vegas, NV, USA, 26 June–1 July 2016; IEEE Computer Society: New York, NY, USA, 2016; pp. 267–275. [Google Scholar]
  49. Fernández Alcantarilla, P.; Bartoli, A.; Davison, A. KAZE Features. In Proceedings of the 12th European Conference on Computer Vision, ECCV 2012, Florence, Italy, 7–13 October 2012; Springer: Berlin/Heidelberg, Germany, 2012; pp. 214–227. [Google Scholar]
  50. Alcantarilla, P.F.; Nuevo, J.; Bartoli, A. Fast Explicit Diffusion for Accelerated Features in Nonlinear Scale Spaces. In Proceedings of the 24th British Machine Vision Conference, BMVC 2013, Bristol, UK, 9–13 September 2013; British Machine Vision Association, BMVA: Bristol, UK, 2013. [Google Scholar]
  51. Wang, B.; Zhang, J.; Lu, L.; Huang, G.; Zhao, Z. A Uniform SIFT-Like Algorithm for SAR Image Registration. IEEE Geosci. Remote Sens. Lett. 2015, 12, 1426–1430. [Google Scholar] [CrossRef] [Scilit]
  52. Soleimani, P.; Capson, D.W.; Li, K.F. Real-Time FPGA-Based Implementation of the AKAZE Algorithm with Nonlinear Scale Space Generation Using Image Partitioning. J. Real-Time Image Process. 2021, 18, 2123–2134. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  53. Toth, C.; Jokow, G. Remote Sensing Platforms and Sensors: A Survey. ISPRS J. Photogramm. Remote Sens. 2016, 115, 22–36. [Google Scholar] [CrossRef] [Scilit]
  54. Colomina, I.; Molina, P. Unmanned Aerial Systems for Photogrammetry and Remote Sensing: A review. ISPRS J. Photogramm. Remote Sens. 2014, 92, 79–97. [Google Scholar] [CrossRef] [Scilit]
  55. Zhuo, X.; Koch, T.; Kurz, F.; Fraundorfer, F.; Reinartz, P. Automatic UAV Image Geo-Registration by Matching UAV Images to Georeferenced Image Data. Remote Sens. 2017, 9, 376. [Google Scholar] [CrossRef] [Scilit]
  56. Yue, K. Multi-Sensor Data Fusion for Autonomous Flight of Unmanned Aerial Vehicles in Complex Flight Environments. Drone Syst. Appl. 2024, 12, 1–12. [Google Scholar] [CrossRef] [Scilit]
  57. Yao, F.; Lan, C.; Wang, L.; Wan, H.; Gao, T.; Wei, Z. GNSS-Denied Geolocalization of UAVs Using Terrain-Weighted Constraint Optimization. Int. J. Appl. Earth Obs. Geoinf. 2024, 135, 104277. [Google Scholar] [CrossRef] [Scilit]
  58. Qiu, X.; Liao, S.; Yang, D.; Li, Y.; Wang, S. High-Precision Visual Geo-Localization of UAV Based on Hierarchical Localization. Expert Syst. Appl. 2025, 267, 126064. [Google Scholar] [CrossRef] [Scilit]
  59. Ye, Q.; Luo, J.; Lin, Y. A Coarse-to-Fine Visual Geo-Localization Method for GNSS-Denied UAV with Oblique-View Imagery. ISPRS J. Photogramm. Remote Sens. 2024, 212, 306–322. [Google Scholar] [CrossRef] [Scilit]
  60. Li, L.; Han, L.; Gao, K.; He, H.; Wang, L.; Li, J. Coarse-to-Fine Matching via Cross Fusion of Satellite Images. Int. J. Appl. Earth Obs. Geoinf. 2023, 125, 103574. [Google Scholar] [CrossRef] [Scilit]
  61. Sarlin, P.-E.; Cadena, C.; Siegwart, R.; Dymczyk, M. From Coarse to Fine: Robust Hierarchical Localization at Large Scale. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018. [Google Scholar]
  62. Cui, Z.; Zhou, P.; Wang, X.; Zhang, Z.; Li, Y.; Li, H.; Zhang, Y. A Novel Geo-Localization Method for UAV and Satellite Images Using Cross-View Consistent Attention. Remote Sens. 2023, 15, 4667. [Google Scholar] [CrossRef] [Scilit]
  63. Qiu, X.; Yang, D.; Liao, S.; Wang, S.; Li, Y. Image Moment Extraction Based Aerial Photo Selection for UAV High-Precision Geolocation without GPS. Meas. J. Int. Meas. Confed. 2024, 226, 114141. [Google Scholar] [CrossRef] [Scilit]
  64. Novikov, D.; Sotirelis, P.; Yilmaz, A. Vehicle Geolocalization from Drone Imagery. ISPRS Ann. Photogramm. Remote Sens. Spat. Inf. Sci. 2024, 10, 171–178. [Google Scholar] [CrossRef] [Scilit]
  65. Li, A.; Cheng, X.; Guan, H.; Feng, T.; Guan, Z. Novel Image Registration Method Based on Local Structure Constraints. IEEE Geosci. Remote Sens. Lett. 2014, 11, 1584–1588. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.