Bi-Level Meta-Learning for Reliable Remote Sensing Image Registration
Highlights
- A saliency-aware bi-level optimization framework is proposed to address the challenges of extreme radiometric and temporal variations in heterogeneous remote sensing image matching.
- A novel Saliency Judgment module is developed to intelligently prioritize geometrically stable landmarks while suppressing spurious matches in weakly informative regions.
- The bi-level meta-learning strategy establishes a new paradigm for UAV localization by enabling effective model training under weakly labeled data regimes, reducing the reliance on large-scale expert annotations.
- The framework significantly enhances registration accuracy and reduces ghosting artifacts, providing a robust “quality-over-quantity” correspondence selection strategy for autonomous aerial localization.
Abstract
1. Introduction
- We propose a semi-supervised learning framework that decouples automated correspondence generation from saliency-based selection, substantially reducing the dependency on costly large-scale dense manual annotations while enabling effective training on high-heterogeneity remote sensing imagery. The parameterized geometric synthesis pipeline generates diverse training pairs with accurate ground-truth correspondences, while the saliency-based selection mechanism filters unreliable samples to improve training efficiency and model robustness.
- We design a Saliency Judgment Network trained through bi-level meta-learning that adaptively identifies and prioritizes reliable correspondences for matcher optimization. The bi-level formulation treats correspondence selection as an inner optimization and matcher performance as an outer objective, enabling the network to learn a principled, data-driven criterion for sample importance without hand-crafted heuristics.
- We introduce RS-Hetero-50K, a large-scale remote sensing dataset encompassing approximately 50,000 high-heterogeneity image pairs across diverse terrain types including urban, suburban, and gobi scenes with extreme cross-domain variations. Extensive experimental evaluations demonstrate that the complete framework reaches 95.4% matching precision and that the plug-and-play saliency module consistently improves LoFTR, SuperGlue, and LightGlue backbones.
2. Related Work
2.1. Deep Learning-Based Image Matching in Natural Scenes
2.2. Image Matching in Remote Sensing and UAV Navigation
3. Materials and Methods
3.1. Dataset Preparation
3.1.1. Initial Correspondence Generation
3.1.2. Preferred Correspondence Selection
Automated Pre-Filtering
Expert-in-the-Loop Refinement
3.2. Network Architecture
3.2.1. CNN-Attention Feature Matching Module
3.2.2. Saliency Judgment and Bi-Level Optimization
| Algorithm 1 Bi-level meta-learning for saliency-aware matching |
|
3.3. Inference Pipeline
4. Experiments and Results
4.1. Datasets
4.2. Implementation Details
4.3. Quantitative Evaluation on RS-Hetero-50K
4.4. Ablation Study on Selection Strategy
4.5. Terrain-Specific Breakdown Analysis
4.6. Robustness and Cross-Domain Generalization
4.7. Computational Efficiency Analysis
5. Discussion
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Scaramuzza, D.; Achtelik, M.C.; Doitsidis, L.; Fraundorfer, F.; Kosmatopoulos, E.; Martinelli, A.; Achtelik, M.W.; Chli, M.; Chatzichristofis, S.; Kneip, L.; et al. Vision-controlled micro flying robots: From system design to autonomous navigation and mapping in GPS-denied environments. IEEE Robot. Autom. Mag. 2014, 21, 26–40. [Google Scholar] [CrossRef]
- Yang, S.; Scherer, S.A.; Zell, A. An onboard monocular vision system for autonomous takeoff, hovering and landing of a micro aerial vehicle. J. Intell. Robot. Syst. 2013, 69, 499–515. [Google Scholar]
- Piasco, N.; Marzat, J.; Sanfourche, M. Vision-based aerial localization using geo-referenced images. In Proceedings of the 2015 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Hamburg, Germany, 28 September–2 October 2015; IEEE: New York, NY, USA, 2015; pp. 3050–3055. [Google Scholar]
- Chen, H.; Luo, Z.; Zhang, J.; Zhou, L.; Bai, X.; Hu, Z.; Tai, C.L.; Quan, L. Learning to match features with seeded graph matching network. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 11–17 October 2021; pp. 6301–6310. [Google Scholar]
- Piasco, N.; Sidibé, D.; Demonceaux, C.; Gouet-Brunet, V. A survey on visual-based localization: On the benefit of heterogeneous data. Pattern Recognit. 2018, 74, 90–109. [Google Scholar] [CrossRef]
- Toft, C.; Stenborg, E.; Hammarstrand, L.; Brynte, L.; Pollefeys, M.; Sattler, T.; Kahl, F. Semantic match consistency for long-term visual localization. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; Springer: Cham, Switzerland, 2018; pp. 383–399. [Google Scholar]
- Li, L.; Han, L.; Ding, M.; Cao, H. Multimodal image fusion framework for end-to-end remote sensing image registration. IEEE Trans. Geosci. Remote Sens. 2023, 61, 5607214. [Google Scholar] [CrossRef]
- Jiang, X.; Ma, J.; Xiao, G.; Shao, Z.; Guo, X. A review of multimodal image matching: Methods and applications. Inf. Fusion 2021, 73, 22–71. [Google Scholar] [CrossRef]
- Li, Z.; Snavely, N. Megadepth: Learning single-view depth prediction from internet photos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018; IEEE: New York, NY, USA, 2018; pp. 2041–2050. [Google Scholar]
- Dai, A.; Chang, A.X.; Savva, M.; Halber, M.; Funkhouser, T.; Niessner, M. ScanNet: Richly-annotated 3D reconstructions of indoor scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; IEEE: New York, NY, USA, 2017; pp. 5828–5839. [Google Scholar]
- Lowe, D.G. Distinctive image features from scale-invariant keypoints. Int. J. Comput. Vis. 2004, 60, 91–110. [Google Scholar] [CrossRef]
- Bay, H.; Ess, A.; Tuytelaars, T.; Van Gool, L. SURF: Speeded up robust features. Comput. Vis. Image Underst. 2008, 110, 346–359. [Google Scholar] [CrossRef]
- Rublee, E.; Rabaud, V.; Konolige, K.; Bradski, G. ORB: An efficient alternative to SIFT or SURF. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Barcelona, Spain, 6–13 November 2011; IEEE: New York, NY, USA, 2011; pp. 2564–2571. [Google Scholar]
- DeTone, D.; Malisiewicz, T.; Rabinovich, A. SuperPoint: Self-supervised interest point detection and description. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Salt Lake City, UT, USA, 18–22 June 2018; IEEE: New York, NY, USA, 2018; pp. 224–236. [Google Scholar]
- Revaud, J.; de Souza, C.R.; Humenberger, M.; Weinzaepfel, P. R2D2: Reliable and repeatable detector and descriptor. In Proceedings of the 33rd International Conference on Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada, 8–14 December 2019; Curran Associates Inc.: Red Hook, NY, USA, 2019; pp. 12405–12415. [Google Scholar]
- Dusmanu, M.; Rocco, I.; Pajdla, T.; Pollefeys, M.; Sivic, J.; Torii, A.; Sattler, T. D2-Net: A trainable CNN for joint description and detection of local features. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019; IEEE: New York, NY, USA, 2019; pp. 8092–8101. [Google Scholar]
- Noh, H.; Araujo, A.; Sim, J.; Weyand, T.; Han, B. Large-scale image retrieval with attentive deep local features. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017; IEEE: New York, NY, USA, 2017; pp. 3456–3465. [Google Scholar]
- Sarlin, P.E.; DeTone, D.; Malisiewicz, T.; Rabinovich, A. SuperGlue: Learning feature matching with graph neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; IEEE: New York, NY, USA, 2020; pp. 4938–4947. [Google Scholar]
- Lindenberger, P.; Sarlin, P.E.; Pollefeys, M. LightGlue: Local feature matching at light speed. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 2–6 October 2023; IEEE: New York, NY, USA, 2023; pp. 17627–17638. [Google Scholar]
- Potje, G.; Cadar, F.; Araujo, A.; Martins, R.; Nascimento, E.R. XFeat: Accelerated features for lightweight image matching. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 17–21 June 2024; IEEE: New York, NY, USA, 2024; pp. 2682–2691. [Google Scholar]
- Sun, J.; Shen, Z.; Wang, Y.; Bao, H.; Zhou, X. LoFTR: Detector-free local feature matching with transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021; IEEE: New York, NY, USA, 2021; pp. 8922–8931. [Google Scholar]
- Wang, Q.; Zhang, J.; Yang, K.; Peng, K.; Stiefelhagen, R. MatchFormer: Interleaving attention in transformers for feature matching. In Proceedings of the Asian Conference on Computer Vision (ACCV), Macau, China, 4–8 December 2022; Springer: Berlin/Heidelberg, Germany, 2022; pp. 2746–2762. [Google Scholar]
- Chen, H.; Luo, Z.; Zhou, L.; Tian, Y.; Zhen, M.; Fang, T.; McKinnon, D.; Tsin, Y.; Quan, L. ASpanFormer: Detector-free image matching with adaptive span transformer. In Proceedings of the European Conference on Computer Vision (ECCV), Tel Aviv, Israel, 23–27 October 2022; Springer: Berlin/Heidelberg, Germany, 2022; pp. 20–36. [Google Scholar]
- Wang, Y.; He, X.; Peng, S.; Tan, D.; Zhou, X. Efficient LoFTR: Semi-dense local feature matching with sparse-like speed. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 17–21 June 2024; IEEE: New York, NY, USA, 2024; pp. 21666–21675. [Google Scholar]
- Ganin, Y.; Lempitsky, V. Unsupervised domain adaptation by backpropagation. In Proceedings of the 32nd International Conference on Machine Learning (ICML), Lille, France, 6–11 July 2015; JMLR: Norfolk, MA, USA, 2015; pp. 1180–1189. [Google Scholar]
- Sun, B.; Saenko, K. Deep CORAL: Correlation alignment for deep domain adaptation. In Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands, 11–14 October 2016; Springer: Cham, Switzerland, 2016; pp. 443–450. [Google Scholar]
- Settles, B. Active Learning Literature Survey; Technical Report 1648; University of Wisconsin-Madison: Madison, WI, USA, 2009. [Google Scholar]
- Sener, O.; Savarese, S. Active learning for convolutional neural networks: A core-set approach. In Proceedings of the International Conference on Learning Representations (ICLR), Vancouver, BC, Canada, 30 April–3 May 2018. [Google Scholar]
- Gal, Y.; Islam, R.; Ghahramani, Z. Deep Bayesian active learning with image data. In Proceedings of the 34th International Conference on Machine Learning (ICML), Sydney, Australia, 6–11 August 2017; JMLR: Norfolk, MA, USA, 2017; pp. 1183–1192. [Google Scholar]
- Bookstein, F.L. Principal warps: Thin-plate splines and the decomposition of deformations. IEEE Trans. Pattern Anal. Mach. Intell. 1989, 11, 567–585. [Google Scholar] [CrossRef]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; IEEE: New York, NY, USA, 2016; pp. 770–778. [Google Scholar]
- Lin, T.Y.; Dollár, P.; Girshick, R.; He, K.; Hariharan, B.; Belongie, S. Feature pyramid networks for object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; IEEE: New York, NY, USA, 2017; pp. 936–944. [Google Scholar]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. In Proceedings of Advances in Neural Information Processing Systems (NeurIPS), Long Beach, CA, USA, 4–9 December 2017; Curran Associates Inc.: Red Hook, NY, USA, 2017; pp. 5998–6008. [Google Scholar]
- Lin, H.; Cheng, X.; Wu, X.; Shen, D. CAT: Cross Attention in Vision Transformer. In Proceedings of the 2022 IEEE International Conference on Multimedia and Expo (ICME), Taipei, Taiwan, 18–22 July 2022; pp. 1–6. [Google Scholar]
- Jang, E.; Gu, S.; Poole, B. Categorical reparameterization with Gumbel-Softmax. In Proceedings of the International Conference on Learning Representations (ICLR), Toulon, France, 24–26 April 2017. [Google Scholar]
- Maddison, C.J.; Mnih, A.; Teh, Y.W. The concrete distribution: A continuous relaxation of discrete random variables. In Proceedings of the International Conference on Learning Representations (ICLR), Toulon, France, 24–26 April 2017. [Google Scholar]
- Fischler, M.A.; Bolles, R.C. Random sample consensus: A paradigm for model fitting with applications to image analysis and automated cartography. Commun. ACM 1981, 24, 381–395. [Google Scholar] [CrossRef]
- Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. PyTorch: An imperative style, high-performance deep learning library. In Proceedings of Advances in Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada, 8–14 December 2019; Curran Associates Inc.: Red Hook, NY, USA, 2019; pp. 8024–8035. [Google Scholar]
- Deng, J.; Dong, W.; Socher, R.; Li, L.J.; Li, K.; Fei-Fei, L. ImageNet: A large-scale hierarchical image database. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Miami, FL, USA, 20–25 June 2009; IEEE: New York, NY, USA, 2009; pp. 248–255. [Google Scholar]
- Kingma, D.P.; Ba, J. Adam: A method for stochastic optimization. In Proceedings of the International Conference on Learning Representations (ICLR), San Diego, CA, USA, 7–9 May 2015. [Google Scholar]








| Symbol | Meaning |
|---|---|
| S, A | Satellite reference image and UAV/aerial orthomosaic. |
| , | Source and target images fed into the matching network. |
| , , | Similarity transform, TPS residual warp, and crop-normalization transform. |
| Composite synthetic mapping used for coarse correspondence generation. | |
| , | Automatically generated coarse dataset and expert-validated meta-dataset. |
| , | Feature matching network and Saliency Judgment Network. |
| , | Putative correspondence set and its l-th correspondence. |
| , , | Dual-softmax confidence, saliency score, and differentiable selection weight. |
| Subset | Pairs | Usage | Geographic Overlap with Test |
|---|---|---|---|
| RS-Hetero-50K training split, including the 500-pair meta subset | 40,000 | Matcher training on ; source of | No |
| RS-Hetero-50K meta subset | 500 | SJN bi-level meta-training only | No |
| RS-Hetero-50K validation split | 5000 | Hyperparameter validation | No |
| RS-Hetero-50K test split | 5000 | Final in-domain evaluation | – |
| FuJian-Mountain | 500 | Independent zero-shot evaluation | No |
| Method | Paradigm | Training Setting | MP (%) | AUC@5 px | AUC@10 px | AUC@20 px |
|---|---|---|---|---|---|---|
| DELF [17] | Local feature | Official pretrained | 69.1 | 14.4 | 32.1 | 57.9 |
| D2-Net [16] | CNN local feature | Fine-tuned on train split | 80.4 | 29.8 | 48.5 | 71.8 |
| XFeat [20] | Lightweight local feature | Fine-tuned on train split | 85.4 | 31.6 | 51.8 | 73.7 |
| LoFTR [21] | Detector-free semi-dense | Fine-tuned on train split | 89.8 | 35.1 | 54.3 | 75.3 |
| Efficient LoFTR [24] | Efficient semi-dense matching | Fine-tuned on train split | 90.2 | 36.4 | 55.6 | 75.9 |
| SP + SG [18] | Sparse learned matching | Fine-tuned on train split | 90.1 | 36.7 | 55.4 | 75.6 |
| LightGlue [19] | Lightweight sparse matching | Fine-tuned on train split | 91.2 | 36.9 | 55.9 | 76.2 |
| Ours (Full) | Bi-level saliency-aware matching | Train split + meta subset | 95.4 | 40.2 | 61.5 | 81.3 |
| Backbone | Saliency Module | MP (%) | AUC@5 px | AUC@10 px |
|---|---|---|---|---|
| LoFTR [21] | None | 89.8 | 35.1 | 54.3 |
| LoFTR | + Proposed | 94.2 | 38.6 | 59.1 |
| SP + SG [18] | None | 90.1 | 36.7 | 55.4 |
| SP + SG | + Proposed | 91.8 | 37.1 | 56.8 |
| Method | Matches | MP (%) | AUC@5 px | AUC@10 px | AUC@20 px |
|---|---|---|---|---|---|
| LoFTR (Baseline) [21] | 2231 | 89.8 | 35.1 | 54.3 | 75.3 |
| LoFTR + Random-Prune | 1562 | 88.4 | 33.2 | 52.1 | 73.8 |
| LoFTR + Proposed | 1577 | 94.2 | 38.6 | 59.1 | 79.6 |
| SP + SG (Baseline) [18] | 1794 | 90.1 | 36.7 | 55.4 | 75.6 |
| SP + SG + Random-Prune | 1256 | 89.5 | 35.1 | 53.8 | 74.2 |
| SP + SG + Proposed | 1099 | 91.8 | 37.1 | 56.8 | 77.9 |
| LightGlue (Baseline) [19] | 1650 | 91.2 | 36.9 | 55.9 | 76.2 |
| LightGlue + Random-Prune | 1155 | 90.6 | 35.4 | 54.3 | 74.9 |
| LightGlue + Proposed | 1210 | 93.5 | 37.8 | 58.3 | 78.5 |
| Terrain | Method | MP (%) | AUC@5 px | AUC@20 px |
|---|---|---|---|---|
| Urban | Baseline | 91.3 | 38.2 | 78.4 |
| + Proposed | 95.7 | 41.3 | 82.1 | |
| Gobi | Baseline | 86.4 | 31.5 | 71.2 |
| + Proposed | 92.1 | 36.8 | 77.5 | |
| Suburban | Baseline | 88.7 | 33.8 | 74.3 |
| + Proposed | 93.4 | 37.1 | 78.9 |
| Method | MP (%) | AUC@5 px | AUC@10 px | AUC@20 px |
|---|---|---|---|---|
| D2-Net [16] | 62.4 | 9.8 | 15.3 | 21.8 |
| LoFTR (Baseline) [21] | 71.5 | 15.1 | 22.4 | 25.3 |
| LoFTR + Proposed | 84.2 | 28.6 | 42.3 | 59.6 |
| SP + SG (Baseline) [18] | 73.8 | 16.7 | 23.1 | 25.6 |
| SP + SG + Proposed | 82.1 | 27.1 | 40.8 | 57.9 |
| Component | Parameter Count |
|---|---|
| ResNet-34 backbone | 21.285 M |
| FPN | 2.607 M |
| CNN-attention interaction module | 2.109 M |
| Saliency Judgment Network (SJN) | 0.428 M |
| Total | 26.429 M |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Shi, L.; Wang, R.; Zhu, X.; An, C.; Zhao, K.; Shu, J.; Yang, D.; Meng, D. Bi-Level Meta-Learning for Reliable Remote Sensing Image Registration. Remote Sens. 2026, 18, 2007. https://doi.org/10.3390/rs18122007
Shi L, Wang R, Zhu X, An C, Zhao K, Shu J, Yang D, Meng D. Bi-Level Meta-Learning for Reliable Remote Sensing Image Registration. Remote Sensing. 2026; 18(12):2007. https://doi.org/10.3390/rs18122007
Chicago/Turabian StyleShi, Lin, Renzhen Wang, Xiaofeng Zhu, Cong An, Kai Zhao, Jun Shu, Dongfang Yang, and Deyu Meng. 2026. "Bi-Level Meta-Learning for Reliable Remote Sensing Image Registration" Remote Sensing 18, no. 12: 2007. https://doi.org/10.3390/rs18122007
APA StyleShi, L., Wang, R., Zhu, X., An, C., Zhao, K., Shu, J., Yang, D., & Meng, D. (2026). Bi-Level Meta-Learning for Reliable Remote Sensing Image Registration. Remote Sensing, 18(12), 2007. https://doi.org/10.3390/rs18122007
