LFODet: Lightweight Few-Shot Object Detection with Meta-Learning in Remote Sensing Images
Abstract
1. Introduction
- We design a lightweight feature fusion network for remote sensing object detection. Instead of simply adding attention modules to the detector, LFODet uses GCA and GSA according to the different roles of multi-level features. GCA mainly enhances channel-level semantic information, while GSA helps preserve fine-grained spatial details. Together with Ghost-based operations, this design improves feature representation while maintaining a lightweight model structure.
- We develop a parameter-constrained three-phase learning procedure for lightweight few-shot detection. Instead of directly applying MAML to the whole detector, LFODet first learns a robust lightweight feature extractor on base classes, then optimizes a meta-learner with the feature extraction network frozen, and finally fine-tunes only the detection head using limited base and novel samples. This strategy constrains the trainable parameter space during few-shot adaptation and improves the stability of novel-class learning.
- We introduce a dual-branch inference strategy to balance novel-class adaptation and base-class retention, rather than simply merging two detection outputs. The base branch preserves base-class knowledge, while the novel branch adapts to few-shot novel classes. A confidence-based gating mechanism selects the more reliable branch using joint classification and objectness confidence, reducing unreliable predictions and alleviating base-class degradation after fine-tuning.
- Extensive comparative and ablation experiments on the DIOR and NWPU VHR-10 datasets demonstrate that LFODet can stably learn the features of novel classes as the number of shots increases, and it outperforms baseline methods under diverse few-shot conditions.
2. Related Work
2.1. Object Detection
2.2. Lightweight Network
2.3. Few-Shot Learning
3. Proposed Method
3.1. Network Architecture
3.1.1. Backbone
3.1.2. Neck
3.1.3. Head
3.2. Three-Phase Learning Procedure
3.2.1. Base Training
3.2.2. Meta-Training
| Algorithm 1 Meta-learning procedure. |
| Require: Model parameters pre-trained in the base training stage; Freeze the feature extraction network; Support set and query set for each task ; Inner-loop learning rate and outer-loop learning rate ; Training epochs: ; Ensure: Updated model parameters ; |
|
3.2.3. Fine-Tuning
3.3. Inference
4. Experiments and Results
4.1. Experimental Platform and Settings
4.2. Dataset and Implementation Details
4.3. Evaluation Metrics
4.4. Results and Comparisons
4.4.1. Numerical Evaluation
4.4.2. Visualization Evaluation
4.5. Ablation Experiments and Discussion
4.5.1. Impact of GSA, GCA, and Fusion
4.5.2. Impact of the Two Parallel Branches and Adding Base Data
4.6. Evaluation of Model Computational Efficiency
4.7. Discussion
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Cheng, G.; Han, J. A survey on object detection in optical remote sensing images. ISPRS J. Photogramm. Remote Sens. 2016, 117, 11–28. [Google Scholar] [CrossRef]
- Xiao, F.; Li, X.; Li, W.; Shi, J.; Zhang, N.; Gao, X. Integrating category-related key regions with a dual-stream network for remote sensing scene classification. J. Vis. Commun. Image Represent. 2024, 100, 104098. [Google Scholar] [CrossRef]
- Fan, Z.; Ma, Y.; Li, Z.; Sun, J. Generalized Few-Shot Object Detection without Forgetting. In Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Virtual, 19–25 June 2021; Computer Vision Foundation (CVF)/IEEE: Piscataway, NJ, USA, 2021; pp. 4525–4534. [Google Scholar] [CrossRef]
- Krizhevsky, A.; Sutskever, I.; Hinton, G.E. ImageNet classification with deep convolutional neural networks. Commun. ACM 2012, 60, 84–90. [Google Scholar] [CrossRef]
- Howard, A.G.; Zhu, M.; Chen, B.; Kalenichenko, D.; Wang, W.; Weyand, T.; Andreetto, M.; Adam, H. MobileNets: Efficient convolutional neural networks for mobile vision applications. arXiv 2017, arXiv:1704.04861. [Google Scholar]
- Sandler, M.; Howard, A.; Zhu, M.; Zhmoginov, A.; Chen, L.C. MobileNetV2: Inverted Residuals and Linear Bottlenecks. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–22 June 2018; Institute of Electrical and Electronics Engineers (IEEE): Piscataway, NJ, USA, 2018; pp. 4510–4520. [Google Scholar] [CrossRef]
- Howard, A.; Sandler, M.; Chu, G.; Chen, L.C.; Chen, B.; Tan, M.; Wang, W.; Zhu, Y.; Pang, R.; Vasudevan, V.; et al. Searching for MobileNetV3. arXiv 2019, arXiv:1905.02244. [Google Scholar] [CrossRef]
- Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You Only Look Once: Unified, Real-Time Object Detection. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NA, USA, 27–30 June 2016; IEEE: Piscataway, NJ, USA, 2016; pp. 779–788. [Google Scholar] [CrossRef]
- Redmon, J.; Farhadi, A. YOLO9000: Better, Faster, Stronger. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; The Institute of Electrical and Electronics Engineers (IEEE): Piscataway, NJ, USA, 2017; pp. 6517–6525. [Google Scholar] [CrossRef]
- Farhadi, A.; Redmon, J. Yolov3: An incremental improvement. In Computer Vision and Pattern Recognition; Springer: Berlin/Heidelberg, Germany, 2018; Volume 1804, pp. 1–6. [Google Scholar]
- Tang, Y.; Han, K.; Guo, J.; Xu, C.; Xu, C.; Wang, Y. GhostNetv2: Enhance cheap operation with long-range attention. Adv. Neural Inf. Process. Syst. 2022, 35, 9969–9982. [Google Scholar] [CrossRef]
- Han, K.; Wang, Y.; Tian, Q.; Guo, J.; Xu, C.; Xu, C. GhostNet: More Features From Cheap Operations. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Virtually, 14–19 June 2020; IEEE (Institute of Electrical and Electronics Engineers): Piscataway, NJ, USA, 2020; pp. 1577–1586. [Google Scholar] [CrossRef]
- Liu, D.; Zhang, J.; Li, T.; Qi, Y.; Wu, Y.; Zhang, Y. A Lightweight Object Detection and Recognition Method Based on Light Global-Local Module for Remote Sensing Images. IEEE Geosci. Remote Sens. Lett. 2023, 20, 6007105. [Google Scholar] [CrossRef]
- Zhu, S.; Miao, M. SCNet: A Lightweight and Efficient Object Detection Network for Remote Sensing. IEEE Geosci. Remote Sens. Lett. 2024, 21, 6001605. [Google Scholar] [CrossRef]
- Fang, X.; Yang, X.; Yang, X.; Li, J. Progressive Multi-Level Feature Fusion network with Global–Local Feature Enhancement for Infrared small target detection. Infrared Phys. Technol. 2025, 150, 105964. [Google Scholar] [CrossRef]
- Zheng, X.; Bi, J.; Li, K.; Zhang, G.; Jiang, P. SMN-YOLO: Lightweight YOLOv8-Based Model for Small Object Detection in Remote Sensing Images. IEEE Geosci. Remote Sens. Lett. 2025, 22, 8001305. [Google Scholar] [CrossRef]
- Zhang, P.; Bai, Y.; Wang, D.; Bai, B.; Li, Y. A Meta-Learning Framework for Few-Shot Classification of Remote Sensing Scene. In Proceedings of the ICASSP 2021—2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Virtually, 6–11 June 2021; Institute of Electrical and Electronics Engineers (IEEE): Piscataway, NJ, USA, 2021; pp. 4590–4594. [Google Scholar] [CrossRef]
- Cheng, M.; Wang, H.; Long, Y. Meta-Learning-Based Incremental Few-Shot Object Detection. IEEE Trans. Circuits Syst. Video Technol. 2022, 32, 2158–2169. [Google Scholar] [CrossRef]
- Peng, P.; Wang, J. How to fine-tune deep neural networks in few-shot learning? arXiv 2020, arXiv:2012.00204. [Google Scholar] [CrossRef]
- Girshick, R.; Donahue, J.; Darrell, T.; Malik, J. Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation. In Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA, 23–28 June 2014; Institute of Electrical and Electronics Engineers (IEEE): Piscataway, NJ, USA, 2014; pp. 580–587. [Google Scholar] [CrossRef]
- Girshick, R. Fast R-CNN. In Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile, 7–13 December 2015; The Institute of Electrical and Electronics Engineers (IEEE): Piscataway, NJ, USA, 2015; pp. 1440–1448. [Google Scholar] [CrossRef]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 39, 1137–1149. [Google Scholar] [CrossRef] [PubMed]
- Shi, Y.; Guo, J.; Wang, X.; Wang, Y. TDENet: Three-branch distillation enhancement network for foggy scene object detection. J. Vis. Commun. Image Represent. 2025, 111, 104534. [Google Scholar] [CrossRef]
- Fu, K.; Zhang, T.; Zhang, Y.; Yan, M.; Chang, Z.; Zhang, Z.; Sun, X. Meta-SSD: Towards Fast Adaptation for Few-Shot Object Detection With Meta-Learning. IEEE Access 2019, 7, 77597–77606. [Google Scholar] [CrossRef]
- Tang, H.; Li, Z.; Zhang, D.; He, S.; Tang, J. Divide-and-Conquer: Confluent Triple-Flow Network for RGB-T Salient Object Detection. IEEE Trans. Pattern Anal. Mach. Intell. 2025, 47, 1958–1974. [Google Scholar] [CrossRef] [PubMed]
- Zhang, X.; Zhou, X.; Lin, M.; Sun, J. ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–22 June 2018; Institute of Electrical and Electronics Engineers (IEEE): Piscataway, NJ, USA, 2018; pp. 6848–6856. [Google Scholar] [CrossRef]
- Ma, N.; Zhang, X.; Zheng, H.T.; Sun, J. ShuffleNet v2: Practical guidelines for efficient cnn architecture design. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; Springer: Cham, Switzerland, 2018; pp. 122–131. [Google Scholar]
- Liang, B.; Luo, H. MEANet: An effective and lightweight solution for salient object detection in optical remote sensing images. Expert Syst. Appl. 2024, 238, 121778. [Google Scholar] [CrossRef]
- Lv, M.; Liu, Y.; Zha, Z.; Zheng, X.; Wang, H.; Wen, Y.; Guo, Z. Si-CA MobileNet: A lightweight and efficient convolutional neural network for distracted driver detection. Neurocomputing 2025, 654, 131281. [Google Scholar] [CrossRef]
- Sun, B.; Li, B.; Cai, S.; Yuan, Y.; Zhang, C. FSCE: Few-Shot Object Detection via Contrastive Proposal Encoding. In Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Virtual, 19–25 June 2021; IEEE (Institute of Electrical and Electronics Engineers): Piscataway, NJ, USA, 2021; pp. 7348–7358. [Google Scholar] [CrossRef]
- Aganian, D.; Eisenbach, M.; Wagner, J.; Seichter, D.; Gross, H.M. Revisiting loss functions for person re-identification. In Proceedings of the Artificial Neural Networks and Machine Learning–ICANN 2021: 30th International Conference on Artificial Neural Networks, Bratislava, Slovakia, 14–17 September 2021; Proceedings, Part V 30; Springer: Berlin/Heidelberg, Germany, 2021; pp. 30–42. [Google Scholar]
- Xin, Z.; Wu, T.; Zou, Y.; Chen, S.; Fu, D.; You, X. Few-Shot Object Detection via Spatial-Channel State Space Model. arXiv 2025, arXiv:2507.15308. [Google Scholar] [CrossRef]
- Zhang, X.; Chen, Z.; Zhang, J.; Liu, T.; Tao, D. Learning General and Specific Embedding with Transformer for Few-Shot Object Detection. Int. J. Comput. Vis. 2024, 133, 968–984. [Google Scholar] [CrossRef]
- Kang, B.; Liu, Z.; Wang, X.; Yu, F.; Feng, J.; Darrell, T. Few-Shot Object Detection via Feature Reweighting. In Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; IEEE: Piscataway, NJ, USA, 2019; pp. 8419–8428. [Google Scholar] [CrossRef]
- Li, X.; Deng, J.; Fang, Y. Few-Shot Object Detection on Remote Sensing Images. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5601614. [Google Scholar] [CrossRef]
- Köhler, M.; Eisenbach, M.; Gross, H.M. Few-Shot Object Detection: A Comprehensive Survey. IEEE Trans. Neural Netw. Learn. Syst. 2024, 35, 11958–11978. [Google Scholar] [CrossRef] [PubMed]
- Hu, J.; Shen, L.; Albanie, S.; Sun, G.; Wu, E. Squeeze-and-Excitation Networks. IEEE Trans. Pattern Anal. Mach. Intell. 2020, 42, 2011–2023. [Google Scholar] [CrossRef] [PubMed]
- Zheng, Z.; Wang, P.; Ren, D.; Liu, W.; Ye, R.; Hu, Q.; Zuo, W. Enhancing Geometric Factors in Model Learning and Inference for Object Detection and Instance Segmentation. IEEE Trans. Cybern. 2022, 52, 8574–8586. [Google Scholar] [CrossRef] [PubMed]
- Lin, T.Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. Focal Loss for Dense Object Detection. IEEE Trans. Pattern Anal. Mach. Intell. 2020, 42, 318–327. [Google Scholar] [CrossRef] [PubMed]
- Chen, W.Y.; Liu, Y.C.; Kira, Z.; Wang, Y.C.F.; Huang, J.B. A Closer Look at Few-shot Classification. arXiv 2019, arXiv:1904.04232. [Google Scholar] [CrossRef]
- Xia, R.; Li, G.; Huang, Z.; Meng, H.; Pang, Y. Bi-path Combination YOLO for Real-time Few-shot Object Detection. Pattern Recognit. Lett. 2023, 165, 91–97. [Google Scholar] [CrossRef]
- Li, K.; Wan, G.; Cheng, G.; Meng, L.; Han, J. Object detection in optical remote sensing images: A survey and a new benchmark. ISPRS J. Photogramm. Remote Sens. 2020, 159, 296–307. [Google Scholar] [CrossRef]
- Cheng, G.; Han, J.; Zhou, P.; Guo, L. Multi-class geospatial object detection and geographic image classification based on collection of part detectors. ISPRS J. Photogramm. Remote Sens. 2014, 98, 119–132. [Google Scholar] [CrossRef]
- Zhang, H.; Cissé, M.; Dauphin, Y.N.; Lopez-Paz, D. mixup: Beyond Empirical Risk Minimization. arXiv 2017, arXiv:1710.09412. [Google Scholar] [CrossRef]
- Nichol, A.; Achiam, J.; Schulman, J. On First-Order Meta-Learning Algorithms. arXiv 2018, arXiv:1803.02999. [Google Scholar] [CrossRef]
- Wang, X.; Huang, T.E.; Darrell, T.; Gonzalez, J.E.; Yu, F. Frustratingly Simple Few-Shot Object Detection. arXiv 2020, arXiv:2003.06957. [Google Scholar] [CrossRef]
- Wu, W.; Jiang, C.; Yang, L.; Wang, W.; Chen, Q.; Zhang, J.; Yang, H.; Chen, Z. Arbitrary Oriented Few-Shot Object Detection in Remote Sensing Images. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2024, 17, 17930–17944. [Google Scholar] [CrossRef]
- Sultan, N.; Hayat, M.; Prom-on, S. DSCH-Net: Diffusion-State-Contextual Hybrid Network for Physics-Inspired and Direction-Aware Dehazing of Remote Sensing Imagery. IEEE Trans. Geosci. Remote Sens. 2026, 64, 4106218. [Google Scholar] [CrossRef]








| Setting | Base Training | Meta-Training | Fine-Tuning |
|---|---|---|---|
| Scheduler | Cosine decay | – | – |
| Optimizer | Adam | Adam | Adam |
| Learning rate | 1.5 × 10−4 initial/ 1 × 10−6 | 1 × 10−3 | 8 × 10−3 |
| Epoch | 120 | 3 | 400 |
| Batch size | 10 | 4 | 16 |
| Method | Backbone | Shots | DIOR | NWPU VHR-10 | |||||
|---|---|---|---|---|---|---|---|---|---|
| nAP | bAP | mAP | nAP | bAP | mAP | ||||
| MAML [45] | MobileNetv3 | 3 | |||||||
| 5 | |||||||||
| 10 | |||||||||
| Fine-Tuning [18] | MobileNetv3 | 3 | |||||||
| 5 | |||||||||
| 10 | |||||||||
| TFA [46] | MobileNetv3 | 3 | |||||||
| 5 | |||||||||
| 10 | |||||||||
| BC-YOLO * [41] | MobileNetv3 | 3 | |||||||
| 5 | |||||||||
| 10 | |||||||||
| BC-YOLO [41] | Darknet-53 | 3 | |||||||
| 5 | |||||||||
| 10 | |||||||||
| AOFS [47] | CSPDarknet-53 | 3 | |||||||
| 5 | |||||||||
| 10 | |||||||||
| LFODet | MobileNetv3 | 3 | |||||||
| 5 | |||||||||
| 10 | |||||||||
| Method | Shot | Airplane | Baseball Field | Tennis Court | Train Station | Wind Mill | nAP |
|---|---|---|---|---|---|---|---|
| MAML [45] | 3 | ||||||
| 5 | |||||||
| 10 | |||||||
| Fine-Tuning [18] | 3 | ||||||
| 5 | |||||||
| 10 | |||||||
| TFA [46] | 3 | ||||||
| 5 | |||||||
| 10 | |||||||
| BC-YOLO * [41] | 3 | ||||||
| 5 | |||||||
| 10 | |||||||
| BC-YOLO [41] | 3 | ||||||
| 5 | |||||||
| 10 | |||||||
| AOFS [47] | 3 | ||||||
| 5 | |||||||
| 10 | |||||||
| LFODet | 3 | ||||||
| 5 | |||||||
| 10 |
| Methods | Base Training | GSA | GCA | Fusion | mAP (DIOR) | mAP (NWPU VHR-10) | FLOPs (G) | Params (M) |
|---|---|---|---|---|---|---|---|---|
| Selected Module(s) | ✓ | ✓ | ✓ | ✓ | 67.5 | 89.6 | 9.654 | 10.06 |
| ✓ | ✓ | ✓ | 60.5 (−7.0) | 80.4 (−9.2) | 8.661 (−0.993) | 8.85 (−1.21) | ||
| ✓ | ✓ | ✓ | 59.6 (−7.9) | 84.1 (−5.5) | 9.144 (−0.510) | 7.44 (−2.62) | ||
| ✓ | ✓ | ✓ | 64.4 (−3.1) | 82.4 (−7.2) | 7.708 (−1.946) | 9.37 (−0.69) | ||
| ✓ | 57.9 (−9.6) | 75.3 (−14.3) | 4.279 (−5.375) | 4.24 (−5.82) |
| Methods | Base Training | GSA | GCA | Fusion | FPS | Detection Time (ms/img) |
|---|---|---|---|---|---|---|
| Selected Module(s) | ✓ | ✓ | ✓ | ✓ | 9.66 | 103.55 |
| ✓ | ✓ | ✓ | 7.27 | 137.55 | ||
| ✓ | ✓ | ✓ | 7.35 | 135.97 | ||
| ✓ | ✓ | ✓ | 5.64 | 177.30 | ||
| ✓ | 4.83 | 207.02 |
| Shots | Parallel Branches | Adding Base Data | DIOR | NWPU VHR-10 | |||||
|---|---|---|---|---|---|---|---|---|---|
| nAP | bAP | mAP | nAP | bAP | mAP | ||||
| 3 | ✓ | ✓ | |||||||
| ✓ | |||||||||
| ✓ | |||||||||
| 5 | ✓ | ✓ | |||||||
| ✓ | |||||||||
| ✓ | |||||||||
| 10 | ✓ | ✓ | |||||||
| ✓ | |||||||||
| ✓ | |||||||||
| Backbone | Vanilla | DSC | Ghost | mAP (DIOR) | mAP (NWPU VHR-10) | FLOPs (G) | Params (M) |
|---|---|---|---|---|---|---|---|
| MobileNetv3 | ✓ | 62.5 | 75.2 | 14.555 | 14.87 | ||
| ✓ | 63.2 | 92.2 | 14.735 | 14.98 | |||
| ✓ | 67.5 | 89.6 | 9.654 | 10.06 | |||
| MobileNetv2 | ✓ | 63.1 | 81.8 | 12.153 | 12.71 | ||
| ShuffleNetv2 | ✓ | 62.4 | 86.5 | 22.377 | 13.80 | ||
| GhostNetv2 | ✓ | 53.6 | 37.8 | 11.293 | 13.30 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Wu, H.; Fang, X.; Xiong, H.; Yang, X. LFODet: Lightweight Few-Shot Object Detection with Meta-Learning in Remote Sensing Images. Sensors 2026, 26, 4371. https://doi.org/10.3390/s26144371
Wu H, Fang X, Xiong H, Yang X. LFODet: Lightweight Few-Shot Object Detection with Meta-Learning in Remote Sensing Images. Sensors. 2026; 26(14):4371. https://doi.org/10.3390/s26144371
Chicago/Turabian StyleWu, Haoran, Xuan Fang, Haonan Xiong, and Xiaomei Yang. 2026. "LFODet: Lightweight Few-Shot Object Detection with Meta-Learning in Remote Sensing Images" Sensors 26, no. 14: 4371. https://doi.org/10.3390/s26144371
APA StyleWu, H., Fang, X., Xiong, H., & Yang, X. (2026). LFODet: Lightweight Few-Shot Object Detection with Meta-Learning in Remote Sensing Images. Sensors, 26(14), 4371. https://doi.org/10.3390/s26144371

