SFSMamba-DETR: Selective Feature Scanning with State Space Models and Dual-Scale Window Attention for Remote Sensing Object Detection
Highlights
- SFSMamba-DETR integrates Selective Feature Scanning, Dual-ScaleWindow Attention, and hierarchical cross-scale feature aggregation to address scale variation, complex backgrounds, and densely distributed small objects in remote sensing imagery.
- On MAR20, UCAS-AOD, and the Jilin-1 Satellite Aircraft Detection Dataset, the proposed method achieved mAP@0.5 scores of 85.7%, 97.6%, and 82.3%, respectively, while maintaining an inference speed of 48.6 FPS on a single NVIDIA A100 GPU.
- The results indicate that selective state space scanning can efficiently model long-range, cross-scale dependencies without relying on additional dense global self-attention.
- The proposed framework provides an effective approach to horizontal-bounding-box detection of small, densely distributed objects in high-resolution remote sensing images with cluttered backgrounds.
Abstract
1. Introduction
- Selective Feature Scanning (SFS) Module: We design a tri-input Mamba-based module that takes main, guide, and auxiliary feature maps from different scales. Through a Vision State Space (VSS) module with 2D Selective Scan (SS2D), SFS enables adaptive long-range dependency modeling with linear complexity, while using cross-scale guidance to enhance feature discrimination.
- Dual-Scale Window Attention (DSWA) Mechanism: We introduce a dual-branch attention module that partitions features into small and large windows simultaneously. A Dual Local–Global Perception Attention (DLGPA) mechanism with multi-kernel convolution bridges the two scales, capturing both fine-grained details critical for small aircraft detection and broader contextual patterns for suppressing background clutter.
- Cross-scale Feature Aggregation Module (CFAM): We design a hierarchical feature fusion architecture within the hybrid encoder that orchestrates SFS modules and learnable fusion blocks across Feature Pyramid Network (FPN) levels, enabling systematic information exchange between scales.
2. Related Work
2.1. CNN-Based Object Detection
2.2. Transformer-Based Detection
2.3. Multi-Scale Feature Fusion in Detection
2.4. State Space Models for Vision
2.5. Object Detection in Remote Sensing
3. Proposed Method
3.1. Overall Architecture

3.2. Selective Feature Scanning Module
3.2.1. Tri-Input Design Rationale
3.2.2. Vision State Space Module
3.2.3. 2D Selective Scan Mechanism
3.3. Dual-Scale Window Attention
Dual Local–Global Perception Attention
3.4. Cross-Scale Feature Aggregation Module
3.5. Loss Function and Training
| Algorithm 1 Training procedure of SFSMamba-DETR. |
| Input: Training dataset , pretrained EfficientNet weights Output: Trained model parameters 1 Initialize backbone with ; randomly initialize encoder, CFAM, DSWA, decoder 2 for epoch to do 3 for each mini-batch do 4 (Multi-scale feature extraction) 5 (Intra-scale feature interaction) 6 (Selective Feature Scanning) 7 8 9 10 (Element-wise aggregation) 11 (Dual-scale attention) 12 (Encoder output) 13 (Query initialization) 14 (Prediction) 15 (Bipartite matching) 16 Compute via Equation (23) 17 Update by 18 end for 19 Apply learning rate scheduling (cosine annealing) 20 end for 21 return |
4. Experiments
4.1. Datasets
| Dataset | Source | #Images | #Categories | #Instances | Annotation |
|---|---|---|---|---|---|
| MAR20 | Google Earth (NWPU) | 3842 | 20 | 22,341 | HBB + OBB |
| UCAS-AOD | Aerial (UCAS) | 1510 | 2 | 14,596 | OBB |
| Jilin-1 | Satellite (∼0.75 m GSD) | 1286 | 1 | 8547 | HBB |
4.2. Implementation Details

4.3. Comparison with State-of-the-Art Methods
4.3.1. Results on MAR20 Dataset
4.3.2. Results on UCAS-AOD Dataset
| Category | Method | mAP@0.5 | mAP@0.5:0.95 | Params (M) | FLOPs (G) | FPS |
|---|---|---|---|---|---|---|
| Two-stage | Faster R-CNN [5] | 72.4 | 42.1 | 41.1 | 134.6 | 18.2 |
| Cascade R-CNN [6] | 75.8 | 45.3 | 69.2 | 186.3 | 14.5 | |
| One-stage | RetinaNet [8] | 70.1 | 40.5 | 36.3 | 119.8 | 21.3 |
| YOLOX-L † [63] | 79.2 | 48.8 | 54.2 | 155.6 | 42.1 | |
| YOLOv5-L † [9] | 78.6 | 48.2 | 46.1 | 109.1 | 52.8 | |
| YOLOv7 † [10] | 80.3 | 49.8 | 36.5 | 103.2 | 48.3 | |
| YOLOv8-L † [11] | 81.5 | 51.2 | 43.6 | 165.2 | 45.6 | |
| YOLOv9-C † [12] | 82.1 | 52.0 | 25.3 | 102.1 | 50.1 | |
| YOLOv10-L † [13] | 82.8 | 52.5 | 24.4 | 120.3 | 46.8 | |
| YOLO11-L † [14] | 83.4 | 53.1 | 25.3 | 86.9 | 55.2 | |
| RTMDet-L † [64] | 81.8 | 51.5 | 52.3 | 160.4 | 40.2 | |
| Anchor-free | FCOS [32] | 71.8 | 41.3 | 32.1 | 108.5 | 23.6 |
| CenterNet [33] | 73.5 | 43.8 | 35.6 | 125.2 | 20.1 | |
| Transformer | DETR [18] | 68.5 | 38.2 | 41.3 | 86.1 | 28.4 |
| Deformable DETR [20] | 76.3 | 46.8 | 39.8 | 173.3 | 22.1 | |
| DAB-DETR [21] | 77.8 | 47.9 | 43.7 | 180.2 | 20.5 | |
| DINO [22] | 80.6 | 50.8 | 47.2 | 194.5 | 18.8 | |
| Co-DETR [36] | 81.2 | 51.3 | 48.1 | 198.2 | 17.5 | |
| Lite DETR [39] | 79.5 | 49.2 | 36.8 | 148.5 | 30.2 | |
| RT-DETR-L [23] | 83.1 | 52.8 | 32.0 | 110.0 | 54.8 | |
| RS-specific | Oriented R-CNN [49] | 77.2 | 47.6 | 41.1 | 139.6 | 16.8 |
| LSKNet-S [53] | 82.0 | 51.7 | 31.0 | 98.5 | 38.5 | |
| Ours | SFSMamba-DETR | 85.7 | 55.9 | 34.8 | 118.5 | 48.6 |
| Category | Method | mAP@0.5 | mAP@0.5:0.95 | Params (M) | FLOPs (G) | FPS | ||
|---|---|---|---|---|---|---|---|---|
| Two-stage | Faster R-CNN [5] | 89.3 | 62.5 | 91.2 | 87.4 | 41.1 | 134.6 | 18.2 |
| Cascade R-CNN [6] | 91.5 | 65.0 | 93.1 | 89.9 | 69.2 | 186.3 | 14.5 | |
| One-stage | YOLOv5-L [9] | 92.8 | 66.8 | 94.5 | 91.1 | 46.1 | 109.1 | 52.8 |
| YOLOv7 [10] | 93.5 | 67.3 | 95.2 | 91.8 | 36.5 | 103.2 | 48.3 | |
| YOLOv8-L [11] | 94.2 | 68.5 | 95.8 | 92.6 | 43.6 | 165.2 | 45.6 | |
| YOLOv9-C [12] | 94.6 | 69.1 | 96.0 | 93.2 | 25.3 | 102.1 | 50.1 | |
| YOLOv10-L [13] | 94.9 | 69.3 | 96.2 | 93.6 | 24.4 | 120.3 | 46.8 | |
| YOLO11-L [14] | 95.3 | 69.6 | 96.5 | 94.1 | 25.3 | 86.9 | 55.2 | |
| Transformer | DETR [18] | 85.4 | 58.1 | 87.6 | 83.2 | 41.3 | 86.1 | 28.4 |
| Deformable DETR [20] | 91.8 | 65.2 | 93.5 | 90.1 | 39.8 | 173.3 | 22.1 | |
| DINO [22] | 93.7 | 67.5 | 95.4 | 92.0 | 47.2 | 194.5 | 18.8 | |
| Co-DETR [36] | 94.4 | 68.9 | 95.9 | 92.9 | 48.1 | 198.2 | 17.5 | |
| RT-DETR-L [23] | 96.2 | 71.2 | 97.2 | 95.3 | 32.0 | 110.0 | 54.8 | |
| RS-specific | LSKNet-S [53] | 94.8 | 69.2 | 96.1 | 93.5 | 31.0 | 98.5 | 38.5 |
| Ours | SFSMamba-DETR | 97.6 | 73.2 | 98.5 | 96.7 | 34.8 | 118.5 | 48.6 |
4.3.3. Results on Jilin-1 Satellite Aircraft Detection Dataset
4.3.4. Supplementary Results on Multi-Category Benchmarks
4.4. Ablation Studies
4.4.1. Module Effectiveness
4.4.2. Effect of the Number of SFS Inputs
4.4.3. Effect of Window Sizes in DSWA

| Small Window | Large Window | mAP@0.5 | mAP@0.5 :0.95 | FPS |
|---|---|---|---|---|
| 5 | 10 | 81.4 | 49.8 | 51.2 |
| 5 | 14 | 81.7 | 50.1 | 49.8 |
| 7 | 14 | 82.3 | 50.7 | 48.6 |
| 7 | 21 | 82.0 | 50.3 | 44.1 |
| 14 | 28 | 81.2 | 49.5 | 40.3 |
4.4.4. Effect of Scanning Directions in SS2D
4.5. Qualitative Analysis
4.5.1. Detection Visualization


4.5.2. Failure Mode Analysis
4.5.3. Performance–Efficiency Trade-Off
4.5.4. Feature Activation Visualization
4.5.5. Feature Representation Analysis
4.5.6. Convergence Analysis


5. Discussion
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| SFS | Selective Feature Scanning |
| DSWA | Dual-Scale Window Attention |
| CFAM | Cross-scale Feature Aggregation Module |
| SSM | State Space Model |
| SS2D | 2D Selective Scan |
| VSS | Vision State Space |
| AIFI | Attention-based Intra-scale Feature Interaction |
| DLGPA | Dual Local–Global Perception Attention |
| FPN | Feature Pyramid Network |
| mAP | mean Average Precision |
| FPS | Frames Per Second |
References
- Li, K.; Wan, G.; Cheng, G.; Meng, L.; Han, J. Object Detection in Optical Remote Sensing Images: A Survey and A New Benchmark. ISPRS J. Photogramm. Remote Sens. 2020, 159, 296–307. [Google Scholar] [CrossRef] [Scilit]
- Cheng, G.; Yuan, X.; Yao, X.; Yan, K.; Zeng, Q.; Xie, X.; Han, J. Towards Large-Scale Small Object Detection: Survey and Benchmarks. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 13467–13488. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Xia, G.S.; Bai, X.; Ding, J.; Zhu, Z.; Belongie, S.; Luo, J.; Datcu, M.; Pelillo, M.; Zhang, L. DOTA: A Large-Scale Dataset for Object Detection in Aerial Images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2018; pp. 3974–3983. [Google Scholar]
- Sun, X.; Wang, P.; Yan, Z.; Xu, F.; Wang, R.; Diao, W.; Chen, J.; Li, J.; Feng, Y.; Xu, T.; et al. FAIR1M: A Benchmark Dataset for Fine-grained Object Recognition in High-Resolution Remote Sensing Imagery. ISPRS J. Photogramm. Remote Sens. 2022, 184, 116–130. [Google Scholar] [CrossRef] [Scilit]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. In Proceedings of the Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2015; Volume 28. [Google Scholar]
- Cai, Z.; Vasconcelos, N. Cascade R-CNN: Delving into High Quality Object Detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2018; pp. 6154–6162. [Google Scholar]
- Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.Y.; Berg, A.C. SSD: Single Shot MultiBox Detector. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2016; pp. 21–37. [Google Scholar]
- Lin, T.Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. Focal Loss for Dense Object Detection. In Proceedings of the IEEE International Conference on Computer Vision; IEEE: Piscataway, NJ, USA, 2017; pp. 2980–2988. [Google Scholar]
- Jocher, G. YOLOv5 by Ultralytics. 2021. Available online: https://github.com/ultralytics/yolov5 (accessed on 16 July 2026).
- Wang, C.Y.; Bochkovskiy, A.; Liao, H.Y.M. YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object Detectors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2023; pp. 7464–7475. [Google Scholar]
- Jocher, G.; Chaurasia, A.; Qiu, J. Ultralytics YOLOv8. 2023. Available online: https://github.com/ultralytics/ultralytics (accessed on 16 July 2026).
- Wang, C.Y.; Yeh, I.H.; Liao, H.Y.M. YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2024; pp. 1–21. [Google Scholar]
- Wang, A.; Chen, H.; Liu, L.; Chen, K.; Lin, Z.; Han, J.; Ding, G. YOLOv10: Real-Time End-to-End Object Detection. Adv. Neural Inf. Process. Syst. 2024, 37, 107984–108011. [Google Scholar] [CrossRef] [Scilit]
- Jocher, G.; Qiu, J. YOLO11 by Ultralytics. 2024. Available online: https://github.com/ultralytics/ultralytics (accessed on 16 July 2026).
- Liu, W.; Li, Y.; Zhang, S.; Mao, R. Towards Smart City Supervision: A Detection Pipeline for Illegal Buildings. Eng. Appl. Artif. Intell. 2026, 163, 113052. [Google Scholar] [CrossRef] [Scilit]
- Yan, Z.; Li, Y. AMSRDet: An Adaptive Multi-Scale UAV Infrared-Visible Remote Sensing Vehicle Detection Network. Sensors 2026, 26, 817. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In Proceedings of the International Conference on Learning Representations, Virtual, 3–7 May 2021. [Google Scholar]
- Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; Zagoruyko, S. End-to-End Object Detection with Transformers. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2020; pp. 213–229. [Google Scholar]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is All You Need. In Proceedings of the Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2017; Volume 30. [Google Scholar]
- Zhu, X.; Su, W.; Lu, L.; Li, B.; Wang, X.; Dai, J. Deformable DETR: Deformable Transformers for End-to-End Object Detection. In Proceedings of the International Conference on Learning Representations, Virtual, 3–7 May 2021. [Google Scholar]
- Liu, S.; Li, F.; Zhang, H.; Yang, X.; Qi, X.; Su, H.; Zhu, J.; Zhang, L. DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR. In Proceedings of the International Conference on Learning Representations, Virtual, 25–29 April 2022. [Google Scholar]
- Zhang, H.; Li, F.; Liu, S.; Zhang, L.; Su, H.; Zhu, J.; Ni, L.M.; Shum, H.Y. DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection. In Proceedings of the International Conference on Learning Representations, Virtual, 1–5 May 2023. [Google Scholar]
- Zhao, Y.; Lv, W.; Xu, S.; Wei, J.; Wang, G.; Dang, Q.; Liu, Y.; Chen, J. DETRs Beat YOLOs on Real-time Object Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2024; pp. 16965–16974. [Google Scholar]
- Gu, A.; Dao, T. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. arXiv 2024, arXiv:2312.00752. [Google Scholar]
- Gu, A.; Goel, K.; Ré, C. Efficiently Modeling Long Sequences with Structured State Spaces. In Proceedings of the International Conference on Learning Representations, Virtual, 25–29 April 2022. [Google Scholar]
- Liu, Y.; Tian, Y.; Zhao, Y.; Yu, H.; Xie, L.; Wang, Y.; Ye, Q.; Liu, Y. VMamba: Visual State Space Model. Adv. Neural Inf. Process. Syst. 2024, 37, 103031–103063. [Google Scholar] [CrossRef] [Scilit]
- Zhu, L.; Liao, B.; Zhang, Q.; Wang, X.; Liu, W.; Wang, X. Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model. In Proceedings of the International Conference on Machine Learning, Virtual, 21–27 July 2024. [Google Scholar]
- Sun, P.; Zheng, Y.; Zhou, Z.; Xu, W.; Liu, Q. MAR20: A Benchmark for Military Aircraft Recognition from Remote Sensing Images. Natl. Remote Sens. Bull. 2024, 28, 996–1009. [Google Scholar]
- Zhu, H.; Chen, X.; Dai, W.; Fu, K.; Ye, Q.; Jiao, J. Orientation Robust Object Detection in Aerial Images Using Deep Convolutional Neural Network. In IEEE International Conference on Image Processing; IEEE: Piscataway, NJ, USA, 2015; pp. 3735–3739. [Google Scholar]
- Chang Guang Satellite Technology Co., Ltd. Jilin-1 Satellite Remote Sensing Image Dataset for Aircraft Detection. IEEE Geosci. Remote Sens. Lett. 2023, 20, 1–5. [Google Scholar]
- Lin, T.Y.; Dollár, P.; Girshick, R.; He, K.; Hariharan, B.; Belongie, S. Feature Pyramid Networks for Object Detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2017; pp. 2117–2125. [Google Scholar]
- Tian, Z.; Shen, C.; Chen, H.; He, T. FCOS: Fully Convolutional One-Stage Object Detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: Piscataway, NJ, USA, 2019; pp. 9627–9636. [Google Scholar]
- Duan, K.; Bai, S.; Xie, L.; Qi, H.; Huang, Q.; Tian, Q. CenterNet: Keypoint Triplets for Object Detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: Piscataway, NJ, USA, 2019; pp. 6569–6578. [Google Scholar]
- Wang, C.; He, W.; Nie, Y.; Guo, J.; Liu, C.; Wang, Y.; Han, K. Gold-YOLO: Efficient Object Detector via Gather-and-Distribute Mechanism. Adv. Neural Inf. Process. Syst. 2024, 36, 51094–51112. [Google Scholar] [CrossRef] [Scilit]
- Cheng, T.; Song, L.; Ge, Y.; Liu, W.; Wang, X.; Shan, Y. YOLO-World: Real-Time Open-Vocabulary Object Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2024; pp. 16901–16911. [Google Scholar]
- Zong, Z.; Song, G.; Liu, Y. DETRs with Collaborative Hybrid Assignments Training. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: Piscataway, NJ, USA, 2023; pp. 6748–6758. [Google Scholar]
- Chen, Q.; Chen, X.; Wang, J.; Zhang, S.; Yao, K.; Feng, H.; Han, J.; Ding, E.; Zeng, G.; Wang, J. Group DETR: Fast DETR Training with Group-Wise One-to-Many Assignment. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: Piscataway, NJ, USA, 2023; pp. 6633–6642. [Google Scholar]
- Chen, S.; Sun, P.; Song, Y.; Luo, P. DiffusionDet: Diffusion Model for Object Detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: Piscataway, NJ, USA, 2023; pp. 19830–19843. [Google Scholar]
- Li, F.; Zeng, A.; Liu, S.; Zhang, H.; Li, H.; Zhang, L.; Ni, L.M. Lite DETR: An Interleaved Multi-Scale Encoder for Efficient DETR. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2023; pp. 18558–18567. [Google Scholar]
- Liu, S.; Qi, L.; Qin, H.; Shi, J.; Jia, J. Path Aggregation Network for Instance Segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2018; pp. 8759–8768. [Google Scholar]
- Tan, M.; Pang, R.; Le, Q.V. EfficientDet: Scalable and Efficient Object Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2020; pp. 10781–10790. [Google Scholar]
- Li, H.; Deng, Z.; Li, Y.; Zhu, H. Rethinking Multi-scale Feature Fusion for Object Detection in Remote Sensing Imagery. In Proceedings of the IEEE International Geoscience and Remote Sensing Symposium; IEEE: Piscataway, NJ, USA, 2024; pp. 2783–2786. [Google Scholar]
- Pei, X.; Huang, T.; Xu, C. EfficientVMamba: Atrous Selective Scan for Light Weight Visual Mamba. arXiv 2024, arXiv:2403.09977. [Google Scholar]
- Yang, C.; Chen, Z.; Espinosa, M.; Ericsson, L.; Wang, Z.; Liu, J.; Crowley, E.J. PlainMamba: Improving Non-Hierarchical Mamba in Visual Recognition. arXiv 2024, arXiv:2403.17695. [Google Scholar]
- Chen, K.; Chen, B.; Liu, C.; Li, W.; Zou, Z.; Shi, Z. RSMamba: Remote Sensing Image Classification with State Space Model. IEEE Geosci. Remote Sens. Lett. 2024, 21, 8002605. [Google Scholar] [CrossRef] [Scilit]
- Li, W.; Zhu, J.; Wen, J.; Wang, D. Mamba-in-Mamba: Centralized Mamba-Cross-Scan in Tokenized Mamba Model for Hyperspectral Image Classification. arXiv 2024, arXiv:2405.12003. [Google Scholar]
- Li, Y.; Wang, T.; Luo, N.; Zhou, L.; Chen, Q. Cgmamba: Intelligent Identification of Counterfeit Goods Based on State Space Models. Int. J. Intell. Syst. 2025, 2025, 9939880. [Google Scholar] [CrossRef] [Scilit]
- Huang, Z.; Chen, N.; Xie, M.; Wan, S. MambaYOLO: SSMs-Based YOLO For Object Detection. arXiv 2024, arXiv:2406.05835. [Google Scholar]
- Xie, X.; Cheng, G.; Wang, J.; Yao, X.; Han, J. Oriented R-CNN for Object Detection in Aerial Images. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: Piscataway, NJ, USA, 2021; pp. 3520–3529. [Google Scholar]
- Ding, J.; Xue, N.; Long, Y.; Xia, G.S.; Lu, Q. Learning RoI Transformer for Oriented Object Detection in Aerial Images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2019; pp. 2849–2858. [Google Scholar]
- Yang, X.; Yan, J.; Feng, Z.; He, T. R3Det: Refined Single-Stage Detector with Feature Refinement for Rotating Object. In Proceedings of the AAAI Conference on Artificial Intelligence, Virtual, 2–9 February 2021; Volume 35, pp. 3163–3171. [Google Scholar] [CrossRef] [Scilit]
- Han, J.; Ding, J.; Xue, N.; Xia, G.S. ReDet: A Rotation-equivariant Detector for Aerial Object Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2021; pp. 2786–2795. [Google Scholar]
- Li, Y.; Hou, Q.; Zheng, Z.; Cheng, M.; Yang, J.; Li, X. Large Selective Kernel Network for Remote Sensing Object Detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: Piscataway, NJ, USA, 2023; pp. 16794–16805. [Google Scholar]
- Qin, Z.; Li, Y. DCAM-DETR: Dual Cross-Attention Mamba Detection Transformer for RGB–Infrared Anti-UAV Detection. Information 2026, 17, 103. [Google Scholar] [CrossRef] [Scilit]
- Yuan, H.; Li, Y. CSFADet: Dual-Modal Anti-UAV Detection via Cross-Spectral Feature Alignment and Adaptive Multi-Scale Refinement. Algorithms 2026, 19, 254. [Google Scholar] [CrossRef] [Scilit]
- Yang, Z.; Li, Y.; Shao, J.; Li, W.; Luo, N. Lightweight Hybrid Attention for Channel, Spatial, and Token-level Enhancement. In Proceedings of the 2025 8th International Conference on Computer Information Science and Artificial Intelligence; IEEE: Piscataway, NJ, USA, 2025; pp. 212–217. [Google Scholar]
- Zhou, D.; Li, Y. SoccerDETR: Real-Time Soccer Object Detection via Visual State Space Models with Semantic-Aware Feature Fusion. Technologies 2026, 14, 142. [Google Scholar] [CrossRef] [Scilit]
- Tan, M.; Le, Q. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. In Proceedings of the International Conference on Machine Learning; PMLR: Long Beach, CA, USA, 2019; pp. 6105–6114. [Google Scholar]
- Ding, X.; Zhang, X.; Ma, N.; Han, J.; Ding, G.; Sun, J. RepVGG: Making VGG-style ConvNets Great Again. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2021; pp. 13733–13742. [Google Scholar]
- Rezatofighi, H.; Tsoi, N.; Gwak, J.; Sadeghian, A.; Reid, I.; Savarese, S. Generalized Intersection over Union: A Metric and A Loss for Bounding Box Regression. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2019; pp. 658–666. [Google Scholar]
- Loshchilov, I.; Hutter, F. Decoupled Weight Decay Regularization. In Proceedings of the International Conference on Learning Representations, New Orleans, LA, USA, 6–9 May 2019. [Google Scholar]
- Lin, T.Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; Zitnick, C.L. Microsoft COCO: Common Objects in Context. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2014; pp. 740–755. [Google Scholar]
- Ge, Z.; Liu, S.; Wang, F.; Li, Z.; Sun, J. YOLOX: Exceeding YOLO Series in 2021. arXiv 2021, arXiv:2107.08430. [Google Scholar]
- Lyu, C.; Zhang, W.; Huang, H.; Zhou, Y.; Wang, Y.; Liu, Y.; Zhang, S.; Chen, K. RTMDet: An Empirical Study of Designing Real-Time Object Detectors. arXiv 2022, arXiv:2212.07784. [Google Scholar]
- van der Maaten, L.; Hinton, G. Visualizing Data using t-SNE. J. Mach. Learn. Res. 2008, 9, 2579–2605. [Google Scholar]






| Category | Method | mAP@0.5 | mAP@0.5:0.95 | Recall | Params (M) | FLOPs (G) | FPS | |
|---|---|---|---|---|---|---|---|---|
| Two-stage | Faster R-CNN [5] | 68.2 | 35.6 | 22.3 | 72.1 | 41.1 | 134.6 | 18.2 |
| Cascade R-CNN [6] | 71.5 | 38.4 | 25.1 | 75.3 | 69.2 | 186.3 | 14.5 | |
| One-stage | YOLOv5-L [9] | 73.8 | 41.6 | 27.5 | 77.6 | 46.1 | 109.1 | 52.8 |
| YOLOv8-L [11] | 76.5 | 44.8 | 30.2 | 80.3 | 43.6 | 165.2 | 45.6 | |
| YOLOv9-C [12] | 77.3 | 45.6 | 31.1 | 81.2 | 25.3 | 102.1 | 50.1 | |
| YOLOv10-L [13] | 77.8 | 46.0 | 31.6 | 81.7 | 24.4 | 120.3 | 46.8 | |
| YOLO11-L [14] | 78.1 | 46.3 | 32.0 | 82.1 | 25.3 | 86.9 | 55.2 | |
| Transformer | DETR [18] | 63.5 | 31.2 | 18.6 | 67.4 | 41.3 | 86.1 | 28.4 |
| Deformable DETR [20] | 72.8 | 40.1 | 26.5 | 76.5 | 39.8 | 173.3 | 22.1 | |
| DINO [22] | 76.2 | 44.3 | 29.8 | 80.0 | 47.2 | 194.5 | 18.8 | |
| Co-DETR [36] | 77.0 | 45.2 | 30.8 | 80.8 | 48.1 | 198.2 | 17.5 | |
| RT-DETR-L [23] | 78.5 | 46.8 | 32.5 | 82.6 | 32.0 | 110.0 | 54.8 | |
| RS-specific | LSKNet-S [53] | 77.5 | 45.8 | 31.3 | 81.4 | 31.0 | 98.5 | 38.5 |
| Ours | SFSMamba-DETR | 82.3 | 50.7 | 36.8 | 86.2 | 34.8 | 118.5 | 48.6 |
| Dataset | Method | mAP@0.5 | mAP@0.5:0.95 | FPS |
|---|---|---|---|---|
| DOTA-v1.0 (HBB) | YOLO11-L [14] | 73.1 | 45.6 | 49.7 |
| RT-DETR-L [23] | 72.8 | 45.1 | 48.9 | |
| SFSMamba-DETR | 74.8 | 47.0 | 43.6 | |
| DIOR | YOLO11-L [14] | 77.0 | 49.2 | 51.4 |
| RT-DETR-L [23] | 77.2 | 49.6 | 50.1 | |
| SFSMamba-DETR | 78.9 | 51.1 | 44.8 |
| SFS | DSWA | CFAM | EfficientNet | mAP@0.5 | mAP@0.5:0.95 |
|---|---|---|---|---|---|
| × | × | × | × | 78.5 | 46.8 |
| × | × | ✓ | × | 79.0 | 47.4 |
| × | × | × | ✓ | 79.2 | 47.6 |
| ✓ | × | × | × | 80.1 | 48.5 |
| × | ✓ | × | × | 79.8 | 48.1 |
| ✓ | ✓ | × | × | 81.0 | 49.3 |
| ✓ | ✓ | ✓ | × | 81.5 | 49.8 |
| ✓ | ✓ | ✓ | ✓ | 82.3 | 50.7 |
| Main | Guide | Auxiliary | mAP@0.5 | mAP@0.5:0.95 |
|---|---|---|---|---|
| ✓ | × | × | 79.3 | 47.6 |
| ✓ | ✓ | × | 79.7 | 48.0 |
| ✓ | ✓ | ✓ | 80.1 | 48.5 |
| Directions | Description | mAP@0.5 | mAP@0.5 :0.95 |
|---|---|---|---|
| 1 | Left-to-right only | 80.5 | 48.3 |
| 2 | Horizontal bidirectional | 81.2 | 49.2 |
| 2 | Vertical bidirectional | 81.0 | 49.0 |
| 4 | All four directions | 82.3 | 50.7 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Cai, Y.; Zhao, J.; Wu, H.; Ma, R. SFSMamba-DETR: Selective Feature Scanning with State Space Models and Dual-Scale Window Attention for Remote Sensing Object Detection. Remote Sens. 2026, 18, 2835. https://doi.org/10.3390/rs18162835
Cai Y, Zhao J, Wu H, Ma R. SFSMamba-DETR: Selective Feature Scanning with State Space Models and Dual-Scale Window Attention for Remote Sensing Object Detection. Remote Sensing. 2026; 18(16):2835. https://doi.org/10.3390/rs18162835
Chicago/Turabian StyleCai, Yuanli, Junchao Zhao, Husheng Wu, and Rui Ma. 2026. "SFSMamba-DETR: Selective Feature Scanning with State Space Models and Dual-Scale Window Attention for Remote Sensing Object Detection" Remote Sensing 18, no. 16: 2835. https://doi.org/10.3390/rs18162835
APA StyleCai, Y., Zhao, J., Wu, H., & Ma, R. (2026). SFSMamba-DETR: Selective Feature Scanning with State Space Models and Dual-Scale Window Attention for Remote Sensing Object Detection. Remote Sensing, 18(16), 2835. https://doi.org/10.3390/rs18162835

