HRL-Det: Hierarchical Reinforcement Learning for Sequential Object Detection in Aerial Imagery
Abstract
1. Introduction
- We propose HRL-Det, a hierarchical reinforcement learning framework for sequential object detection that reduces redundant dense-image evaluation in high-resolution UAV imagery.
- We design a Neural ODE-driven Continuous-Time Bellman State Evolution Module that models the agent’s internal state dynamics as a continuous-time stochastic process governed by neural stochastic differential equations. The resulting evolved state representation is used by a standard Dueling Double DQN for discrete action selection, improving the temporal resolution of the search policy and enhancing robustness to visual ambiguity.
- We introduce a Lyapunov-Guided Entropy-Regularized Reward Shaping Mechanism, which provides convergence-promoting dense reward signals informed by Lyapunov stability analysis and integrates maximum entropy policy optimization. It effectively addresses the challenges of reward sparsity and premature policy collapse, thereby promoting stable and accelerated training convergence.
- Extensive experiments on the VisDrone2019, DroneVehicle, and MS COCO 2017 datasets demonstrate that our HRL-Det outperforms existing RL-based detectors in terms of both detection precision and inference efficiency.
2. Related Work
2.1. Deep Learning for Object Detection in UAV Aerial Imagery
2.2. Reinforcement Learning for Object Detection
3. Proposed Method: HRL-Det
3.1. Markov Decision Process Formulation
3.1.1. State Space
3.1.2. Action Space
3.1.3. Reward Function
3.1.4. State Transition
3.1.5. Episode Termination
- The agent selects the trigger action;
- The maximum search horizon is reached;
- The bounding box becomes invalid or moves outside the image boundary.
3.1.6. Relationship Between Continuous-Time Modeling and Discrete Decision-Making
3.2. Neural ODE-Driven Continuous-Time Bellman State Evolution
3.3. Lyapunov-Guided Entropy-Regularized Reward Shaping
3.4. From Continuous-Time Theory to Discrete-Time Agent
4. Experimental Results and Analysis
4.1. Datasets and Evaluation Metrics
4.2. Implementation Details
4.3. Quantitative Comparison with State-of-the-Arts
4.4. Ablation Study
4.5. Training Convergence Assessment
4.6. Hierarchical Search Process and Spatial Attention
4.7. Qualitative Analysis and Localization Accuracy
5. Discussion
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| HRL-Det | Hierarchical Reinforcement Learning Detector |
| DRL | Deep Reinforcement Learning |
| DQN | Deep Q-Network |
| MDP | Markov Decision Process |
| ODE | Ordinary Differential Equation |
| SDE | Stochastic Differential Equation |
| HJB | Hamilton–Jacobi–Bellman |
| IoU | Intersection over Union |
| mAP | mean Average Precision |
| UAV | Unmanned Aerial Vehicle |
| NODE-BSE | Neural ODE Bellman State Evolution |
| GAP | Global Average Pooling |
| GRU | Gated Recurrent Unit |
| PER | Prioritized Experience Replay |
| TD | Temporal Difference |
References
- Bouguettaya, A.; Zarzour, H.; Kechida, A.; Taberkit, A.M. Vehicle Detection from UAV Imagery with Deep Learning: A Review. IEEE Trans. Neural Netw. Learn. Syst. 2022, 33, 6047–6067. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lygouras, E.; Santavas, N.; Taitzoglou, A.; Tarchanidis, K.; Mitropoulos, A.; Gasteratos, A. Unsupervised Human Detection with an Embedded Vision System on a Fully Autonomous UAV for Search and Rescue Operations. Sensors 2019, 19, 3542. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Nguyen, V.N.; Jenssen, R.; Roverso, D. Automatic Autonomous Vision-Based Power Line Inspection: A Review of Current Status and the Potential Role of Deep Learning. Int. J. Electr. Power Energy Syst. 2018, 99, 107–120. [Google Scholar] [CrossRef] [Scilit]
- Wei, X.; Yang, Z.; Liu, Y.; Wei, D.; Jia, L.; Li, Y. Railway Track Fastener Defect Detection Based on Image Processing and Deep Learning Techniques: A Comparative Study. Eng. Appl. Artif. Intell. 2019, 80, 66–81. [Google Scholar] [CrossRef] [Scilit]
- Du, D.; Zhu, P.; Wen, L.; Bian, X.; Lin, H.; Hu, Q.; Peng, T.; Zheng, J.; Wang, X.; Zhang, Y.; et al. VisDrone-DET2019: The Vision Meets Drone Object Detection in Image Challenge Results. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), Seoul, Republic of Korea, 27–28 October 2019; pp. 213–226. [Google Scholar]
- Zhu, P.; Wen, L.; Du, D.; Bian, X.; Fan, H.; Hu, Q.; Ling, H. Detection and Tracking Meet Drones Challenge. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 44, 7380–7399. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 39, 1137–1149. [Google Scholar] [PubMed]
- Cai, Z.; Vasconcelos, N. Cascade R-CNN: Delving into High Quality Object Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–23 June 2018; pp. 6154–6162. [Google Scholar]
- Jocher, G.; Chaurasia, A.; Stoken, A.; Borovec, J.; NanoCode012; Kwon, Y.; Michael, K.; Xie, T.; Fang, J.; imyhxy; et al. Ultralytics/YOLOv5: V7.0—YOLOv5 SOTA Realtime Instance Segmentation. Zenodo 2022. [Google Scholar] [CrossRef]
- Jocher, G.; Chaurasia, A.; Qiu, J. Ultralytics YOLOv8. Available online: https://github.com/ultralytics/ultralytics (accessed on 15 January 2025).
- Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.-Y.; Berg, A.C. SSD: Single Shot MultiBox Detector. In Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands, 11–14 October 2016; pp. 21–37. [Google Scholar]
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 11–17 October 2021; pp. 10012–10022. [Google Scholar]
- Zhou, X.; Wang, D.; Krahenbuhl, P. Objects as Points. arXiv 2019, arXiv:1904.07850. [Google Scholar]
- Li, C.; Yang, T.; Zhu, S.; Chen, C.; Guan, S. Density Map Guided Object Detection in Aerial Images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Seattle, WA, USA, 14–19 June 2020; pp. 737–746. [Google Scholar]
- Mittal, P.; Singh, R.; Sharma, A. Deep Learning-Based Object Detection in Low-Altitude UAV Datasets: A Survey. Image Vis. Comput. 2020, 104, 104046. [Google Scholar]
- Cheng, G.; Yuan, X.; Yao, X.; Yan, K.; Xie, X.; Zeng, Q.; Han, J. Towards Large-Scale Small Object Detection: Survey and Benchmarks. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 13467–13488. [Google Scholar] [PubMed]
- Caicedo, J.C.; Lazebnik, S. Active Object Localization with Deep Reinforcement Learning. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Santiago, Chile, 7–13 December 2015; pp. 2488–2496. [Google Scholar]
- Bellver, M.; Giro-i-Nieto, X.; Marques, F.; Torres, J. Hierarchical Object Detection with Deep Reinforcement Learning. In Proceedings of the NIPS 2016 Deep Reinforcement Learning Workshop, Barcelona, Spain, 5–10 December 2016. [Google Scholar]
- Le, N.; Rathour, V.S.; Yamazaki, K.; Luu, K.; Savvides, M. Deep Reinforcement Learning in Computer Vision: A Comprehensive Survey. Artif. Intell. Rev. 2022, 55, 2733–2819. [Google Scholar]
- Uzkent, B.; Yoon, S. Learning When and Where to Zoom with Deep Reinforcement Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 12345–12354. [Google Scholar]
- Liu, S.; Huang, D.; Wang, Y. Pay Attention to Them: Deep Reinforcement Learning-Based Cascade Object Detection. IEEE Trans. Neural Netw. Learn. Syst. 2020, 31, 2544–2556. [Google Scholar] [PubMed]
- van Hasselt, H.; Guez, A.; Silver, D. Deep Reinforcement Learning with Double Q-Learning. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), Phoenix, AZ, USA, 12–17 February 2016; pp. 2094–2100. [Google Scholar]
- Zhou, X.; Chen, Y.; Wang, H.; Li, J. Hybrid DQN-Based Low-Computational Reinforcement Learning Object Detection with Adaptive Dynamic Reward Function and ROI Align-Based Bounding Box Regression. IEEE Trans. Image Process. 2025, 34, 1024–1038. [Google Scholar]
- Kong, X.; Xin, B.; Wang, Y.; Hua, G. Collaborative Deep Reinforcement Learning for Joint Object Search. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 1695–1704. [Google Scholar]
- Lin, T.-Y.; Goyal, P.; Girshick, R.; He, K.; Dollar, P. Focal Loss for Dense Object Detection. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017; pp. 2980–2988. [Google Scholar]
- Wang, A.; Chen, H.; Liu, L.; Chen, K.; Lin, Z.; Han, J.; Ding, G. YOLOv10: Real-Time End-to-End Object Detection. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada, 9–15 December 2024. [Google Scholar]
- Wang, C.-Y.; Yeh, I.-H.; Liao, H.-Y.M. YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information. In Proceedings of the European Conference on Computer Vision (ECCV), Milan, Italy, 29 September–4 October 2024; pp. 1–21. [Google Scholar]
- Zhu, X.; Su, W.; Lu, L.; Li, B.; Wang, X.; Dai, J. Deformable DETR: Deformable Transformers for End-to-End Object Detection. In Proceedings of the International Conference on Learning Representations (ICLR), Virtual, 3–7 May 2021. [Google Scholar]
- Lv, W.; Xu, S.; Zhao, Y.; Wang, G.; Wei, J.; Cui, C.; Du, Y.; Dang, Q.; Liu, Y. DETRs Beat YOLOs on Real-Time Object Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 17–21 June 2024; pp. 16965–16974. [Google Scholar]
- Zhang, H.; Li, F.; Liu, S.; Zhang, L.; Su, H.; Zhu, J.; Ni, L.M.; Shum, H.-Y. DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection. arXiv 2022, arXiv:2203.03605. [Google Scholar]
- Mathe, S.; Pirinen, A.; Sminchisescu, C. Reinforcement Learning for Visual Object Detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 2894–2902. [Google Scholar]
- Jie, Z.; Liang, X.; Feng, J.; Jin, X.; Lu, W.; Yan, S. Tree-Structured Reinforcement Learning for Sequential Object Localization. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Barcelona, Spain, 5–10 December 2016; pp. 127–135. [Google Scholar]
- Pirinen, A.; Sminchisescu, C. Deep Reinforcement Learning of Region Proposal Networks for Object Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–23 June 2018; pp. 6945–6954. [Google Scholar]
- Ding, W.; Majcherczyk, N.; Deshpande, M.; Qi, X.; Zhao, D.; Madhivanan, R.; Sen, A. Learning to View: Decision Transformers for Active Object Detection. In Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), London, UK, 29 May–2 June 2023; pp. 7140–7146. [Google Scholar]
- Zhang, J.; Yang, X.; He, W.; Ren, J.; Zhang, Q.; Zhao, Y.; Bai, R.; He, X.; Liu, J. Scale Optimization Using Evolutionary Reinforcement Learning for Object Detection on Drone Imagery. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), Vancouver, BC, Canada, 20–27 February 2024; Volume 38, pp. 410–418. [Google Scholar]
- Akhloufi, M.A.; Arola, S.; Bonnet, A. Drones Chasing Drones: Reinforcement Learning and Deep Search Area Proposal. Drones 2019, 3, 58. [Google Scholar] [CrossRef] [Scilit]
- Alpdemir, M.N.; Sezgin, M. A Reinforcement Learning (RL)-Based Hybrid Method for Ground Penetrating Radar (GPR)-Driven Buried Object Detection. Neural Comput. Appl. 2024, 36, 8199–8219. [Google Scholar] [CrossRef] [Scilit]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 770–778. [Google Scholar]
- Dormand, J.R.; Prince, P.J. A Family of Embedded Runge-Kutta Formulae. J. Comput. Appl. Math. 1980, 6, 19–26. [Google Scholar] [CrossRef] [Scilit]
- Ng, A.Y.; Harada, D.; Russell, S. Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping. In Proceedings of the International Conference on Machine Learning (ICML), Bled, Slovenia, 27–30 June 1999; pp. 278–287. [Google Scholar]
- Haarnoja, T.; Zhou, A.; Abbeel, P.; Levine, S. Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. In Proceedings of the International Conference on Machine Learning (ICML), Stockholm, Sweden, 10–15 July 2018; pp. 1861–1870. [Google Scholar]
- Sun, Y.; Cao, B.; Zhu, P.; Hu, Q. DroneVehicle: Drone-Based RGB-Infrared Vehicle Detection Benchmark. IEEE Trans. Circuits Syst. Video Technol. 2022, 32, 6734–6745. [Google Scholar]
- Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollar, P.; Zitnick, C.L. Microsoft COCO: Common Objects in Context. In Proceedings of the European Conference on Computer Vision (ECCV), Zurich, Switzerland, 6–12 September 2014; pp. 740–755. [Google Scholar]
- Zong, Z.; Song, G.; Liu, Y. DETRs with Collaborative Hybrid Assignments Training. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 1–6 October 2023; pp. 6748–6758. [Google Scholar]










| Method (Year) | VisDrone2019 | DroneVehicle | MS COCO 2017 | FLOPs (G) | Steps ↓ | FPS ↑ | |||
|---|---|---|---|---|---|---|---|---|---|
| mAP@0.5 | mAP@0.5:0.95 | mAP@0.5 | mAP@0.5:0.95 | mAP@0.5 | mAP@0.5:0.95 | ||||
| Caicedo & Lazebnik (2015) † [17] | 0.271 | 0.131 | 0.634 | 0.397 | 0.521 | 0.318 | 3.82 | 11.4 | 47 |
| Mathe et al. (2016) † [31] | 0.289 | 0.144 | 0.648 | 0.409 | 0.543 | 0.331 | 4.17 | 10.8 | 51 |
| Bellver et al. (2016) † [18] | 0.314 | 0.168 | 0.671 | 0.428 | 0.578 | 0.362 | 3.54 | 9.2 | 59 |
| Jie et al. (2016) † [32] | 0.302 | 0.157 | 0.658 | 0.416 | 0.559 | 0.347 | 5.06 | 10.1 | 54 |
| Kong et al. (2017) † [24] | 0.311 | 0.163 | 0.665 | 0.422 | 0.568 | 0.355 | 4.73 | 9.6 | 52 |
| Pirinen & Sminchisescu (2018) † [33] | 0.328 | 0.178 | 0.687 | 0.441 | 0.601 | 0.381 | 3.91 | 8.7 | 62 |
| Liu et al. (2020) † [21] | 0.349 | 0.198 | 0.718 | 0.473 | 0.631 | 0.412 | 3.28 | 7.9 | 64 |
| Uzkent & Yoon (2020) † [20] | 0.341 | 0.189 | 0.703 | 0.459 | 0.617 | 0.398 | 2.93 | 8.1 | 68 |
| Ding et al. (2023) † [34] | 0.363 | 0.215 | 0.744 | 0.511 | 0.671 | 0.441 | 2.87 | 7.2 | 71 |
| Zhang et al. (2024) † [35] | 0.369 | 0.219 | 0.752 | 0.527 | 0.685 | 0.452 | 3.14 | 7.5 | 59 |
| LHAR-RLD (2025) † [23] | 0.374 | 0.224 | 0.760 | 0.543 | 0.699 | 0.463 | 1.43 | 8.7 | 79 |
| Ours—HRL-Det (2026) † | 0.412 | 0.251 | 0.812 | 0.578 | 0.735 | 0.512 | 2.61 | 6.3 | 86 |
| Method (Year) | Params (M) | VisDrone2019 | DroneVehicle | MS COCO 2017 | FPS ↑ | |||
|---|---|---|---|---|---|---|---|---|
| mAP@0.5 | mAP@0.5:0.95 | mAP@0.5 | mAP@0.5:0.95 | mAP@0.5 | mAP@0.5:0.95 | |||
| Faster R-CNN (2017) [7] | 41.8 | 0.256 | 0.118 | 0.621 | 0.384 | 0.565 | 0.362 | 31 |
| SSD (2016) [11] | 26.3 | 0.198 | 0.089 | 0.549 | 0.312 | 0.412 | 0.232 | 82 |
| RetinaNet (2017) [25] | 37.7 | 0.274 | 0.131 | 0.648 | 0.401 | 0.593 | 0.387 | 37 |
| Deformable DETR (2021) [28] | 40.1 | 0.331 | 0.187 | 0.712 | 0.481 | 0.643 | 0.449 | 29 |
| YOLOv5 (2022) [9] | 46.5 | 0.352 | 0.201 | 0.698 | 0.447 | 0.624 | 0.421 | 108 |
| YOLOv8 (2023) [10] | 43.7 | 0.378 | 0.228 | 0.741 | 0.521 | 0.668 | 0.478 | 131 |
| RT-DETR (2023) [29] | 42.0 | 0.389 | 0.236 | 0.756 | 0.532 | 0.693 | 0.501 | 114 |
| YOLOv9 (2024) [27] | 57.3 | 0.382 | 0.231 | 0.748 | 0.527 | 0.674 | 0.483 | 87 |
| DINO (2022) [30] | 47.0 | 0.403 | 0.244 | 0.779 | 0.561 | 0.720 | 0.511 | 22 |
| YOLOv10 (2024) [26] | 38.4 | 0.391 | 0.238 | 0.762 | 0.538 | 0.681 | 0.491 | 148 |
| Co-DETR (2023) [44] | 146.0 | 0.407 | 0.248 | 0.786 | 0.569 | 0.724 | 0.519 | 14 |
| Ours—HRL-Det (2026) † | 17.3 | 0.412 | 0.251 | 0.812 | 0.578 | 0.735 | 0.512 | 86 ‡ |
| Configuration | NODE-BSE | Lyap. Reward | Ent.-Masking | mAP@0.5 | mAP@0.5:0.95 |
|---|---|---|---|---|---|
| w/o NODE-BSE | – | ✓ | ✓ | 0.364 | 0.219 |
| w/o Lyap. Reward | ✓ | – | ✓ | 0.381 | 0.227 |
| w/o Ent.-Masking | ✓ | ✓ | – | 0.390 | 0.235 |
| Full HRL-Det (Ours) | ✓ | ✓ | ✓ | 0.412 | 0.251 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Li, M.; Hu, Y. HRL-Det: Hierarchical Reinforcement Learning for Sequential Object Detection in Aerial Imagery. Sensors 2026, 26, 4232. https://doi.org/10.3390/s26134232
Li M, Hu Y. HRL-Det: Hierarchical Reinforcement Learning for Sequential Object Detection in Aerial Imagery. Sensors. 2026; 26(13):4232. https://doi.org/10.3390/s26134232
Chicago/Turabian StyleLi, Meng, and Yaowen Hu. 2026. "HRL-Det: Hierarchical Reinforcement Learning for Sequential Object Detection in Aerial Imagery" Sensors 26, no. 13: 4232. https://doi.org/10.3390/s26134232
APA StyleLi, M., & Hu, Y. (2026). HRL-Det: Hierarchical Reinforcement Learning for Sequential Object Detection in Aerial Imagery. Sensors, 26(13), 4232. https://doi.org/10.3390/s26134232

