HSMD-YOLO: An Anti-Aliasing Feature-Enhanced Network for High-Speed Microbubble Detection
Abstract
1. Introduction
2. Related Work
2.1. Bubble Detection and Measurement
2.2. Single-Stage and End-to-End High-Resolution Small Object Detection
2.3. Feature Fusion Mechanism
3. Methods
3.1. Scale Switch Block (SSB)
3.2. Global-Local Refine Block (GLRB)
3.2.1. Multi-DConv Head Transposed Self-Attention (MDTA)
3.2.2. Gated-Dconv Feed-Forward Network (GDFN)
3.3. Bidirectional Exponential Moving Attention Fusion (BEMAF)
4. Experiment and Analysis
4.1. Dataset
4.2. Experiment Settings
4.3. Evaluation Metrics
4.4. Compared to State-of-the-Art Methods
4.5. Ablation Study
4.6. Visualization Analysis
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Reed, A.M.; Milgram, J.H. Ship wakes and their radar images. Annu. Rev. Fluid Mech. 2002, 34, 469–502. [Google Scholar] [CrossRef]
- Trevorrow, M.V.; Vagle, S.; Farmer, D.M. Acoustical measurements of microbubbles within ship wakes. J. Acoust. Soc. Am. 1994, 95, 1922–1932. [Google Scholar] [CrossRef]
- Mazzeo, A.; Renga, A.; Graziano, M.D. A Systematic Review of Ship Wake Detection Methods in Satellite Imagery. Remote Sens. 2024, 16, 3775. [Google Scholar] [CrossRef]
- Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You Only Look Once: Unified, Real-Time Object Detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2016; pp. 779–788. [Google Scholar] [CrossRef]
- Bochkovskiy, A.; Wang, C.Y.; Liao, H.Y.M. YOLOv4: Optimal Speed and Accuracy of Object Detection. arXiv 2020, arXiv:2004.10934. [Google Scholar] [CrossRef]
- Hibiki, T.; Ishii, M. Active nucleation site density in boiling systems: Part II—subcooled boiling flow. Int. J. Heat Mass Transf. 2003, 46, 2603–2615. [Google Scholar] [CrossRef]
- Odena, A.; Dumoulin, V.; Olah, C. Deconvolution and Checkerboard Artifacts. Distill 2016, 1, e3. [Google Scholar] [CrossRef]
- Zhang, R. Making Convolutional Networks Shift-Invariant Again. arXiv 2019, arXiv:1904.11486. [Google Scholar] [CrossRef]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention Is All You Need. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS); Curran Associates Inc.: Red Hook, NY, USA, 2017; pp. 5998–6008. [Google Scholar]
- Katharopoulos, A.; Vyas, A.; Pappas, N.; Fleuret, F. Transformers are RNNs: Fast autoregressive transformers with linear attention. In Proceedings of the International Conference on Machine Learning (ICML); PMLR: Cambridge, MA, USA, 2020; pp. 5156–5165. [Google Scholar]
- Choromanski, K.; Likhosherstov, V.; Dohan, D.; Song, X.; Gane, A.; Sarlos, T.; Hawkins, P.; Davis, J.; Mohiuddin, A.; Kaiser, L.; et al. Rethinking attention with performers. In Proceedings of the International Conference on Learning Representations (ICLR); ICLR: Vienna, Austria, 2021. [Google Scholar]
- Yang, Z.; Liu, S.; Hu, H.; Wang, L.; Lin, S. RepPoints: Point set representation for object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: New York, NY, USA, 2019; pp. 9657–9666. [Google Scholar] [CrossRef]
- Rezatofighi, H.; Tsoi, N.; Gwak, J.; Sadeghian, A.; Reid, I.; Savarese, S. Generalized Intersection over Union: A Metric and a Loss for Bounding Box Regression. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2019; pp. 658–666. [Google Scholar] [CrossRef]
- Li, X.; Wang, W.; Wu, L.; Chen, S.; Hu, X.; Li, J.; Tang, J.; Yang, J. Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object Detection. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS); Curran Associates Inc.: Red Hook, NY, USA, 2020. [Google Scholar] [CrossRef]
- Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional Networks for Biomedical Image Segmentation. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI); Springer: Cham, Switzerland, 2015; pp. 234–241. [Google Scholar] [CrossRef]
- Chen, L.C.; Zhu, Y.; Papandreou, G.; Schroff, F.; Adam, H. Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation. In Proceedings of the European Conference on Computer Vision (ECCV); Springer: Cham, Switzerland, 2018; pp. 801–818. [Google Scholar] [CrossRef]
- He, K.; Gkioxari, G.; Dollár, P.; Girshick, R. Mask R-CNN. In Proceedings of the IEEE International Conference on Computer Vision (ICCV); IEEE: New York, NY, USA, 2017; pp. 2961–2969. [Google Scholar] [CrossRef]
- Zhao, Y.; Lv, W.; Xu, S.; Wei, J.; Wang, G.; Dang, Q.; Liu, Y.; Chen, J. DETRs Beat YOLOs on Real-Time Object Detection. arXiv 2023, arXiv:2304.08069. [Google Scholar]
- Zhou, W.; Miwa, S.; Tsujimura, R.; Nguyen, T.B.; Okawa, T.; Okamoto, K. Bubble feature extraction in subcooled flow boiling using AI-based object detection and tracking techniques. Int. J. Heat Mass Transf. 2024, 222, 125188. [Google Scholar] [CrossRef]
- Zhu, X.; Hu, H.; Lin, S.; Dai, J. Deformable ConvNets v2: More Deformable, Better Results. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2019; pp. 9308–9316. [Google Scholar] [CrossRef]
- Zhu, X.; Su, W.; Lu, L.; Li, B.; Wang, X.; Dai, J. Deformable DETR: Deformable Transformers for End-to-End Object Detection. In Proceedings of the International Conference on Learning Representations (ICLR); ICLR: Vienna, Austria, 2021. [Google Scholar] [CrossRef]
- Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.Y.; Berg, A.C. SSD: Single Shot MultiBox Detector. In Proceedings of the European Conference on Computer Vision (ECCV); Springer: Cham, Switzerland, 2016; pp. 21–37. [Google Scholar]
- Lin, T.Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. Focal Loss for Dense Object Detection. In Proceedings of the IEEE International Conference on Computer Vision (ICCV); IEEE: New York, NY, USA, 2017; pp. 2980–2988. [Google Scholar] [CrossRef]
- Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; Zagoruyko, S. End-to-End Object Detection with Transformers. In Proceedings of the European Conference on Computer Vision (ECCV); Springer: Berlin/Heidelberg, Germany, 2020; pp. 213–229. [Google Scholar]
- Liu, S.; Li, F.; Zhang, H.; Yang, X.; Qi, X.; Su, H.; Zhu, J.; Zhang, L. DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR. In Proceedings of the International Conference on Learning Representations (ICLR); ICLR: Online, 2022. [Google Scholar]
- Li, F.; Zhang, H.; Liu, S.; Guo, J.; Ni, L.M.; Zhang, L. DN-DETR: Accelerate DETR Training by Introducing Query DeNoising. arXiv 2022, arXiv:2203.01305. [Google Scholar] [CrossRef]
- Zhang, H.; Li, F.; Liu, S.; Zhang, L.; Su, H.; Zhu, J.; Ni, L.M.; Shum, H.Y. DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection. In Proceedings of the International Conference on Learning Representations (ICLR); ICLR: Kigali, Rwanda, 2023. [Google Scholar]
- Lin, T.Y.; Dollár, P.; Girshick, R.; He, K.; Hariharan, B.; Belongie, S. Feature Pyramid Networks for Object Detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2017; pp. 2117–2125. [Google Scholar]
- Liu, S.; Qi, L.; Qin, H.; Shi, J.; Jia, J. Path Aggregation Network for Instance Segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2018; pp. 8759–8768. [Google Scholar]
- Ghiasi, G.; Lin, T.Y.; Le, Q.V. NAS-FPN: Learning Scalable Feature Pyramid Architecture for Object Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2019; pp. 7036–7045. [Google Scholar]
- Tan, M.; Pang, R.; Le, Q.V. EfficientDet: Scalable and Efficient Object Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2020; pp. 10781–10790. [Google Scholar]
- Qiao, S.; Chen, L.C.; Yuille, A. DetectoRS: Detecting Objects with Recursive Feature Pyramid and Switchable Atrous Convolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2021; pp. 10213–10224. [Google Scholar]
- Sun, K.; Xiao, B.; Liu, D.; Wang, J. High-Resolution Representations for Labeling Pixels and Regions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2019; pp. 3370–3379. [Google Scholar]
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: New York, NY, USA, 2021; pp. 10012–10022. [Google Scholar]
- Wang, W.; Xie, E.; Li, X.; Fan, D.P.; Song, K.; Liang, D.; Lu, T.; Luo, P.; Shao, L. Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: New York, NY, USA, 2021; pp. 548–558. [Google Scholar]
- Cheng, G.; Yuan, X.; Yao, X.; Yan, K.; Zeng, Q.; Xie, X.; Han, J. Towards Large-Scale Small Object Detection: Survey and Benchmarks. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 13467–13488. [Google Scholar] [CrossRef] [PubMed]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS); MIT Press: Cambridge, MA, USA, 2015; Volume 28. [Google Scholar]
- Zhou, X.; Wang, D.; Krähenbühl, P. Objects as Points. arXiv 2019, arXiv:1904.07850. [Google Scholar] [PubMed]
- Zhu, X.; Lyu, S.; Wang, X.; Zhao, Q. TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-Captured Scenarios. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (ICCVW); IEEE: New York, NY, USA, 2021; pp. 2778–2788. [Google Scholar]
- Jocher, G.; Chaurasia, A.; Qiu, J. Ultralytics YOLOv8. 2023. Available online: https://github.com/ultralytics/ultralytics (accessed on 12 January 2026).
- Wang, C.Y.; Yeh, I.H.; Liao, H.Y.M. YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information. arXiv 2024, arXiv:2402.13616. [Google Scholar] [CrossRef]
- Wang, A.; Chen, H.; Liu, L.; Chen, K.; Lin, Z.; Han, J.; Ding, G. YOLOv10: Real-Time End-to-End Object Detection. arXiv 2024, arXiv:2405.14458. [Google Scholar]
- Jocher, G.; Qiu, J. YOLO11 by Ultralytics. 2024. Available online: https://github.com/ultralytics/ultralytics (accessed on 12 January 2026).
- Tian, Y.; Ye, Q.; Doermann, D. YOLOv12: Attention-Centric Real-Time Object Detectors. arXiv 2025, arXiv:2502.12524. [Google Scholar]
- Liu, M.; Wang, J.; Wang, S.; Li, Y. YOLOv13: Real-Time Object Detection with Hypergraph-Enhanced Adaptive Aggregation. arXiv 2025, arXiv:2506.17733. [Google Scholar]
- Wang, C.; He, W.; Nie, Y.; Guo, J.; Liu, C.; Wang, Y.; Han, K. Gold-YOLO: Efficient Object Detector via Gather-and-Distribute Mechanism. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS); Curran Associates Inc.: Red Hook, NY, USA, 2023; Volume 36. [Google Scholar]
- Wang, J.; Xu, C.; Yang, W.; Yu, L. A Normalized Gaussian Wasserstein Distance for Tiny Object Detection. arXiv 2021, arXiv:2110.13389. [Google Scholar]













| Model | Precision | Recall | mAP@0.5 | mAP@50:95 | Params (M) | FLOPs (G) |
|---|---|---|---|---|---|---|
| Faster R-CNN [37] | 0.865 | 0.830 | 0.845 | 0.610 | 41.5 | 180.0 |
| CenterNet [38] | 0.872 | 0.842 | 0.856 | 0.625 | 32.6 | 140.5 |
| SSD [22] | 0.820 | 0.763 | 0.803 | 0.359 | 26.3 | 34.4 |
| TPH-YOLOv5 [39] | 0.877 | 0.851 | 0.878 | 0.618 | 7.5 | 17.0 |
| YOLOv8n [40] | 0.906 | 0.867 | 0.905 | 0.690 | 3.2 | 8.7 |
| YOLOv9t [41] | 0.901 | 0.868 | 0.906 | 0.688 | 2.0 | 7.7 |
| YOLOv10n [42] | 0.892 | 0.853 | 0.887 | 0.663 | 2.3 | 6.7 |
| YOLOv11n [43] | 0.900 | 0.868 | 0.904 | 0.687 | 9.0 | 6.6 |
| YOLOv12n [44] | 0.894 | 0.870 | 0.903 | 0.687 | 8.8 | 6.4 |
| YOLOv13n [45] | 0.893 | 0.870 | 0.900 | 0.681 | 8.5 | 6.2 |
| RT-DETR-R18 [18] | 0.863 | 0.831 | 0.833 | 0.633 | 20.0 | 60.0 |
| Gold-YOLO-n [46] | 0.896 | 0.862 | 0.899 | 0.675 | 5.6 | 12.1 |
| HSMD-YOLO (ours) | 0.908 | 0.874 | 0.911 | 0.703 | 11.4 | 39.3 |
| Model | Precision | Recall | mAP@ | mAP@50:95 |
|---|---|---|---|---|
| YOLOv8n | 0.824 | 0.751 | 0.836 | 0.387 |
| YOLOv9n | 0.840 | 0.770 | 0.843 | 0.386 |
| YOLOv10n | 0.782 | 0.726 | 0.808 | 0.373 |
| YOLOv11n | 0.813 | 0.779 | 0.843 | 0.386 |
| YOLOv12n | 0.820 | 0.757 | 0.834 | 0.374 |
| YOLOv13n | 0.825 | 0.749 | 0.826 | 0.372 |
| HSMD-YOLOn (Ours) | 0.828 | 0.785 | 0.854 | 0.398 |
| YOLOv8s | 0.841 | 0.805 | 0.872 | 0.415 |
| YOLOv9s | 0.841 | 0.807 | 0.871 | 0.414 |
| YOLOv10s | 0.803 | 0.777 | 0.843 | 0.401 |
| YOLOv11s | 0.827 | 0.806 | 0.859 | 0.411 |
| YOLOv12s | 0.847 | 0.778 | 0.858 | 0.406 |
| YOLOv13s | 0.839 | 0.774 | 0.853 | 0.397 |
| HSMD-YOLOs (Ours) | 0.850 | 0.810 | 0.880 | 0.422 |
| YOLOv8m | 0.857 | 0.802 | 0.874 | 0.426 |
| YOLOv9m | 0.842 | 0.809 | 0.869 | 0.417 |
| YOLOv10m | 0.828 | 0.768 | 0.849 | 0.412 |
| YOLOv11m | 0.824 | 0.804 | 0.899 | 0.425 |
| YOLOv12m | 0.842 | 0.806 | 0.866 | 0.419 |
| YOLOv13m | 0.834 | 0.812 | 0.874 | 0.417 |
| HSMD-YOLOm (Ours) | 0.843 | 0.824 | 0.891 | 0.430 |
| Model | SSB | GLRB | BEMAF | Precission | Recall | mAP@ | mAP@50:95 |
|---|---|---|---|---|---|---|---|
| exp1 | 0.904 | 0.860 | 0.893 | 0.672 | |||
| exp1 | ✓ | 0.909 | 0.867 | 0.898 | 0.685 | ||
| exp3 | ✓ | 0.901 | 0.865 | 0.897 | 0.692 | ||
| exp4 | ✓ | 0.910 | 0.866 | 0.894 | 0.676 | ||
| exp5 | ✓ | ✓ | 0.913 | 0.866 | 0.906 | 0.694 | |
| exp6 | ✓ | ✓ | 0.909 | 0.862 | 0.900 | 0.693 | |
| exp7 | ✓ | ✓ | 0.908 | 0.870 | 0.903 | 0.695 | |
| exp8 (Ours) | ✓ | ✓ | ✓ | 0.908 | 0.874 | 0.911 | 0.703 |
| NWDloss | Precission | Recall | mAP@ | mAP@50:95 |
|---|---|---|---|---|
| 0.2 | 0.915 | 0.872 | 0.903 | 0.692 |
| 0.3 | 0.916 | 0.867 | 0.907 | 0.702 |
| 0.4 | 0.908 | 0.874 | 0.911 | 0.703 |
| 0.5 | 0.905 | 0.871 | 0.905 | 0.699 |
| 0.6 | 0.912 | 0.867 | 0.907 | 0.703 |
| 0.7 | 0.907 | 0.866 | 0.892 | 0.677 |
| 0.8 | 0.907 | 0.868 | 0.909 | 0.698 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Luo, W.; Li, Y.; Zong, S. HSMD-YOLO: An Anti-Aliasing Feature-Enhanced Network for High-Speed Microbubble Detection. Algorithms 2026, 19, 234. https://doi.org/10.3390/a19030234
Luo W, Li Y, Zong S. HSMD-YOLO: An Anti-Aliasing Feature-Enhanced Network for High-Speed Microbubble Detection. Algorithms. 2026; 19(3):234. https://doi.org/10.3390/a19030234
Chicago/Turabian StyleLuo, Wenda, Yongjie Li, and Siguang Zong. 2026. "HSMD-YOLO: An Anti-Aliasing Feature-Enhanced Network for High-Speed Microbubble Detection" Algorithms 19, no. 3: 234. https://doi.org/10.3390/a19030234
APA StyleLuo, W., Li, Y., & Zong, S. (2026). HSMD-YOLO: An Anti-Aliasing Feature-Enhanced Network for High-Speed Microbubble Detection. Algorithms, 19(3), 234. https://doi.org/10.3390/a19030234

