Next Article in Journal
Solving the Multi-Skill Resource-Constrained Project Scheduling Problem Considering Skill Depth and Breadth Using ACO Algorithm and LINGO
Previous Article in Journal
Explainable AI for Conditional Stroke Probability Estimation: A Hybrid Bayesian Network Approach
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

MD-YOLO: An Improved YOLO26-Based Model for Small-Object Detection in UAV Aerial Imagery

1
School of Computer and Information Engineering, Shanghai Polytechnic University, Shanghai 201209, China
2
School of Electronics and Information Engineering, Tongji University, Shanghai 201804, China
3
Department of Engineering, Durham University, Durham DH1 3LE, UK
4
Ningbo Whale Flare Information Technology Co., Ltd., Ningbo 315000, China
*
Author to whom correspondence should be addressed.
Algorithms 2026, 19(9), 737; https://doi.org/10.3390/a19090737
Submission received: 29 July 2026 / Revised: 27 August 2026 / Accepted: 28 August 2026 / Published: 1 September 2026

Abstract

Object detection in UAV aerial imagery plays a vital role in applications such as traffic surveillance, urban management, and low-altitude inspection. However, aerial images typically present challenges including small object scales, dense distributions, severe occlusion, and cluttered backgrounds. Existing YOLO-series detectors still exhibit limitations in small-object feature representation, multi-scale contextual modeling, and downsampling detail preservation. To address these issues, this paper proposes MD-YOLO, an improved object detection model tailored for UAV scenarios, built upon the YOLO26 baseline. MD-YOLO incorporates three lightweight modules—IMO, DS-SPPF, and HPConv to optimize backbone feature extraction, multi-scale contextual aggregation, and Neck downsampling, respectively, thereby enhancing the model’s detection capability for small objects in complex UAV scenarios. Experimental results demonstrate that, compared with the baseline model, MD-YOLO achieves improvements of 3.9, 3.2, 2.9, and 3.1 percentage points in Precision, Recall, mAP50, and mAP50-95 on the VisDrone-2019 dataset, while maintaining a parameter count of only 9.3 M, thereby striking a favorable balance between accuracy and complexity. Independent evaluations on the UAVDT and NWPU VHR-10 datasets further confirm the consistent effectiveness of the model across diverse UAV imaging conditions, with mAP50 improvements of 8.3 and 1.4 percentage points, respectively.

1. Introduction

As one of the important research directions in computer vision, Unmanned Aerial Vehicle (UAV)-based aerial object detection has demonstrated broad application value in a variety of real-world scenarios, including intelligent transportation [1], urban security surveillance, emergency rescue, agricultural monitoring [2], and power-line inspection [3]. Compared with conventional ground-view images, UAV aerial imagery offers a wider field of view, more flexible viewing angles, and richer scene information, thereby providing effective data support for large-scale target monitoring and real-time intelligent perception [4]. Nevertheless, aerial images typically exhibit a pronounced top-down perspective, small object scales, dense object distributions, dramatic scale variations, and complex background clutter; these characteristics make UAV-based object detection substantially more challenging than object detection in ordinary natural scenes.
With the rapid development of deep learning-based object detection, representative detection frameworks such as Faster R-CNN [5,6], SSD [7], and the YOLO family [8,9,10,11,12,13], together with backbone networks such as ResNet [14], have been widely applied to aerial object detection tasks. Among them, YOLO detectors, with their end-to-end architecture, fast inference speed, and strong multi-scale detection capability, show considerable potential in real-time UAV detection scenarios. Existing studies typically improve aerial object detection performance by enhancing backbone networks [15], introducing attention mechanisms, fusing multi-scale features [16], and designing lightweight detector architectures [17], thereby alleviating missed detections of small objects and false positives caused by complex backgrounds to some extent.
Although existing CNN- and YOLO-based detectors for UAV aerial imagery have attained satisfactory real-time performance, several notable limitations remain. First, the backbones of mainstream detectors exploit shallow-level fine-grained textures and edge cues only to a limited extent and exhibit weak capability for modeling long-range global-context dependencies in deep features; consequently, when confronted with complex aerial scenes involving large-scale variations and dense occlusions, they are prone to missing small objects and producing insufficient feature representations. Second, conventional spatial pyramid pooling modules aggregate features through fixed receptive fields and are unable to adaptively adjust their fusion strategy according to the actual object size, thereby failing to simultaneously accommodate the feature-extraction requirements of both large and small objects. Third, in the bottom-up feature fusion path of the Neck, most YOLO-series detectors perform scale transformation with a single convolutional branch; although this design is structurally simple and computationally efficient, it tends to discard the edges, textures, and weak-response cues of small objects as the feature-map resolution decreases, so that shallow-level details cannot effectively propagate back to the mid- and deep-level features, ultimately undermining the detection stability for small objects in complex UAV scenarios. Finally, most improved algorithms achieve higher accuracy at the cost of a substantial increase in parameters and computational overhead, making it difficult to strike a balance among detection accuracy, inference speed, and lightweight-deployment requirements; moreover, these methods generalize poorly across various publicly available aerial benchmarks and perform unsatisfactorily on extremely small categories such as non-motorized vehicles.
To address the aforementioned limitations, the proposed MD-YOLO algorithm introduces targeted improvements at three levels: the backbone, the feature-aggregation module, and the Neck downsampling operator. First, an Improved MambaOut [18] backbone (IMO) is constructed to compensate for the deficiencies of the original network in small-object feature extraction and long-range dependency modeling. Second, a Dynamic Scale SPPF (DS-SPPF) module is designed, which leverages an adaptive gating mechanism to dynamically allocate multi-scale receptive fields, thereby flexibly accommodating the highly variable object scales encountered in aerial imagery. Finally, a Hybrid Pooling Convolution (HPConv) downsampling module is embedded in the Neck. By jointly modeling a pooling branch and a multi-scale convolutional branch, HPConv better preserves salient responses, edge and texture cues, and contextual information while performing spatial-scale compression, effectively mitigating the degradation of small-object features during downsampling. Together, these designs substantially enhance the multi-scale feature-fusion capability of the model in aerial scenarios that involve severe occlusions and cluttered backgrounds.
The overall architecture optimizes and refines the inherent lightweight design of YOLO26 [13], aiming to improve detection stability in scenarios involving dense occlusions and challenging illumination conditions at the cost of only a modest increase in parameters and computational overhead. We adopt YOLO26 as our baseline because it (i) provides the most recent unified end-to-end architecture in the Ultralytics YOLO series (2026), removing the NMS bottleneck of earlier versions; (ii) already achieves a strong accuracy–efficiency trade-off at the “s” scale; and (iii) benefits from Task-Aligned Learning and Distribution Focal Loss, which are well-suited to the extreme scale variations found in UAV aerial imagery. In summary, the proposed method effectively alleviates three core challenges in UAV aerial object detection—namely, the difficulty of recognizing small targets, the inefficiency of multi-scale feature fusion, and the trade-off between accuracy and lightweight deployment—and achieves significant performance gains on the VisDrone-2019 [19], UAVDT [20], and NWPU VHR-10 [21] public benchmarks, while simultaneously delivering competitive detection accuracy, stable convergence behavior, and good potential for practical deployment.
In summary, the main contributions of this paper are as follows:
  • A lightweight backbone-optimization strategy is proposed. For UAV aerial object detection, an Improved MambaOut backbone (IMO) is designed. A local detail-enhancement branch is introduced at the shallow layers, and a linear-attention-based adapter module is embedded at the deeper layers, so that the network simultaneously strengthens the extraction of fine-grained small-object features and the modeling of global contextual dependencies. This design substantially improves the feature-representation capability in complex aerial scenes while preserving favorable lightweight characteristics.
  • An adaptive multi-scale feature-aggregation module is developed. To overcome the limitation of the fixed-scale aggregation used in the conventional SPPF, a Dynamic Scale SPPF (DS-SPPF) module is constructed by integrating multiple parallel branches with channel-wise adaptive gating weights. Guided by the scale distribution of objects within the aerial scene, the module dynamically allocates the weights of different receptive fields and thereby enables flexible multi-scale spatial feature fusion, effectively boosting the detection performance for multi-scale objects in complex scenes with negligible additional computational overhead.
  • An HPConv downsampling module tailored for the feature-fusion stage is proposed. To mitigate the fine-grained information loss caused by vanilla convolutional downsampling in the Neck, a Hybrid Pooling Conv (HPConv) module is designed to replace the original downsampling convolution. Through the joint modeling of a pooling branch and a multi-scale convolutional branch, HPConv preserves edge textures and weak small-object responses, further enhancing the propagation of shallow-level details toward the mid- and deep-level features, thereby aiming to improve the detection robustness for small objects and occluded targets.
  • A high-performance UAV aerial small-object detection framework is built and comprehensively validated. Building upon YOLO26, the proposed improved backbone, the DS-SPPF module, and the HPConv downsampling module are integrated to form the MD-YOLO detection framework. Systematic experimental validation is conducted on three widely used UAV and remote-sensing benchmarks, namely VisDrone-2019, UAVDT, and NWPU VHR-10. The experimental results demonstrate that the proposed algorithm consistently outperforms the baseline YOLO26 model in terms of precision, recall, mAP50, and mAP50-95, with particularly notable improvements on tiny categories such as bicycles and tricycles, while simultaneously exhibiting stable training convergence, consistent effectiveness across multiple UAV datasets, and a favorable accuracy–efficiency trade-off.

2. Related Work

2.1. UAV/Aerial Object Detection

Aerial object detection from Unmanned Aerial Vehicles (UAVs) has become one of the important research directions within the object detection community and has been widely applied in scenarios such as traffic monitoring, urban security surveillance, emergency rescue, agricultural inspection, and low-altitude remote sensing. Compared with imagery captured in ordinary natural scenes, aerial images typically exhibit a predominantly top-down viewing perspective, pronounced object-scale variations, cluttered background textures, densely distributed objects, and frequent occlusions; these characteristics impose considerably higher demands on detection models in terms of object localization, small-object recognition, and complex-background suppression. In particular, under long-range imaging conditions, objects often occupy only a small number of pixels, and their edge, texture, and local structural cues are inherently weak; such cues are readily attenuated during deep feature extraction and successive downsampling operations, which in turn leads to missed detections, false detections, and inaccurate localization.
Researchers have explored a variety of strategies to address the detection challenges arising in UAV aerial scenarios. To cope with complex illumination, occlusion, and the weak response of small objects, Zheng et al. [22] incorporated an attention-enhancement mechanism together with a dedicated small-object detection branch, thereby strengthening the model’s perception of salient target regions. From the perspective of multi-scale feature fusion, Li et al. [23] enhanced the representation of densely distributed and blurred objects through finer-grained cross-scale interactions. Focusing on the trade-off between accuracy and efficiency, Li et al. [16] reduced redundant computations via structural optimization while simultaneously reinforcing the feature representations of small and densely distributed objects. Wang et al. [24] further refined cross-layer feature propagation by optimizing upsampling, downsampling, and multi-scale feature interactions, thereby alleviating the information imbalance between features of different resolutions. Approaching the problem from the perspective of inter-object and object–background relationship modeling, Moser et al. [25] introduced a relation-aware mechanism to enhance target discrimination under cluttered backgrounds. Zang et al. [26] incorporated open-vocabulary cues and semantic constraints into aerial object detection, offering a new perspective on generalized detection in complex open-world scenarios.
Overall, existing studies have gradually expanded from conventional detection-head refinements toward a broader spectrum of directions, including attention enhancement, multi-scale feature fusion, lightweight architecture design, relation modeling, and generalization to open-world scenarios. Although these methods have improved UAV aerial object detection performance to a certain extent, they still exhibit several notable weaknesses in scenarios in which extremely small objects, dense occlusions, and cluttered backgrounds coexist—namely, insufficient preservation of fine-grained information, inadequate modeling of global contextual dependencies, and difficulty in reconciling real-time efficiency with detection accuracy. Therefore, how to simultaneously enhance the local-detail representation of small objects and the contextual modeling capability in complex scenes, while preserving computational efficiency, remains a key open problem in UAV aerial object detection.

2.2. YOLO-Based Real-Time Detectors

Owing to its end-to-end architecture and high inference speed, the YOLO family has become one of the representative approaches for real-time object detection. Early YOLO models reformulated object detection as a single-stage regression problem, thereby achieving fast detection responses; subsequent versions have continuously evolved in terms of the backbone, feature-fusion structure, detection head, and training strategy, progressively striking a favorable balance between detection accuracy and real-time efficiency. Specifically, YOLOv8 [8] further refined feature representation and fusion within its convolutional modules and Neck, and incorporated a decoupled detection head together with Task-Aligned Learning (TAL) to boost detection performance. YOLO11 [10] further optimized the backbone and Neck upon its predecessors, strengthening feature extraction while achieving a more favorable speed–accuracy trade-off with a compact parameter footprint. YOLOv12 [11] adopts attention as its central design principle, introducing efficient area attention and an improved feature-aggregation structure so that attention-based models can jointly satisfy the requirements of global context modeling and inference efficiency in real-time detection. YOLOv13 [12] incorporates an improved multi-scale feature-fusion Neck together with an optimized attention-balancing mechanism, substantially enhancing detection robustness under cluttered backgrounds while preserving efficient real-time inference. Unlike YOLOv8 through YOLOv13, which progressively refine the backbone, Neck, and detection head but still rely on Non-Maximum Suppression (NMS) post-processing and Distribution Focal Loss (DFL), YOLO26 [13] departs from this paradigm at the architectural level. It is designed with a stronger focus on edge deployment and low-power devices: it employs native end-to-end NMS-free inference, removes the Distribution Focal Loss (DFL) to simplify hardware adaptation, and integrates training-optimization strategies such as ProgLoss, STAL, and MuSGD to comprehensively improve convergence stability, small-object detection capability, and practical deployment efficiency. In parallel with architectural improvements, the practical deployment of lightweight YOLO detectors on resource-constrained UAV platforms has recently received systematic attention. Bakirci [27] benchmarked YOLO-series detectors on a Jetson Xavier-equipped drone platform, analyzing the joint trade-offs among detection accuracy, inference speed, and on-board resource utilization. Rey et al. [28] further extended this evaluation to several edge platforms such as Jetson Orin NX, and reported that INT8 post-training quantization can substantially improve the on-device inference throughput of YOLOv8-series detectors without significant accuracy degradation.

2.3. Efficient Feature Modeling and Fusion

Efficient feature modeling has long been an important research direction for lightweight object detection networks. Howard et al. [15] employed depthwise separable convolutions to reduce the computational cost of convolutional operations, enabling lightweight networks to retain competitive visual feature-extraction capability in mobile and embedded scenarios. Liu et al. [29] further demonstrated that convolutional networks equipped with carefully redesigned structures still possess strong local modeling and hierarchical feature-representation capabilities, and can therefore serve as competent backbones for dense-prediction tasks such as detection and segmentation. To strengthen the feature-selection capability of convolutional structures, Rao et al. [30] introduced the notion of gated convolution, realizing high-order spatial interactions through a gating mechanism and thereby enhancing spatial-feature modeling at a modest computational cost. From the perspective of vision backbones, Yu et al. [18] further verified the effectiveness of gated-convolution-style structures on image-recognition tasks, showing that competitive visual representations can be obtained without relying on complex sequence-modeling modules. Convolutional and gated-convolutional operators, however, place greater emphasis on local spatial modeling, and their global-context modeling capability remains limited when confronted with densely distributed objects, occluded regions, and cluttered backgrounds in UAV aerial imagery. To alleviate this limitation, Katharopoulos et al. [31] reduced the long-sequence modeling complexity of self-attention through linear attention, whereas Su et al. [32] and Dong et al. [33] refined the representation of positional information from the perspectives of rotary position encoding and locally enhanced positional cues, respectively, offering new insights for lightweight global-context modeling.
Multi-scale context modeling is an important means of enhancing the scale-adaptation capability of object detection models. To address the limited scale-representation ability caused by the fixed pooling structure of the conventional SPPF, existing studies have introduced improvements from several perspectives. Wu et al. [34] integrated the multi-scale feature-extraction ability of SPP with the feature-fusion idea of CSP as a substitute for the original SPPF, aiming to improve contextual perception in complex traffic scenes. Yue et al. [17] replaced the SPPF with a multi-scale feature-aggregation module and, in combination with a dynamic scale-information fusion strategy, enhanced the model’s adaptability to weak small objects and cluttered backgrounds. Together, these studies indicate that existing improvements around the SPPF concentrate primarily on enlarging the receptive field, incorporating attention mechanisms, strengthening multi-scale fusion, and suppressing background clutter. Nevertheless, most of these designs still rely on relatively fixed structures and lack finer-grained adaptive modulation of the importance of different scale branches across spatial positions and channels.
Beyond multi-scale context aggregation, the preservation of fine details during downsampling also has a direct bearing on object detection performance. Prior studies have shown that although ordinary pooling and strided convolution can rapidly reduce the resolution of feature maps, they may simultaneously introduce aliasing artifacts, the loss of discriminative details, and the attenuation of small-object responses. Zhang [35] pointed out that downsampling operations without effective anti-aliasing treatment tend to compromise feature stability. Gao et al. [36] further argued that average pooling, max pooling, and strided convolution may overlook locally discriminative regions when compressing spatial information, and thus important features should be adaptively preserved according to the input content. Sunkara et al. [37] analyzed the information loss induced by strided convolution and pooling from the perspectives of low-resolution imagery and small-object detection, emphasizing that the disruption of fine-grained features during downsampling should be minimized. Taken together, these studies indicate that downsampling modules are responsible not only for scale conversion but also directly govern whether shallow-level details can be effectively propagated toward the mid- and deep-level features. Motivated by these observations, this paper proposes MD-YOLO, an improved object detection model tailored for UAV aerial scenarios that achieves a well-balanced trade-off between detection accuracy and inference efficiency.

3. Methods

YOLO26 is the most recent object detection model in the YOLO family and demonstrates superior overall performance compared with its predecessors; nevertheless, it still leaves room for improvement in the context of UAV aerial object detection. Given the characteristics of aerial imagery—small object scales, cluttered backgrounds, and frequent occlusions—the original YOLO26 remains limited in fine-grained feature extraction, multi-scale context modeling, and detail preservation during downsampling. To further strengthen the ability of the model to detect small and background-cluttered objects, this paper introduces three key modifications on top of YOLO26. First, an Improved MambaOut backbone (IMO) is adopted to reinforce the extraction of multi-level features. Second, a Dynamic Scale SPPF (DS-SPPF) module is designed to achieve dynamic multi-scale context enhancement. Third, an HPConv module is embedded in the downsampling path of the Neck to mitigate the loss of shallow-level small-object information during scale transformation. Working in concert, these enhancement modules jointly strengthen the feature-representation capability and detection accuracy of the model while maintaining a low computational cost. The overall architecture is illustrated in Figure 1.

3.1. Improved MambaOut Backbone Network

To strengthen the ability of the backbone to represent small objects in UAV aerial imagery, an improved variant of MambaOut, denoted as IMO, is introduced as the backbone of YOLO26. MambaOut employs the GatedCNN Block as its basic feature-extraction unit, which selects and modulates the input features through a gating mechanism and thereby achieves high computational efficiency. However, since the original GatedCNN Block relies mainly on the gating mechanism to perform feature selection and channel-wise interaction, it exploits the edge, texture, and local-detail cues of small objects only to a limited extent in aerial scenarios; moreover, such cues are prone to being attenuated during deep feature transformations.
To address this issue, a 3 × 3 depthwise separable convolutional branch is added in parallel within the GatedCNN Block to supplement the local-neighborhood spatial information and to strengthen the network’s perception of small-object contours, boundary textures, and fine-grained structures. Unlike depthwise separable convolutions in MobileNets, lightweight linear operations in Ghost modules, or spatial reorganization in SPD-Conv, our local-enhancement branch preserves the spatial resolution of the input while actively strengthening high-frequency local signals via parallel 3 × 3 depthwise convolutions. This design avoids the excessive attenuation of small-object edges and textures that typically occurs during channel compression or spatial rearrangement, and therefore allows IMO to more reliably retain the fine-grained responses critical for small-object detection under the densely distributed and cluttered-background conditions common in UAV aerial imagery. The improved structure is illustrated in Figure 2. For this 3 × 3 depthwise separable convolutional branch, given an input feature X , the local-enhancement branch is formulated as in Equation (1):
F l o c a l = D W C o n v 3 × 3 X
where F l o c a l denotes the output feature of the local-enhancement branch, and D W C o n v 3 × 3 denotes the depthwise convolution operation with a 3 × 3 kernel. Since this branch performs local spatial modeling independently within each channel, it can extract high-frequency detail cues, such as edges and textures, at a low computational cost.
In parallel, the input feature X is still processed by the original GatedCNN main branch to undergo the gated feature transformation. Finally, the local-enhancement branch, the gated main branch, and the original residual feature are fused together to produce the output of the improved GatedCNN Block, as formulated in Equation (2):
Y = X + F g a t e X + F l o c a l
where F g a t e X denotes the output feature of the GatedCNN main branch, and Y denotes the final output feature of the improved GatedCNN Block. With this design, the improved GatedCNN Block preserves the original gated feature-selection capability while additionally incorporating a local-detail branch, endowing the IMO backbone with a stronger small-object feature-extraction capability at every stage. Compared with directly stacking plain convolutions, this structure enhances local representation while keeping the parameter count and computational complexity low, making it particularly suitable for small-object detection in UAV aerial scenarios.
As the backbone progressively deepens, the resolution of the feature maps decreases accordingly. Although deep features possess stronger semantic representation capability, their ability to model long-range spatial dependencies remains limited, owing to the inherently local nature of convolutional structures. In UAV aerial scenarios, however, small objects in complex UAV scenarios and cluttered background regions typically require broader contextual information for reliable discrimination. To strengthen the global perception capability of deep features, inspired by Han et al. [38], a MILA (Mamba-Inspired Linear Attention) Adapter is introduced after the last two deep stages of the IMO, enabling the network to supplement the model’s long-range spatial dependencies at a low computational cost. The structure of the adapter is illustrated in Figure 3. Given an input feature X ∈ R B × H × W × C , the MILA Adapter first applies LayerNorm and Flatten sequentially, mapping the two-dimensional feature into a sequential representation, as formulated in Equation (3):
X s = F l a t t e n L N X
where L N denotes the LayerNorm operation, and F l a t t e n denotes the operation that unfolds the spatial dimensions H × W into a feature sequence of length N = H × W . The resulting sequence is then fed into a linear attention module that integrates Rotary Position Embedding (RoPE), Locally Enhanced Positional Encoding (LePE), and an ELU+1 feature mapping, thereby modeling spatial-context relations while reducing the computational complexity of self-attention. The attention output is subsequently reshaped back to its original spatial form and modulated by a learnable scaling factor γ , which controls its contribution to the original feature and thus preserves training stability. The overall computation is formulated as in Equation (4):
F m i l a = γ ⋅ R e s h a p e L A X s
where L A denotes the linear attention operation. Compared with standard self-attention, linear attention substantially reduces the computational burden on high-resolution feature maps and is therefore better aligned with the lightweight-deployment requirements of UAV detection tasks. Finally, the MILA Adapter combines the enhanced contextual feature with the original input through a residual connection. Benefiting from this design, the MILA Adapter preserves the stability of the original features while effectively strengthening the capacity of deep features to model small objects amid cluttered backgrounds. Distinct from existing MambaOut variants that primarily target standard image recognition, our IMO integrates a resolution-preserving local-enhancement branch parallel to the gated convolution and a lightweight MILA adapter, explicitly aimed at retaining small-object cues that are typically attenuated by pure gating or attention backbones in cluttered UAV scenes.

3.2. Dynamic Scale SPPF Module

The conventional SPPF module mainly relies on pooling operations at fixed scales to enlarge the receptive field, which enhances the contextual representation capability of deep features to some extent. However, UAV aerial imagery exhibits pronounced object-scale variations, and a single image may simultaneously contain distant small objects, nearby large objects, and densely distributed objects; the feature-aggregation strategy adopted by fixed-receptive-field pooling is unable to adaptively select an appropriate contextual scope for different object regions, and consequently tends to yield insufficient detail responses for small objects or inadequate contextual representations for large-scale objects. To this end, a Dynamic Scale SPPF (DS-SPPF) module is designed to replace the original SPPF with a dynamic multi-scale feature-aggregation scheme; its structure is illustrated in Figure 4.
Given an input feature X , the DS-SPPF first performs channel compression via a 1 × 1 convolution to obtain a hidden feature F. Based on F, the module then constructs multiple parallel scale branches, namely the identity branch F 0 , the max-pooling branch F 1 , and two depthwise convolutional branches F 2 and F 3 with different dilation rates. The computation of each branch is formulated as in Equation (5):
F 0 = F ,   F 1 = M a x P o o l F ,   F 2 = D W C o n v 3 × 3 d = 2 F ,   F 3 = D W C o n v 3 × 3 d = 3 F
where F 0 preserves the original deep-feature information, F 1 highlights locally salient responses, and F 2 and F 3 enlarge the receptive field with different dilation rates to capture contextual information at diverse spatial ranges. Compared with a single pooling path, this multi-branch structure simultaneously accommodates fine local details and long-range semantic context.
To prevent the multiple scale branches from being fused under a fixed strategy, DS-SPPF further introduces a dynamic weight-generation mechanism. The weight generator consists of a 3 × 3 convolution, a SiLU activation function, and a 1 × 1 convolution cascaded sequentially, which adaptively produces a weight score for each branch conditioned on the input feature F , as formulated in Equation (6):
W = S o f t m a x R e s h a p e C o n v 1 × 1 S i L U C o n v 3 × 3 F
where W i denotes the dynamic weight corresponding to the i -th branch, enabling per-location adaptive scale selection as reflected in the ablation gains reported in Section 4; a direct visualization of these weights is left for future work. This weight is both spatially adaptive and channel-wise, enabling the model to allocate the importance of each scale branch adaptively across different spatial locations and channels. Finally, DS-SPPF fuses the branch features by performing element-wise multiplication with the corresponding weights followed by concatenation, as formulated in Equation (7):
Y = C o n v 1 × 1 C o n c a t W 0 F 0 , W 1 F 1 , W 2 F 2 , W 3 F 3
where Y denotes the output feature of DS-SPPF. Through the above design, DS-SPPF dynamically adjusts the contributions of the multi-scale branches according to the scale and contextual requirements of different target regions, enabling the deep features to simultaneously capture local saliency and long-range contextual representations. This improves the model’s adaptability to challenging scenarios in UAV aerial imagery, including targets with drastic scale variations and objects embedded in cluttered backgrounds. Compared with existing adaptive SPP variants that produce only channel-wise weights, DS-SPPF generates spatially- and channel-adaptive weights via a lightweight two-layer generator, enabling per-pixel scale selection critical for the drastic scale variations found in UAV imagery.

3.3. Hybrid Pooling Conv Module

In the bottom-up feature fusion path of the Neck in the YOLO series, plain 3 × 3 stride-2 convolutions are conventionally used to perform downsampling between features of different scales. Although this operation is structurally simple and computationally efficient, its single-path compression tends to cause a loss of shallow-layer fine-grained information. This issue is particularly pronounced in UAV aerial imagery, where the edges, textures, and weak-response features of small objects are further attenuated as the resolution is reduced. To address this, we design the Hybrid Pooling Conv (HPConv) module to replace the plain downsampling convolution in the Neck, allowing shallow-layer features to retain more salient responses and contextual information when propagated to the middle and deep layers. In contrast to anti-aliased pooling and SPD-Conv that adopt a single detail-preservation strategy, HPConv jointly combines average and max pooling in parallel with multi-scale large-kernel convolutions, simultaneously preserving aggregate regional response, locally salient features, and enlarged receptive-field context without altering channel dimensionality. Its structure is illustrated in Figure 5.
Given an input feature X , HPConv first splits it along the channel dimension into two sub-features X 1 and X 2 , which correspond to two distinct downsampling modeling branches. For X 1 , the module first applies average pooling to preserve the overall regional response, and then further divides it into two sub-branches that extract local contextual information through convolutions of different kernel sizes, as formulated in Equations (8)–(10):
X 3 , X 4 = S p l i t A v g P o o l X 1
F 1 = C o n v 7 × 7 s = 2 X 3
F 2 = C o n v 5 × 5 s = 2 X 4
where C o n v 7 × 7 s = 2 and C o n v 5 × 5 s = 2 denote large-kernel convolutions with a stride of 2, respectively. Larger convolutional kernels enlarge the local receptive field during downsampling, allowing the module to preserve more contextual information around the targets while reducing the spatial resolution of the feature map.
For the other branch X 2 , HPConv adopts max pooling to highlight locally salient responses, and then employs a 1 × 1 convolution to perform channel projection and feature integration, as formulated in Equation (11):
F 3 = C o n v 1 × 1 M a x P o o l 3 × 3 s = 2 X 2
Max pooling reinforces locally high-response regions and helps preserve the edges and salient textures of small objects. Finally, HPConv first concatenates the outputs of the two convolutional branches on the left, and then concatenates the resulting feature with the max-pooling branch on the right to yield the final downsampled output.
Through the collaborative modeling of average pooling, max pooling, and multi-scale convolutional branches, HPConv achieves spatial-scale compression while simultaneously accounting for the overall regional response, locally salient responses, and contextual representation. We apply HPConv at the P3 → P4 and P4 → P5 scale-transition positions in the Neck, allowing the fine-grained details of small objects in shallow high-resolution features to be propagated more effectively into the middle and deep detection branches. This helps improve the detection accuracy and robustness of the model for small objects in complex UAV scenarios and objects in cluttered backgrounds.

4. Experiments and Analysis

4.1. Datasets

We conduct experiments on three publicly available UAV aerial datasets: VisDrone-2019 [19], UAVDT [20], and NWPU VHR-10 [21]. VisDrone-2019, released by the AISKYEYE team at Tianjin University, is a widely used benchmark in the field of UAV object detection. We strictly follow the official split released by the AISKYEYE team: 6471 images for training, 548 images for validation, and 1610 images for the test set, covering 10 object categories (pedestrian, people, bicycle, car, van, truck, tricycle, awning-tricycle, bus, and motor). The dataset is characterized by small object scales, dense object distributions, severe occlusion, cluttered backgrounds, and drastic viewpoint variations, which faithfully reproduce the key challenges encountered in real-world UAV scenarios. Accordingly, we adopt VisDrone-2019 as the primary experimental dataset to evaluate the detection performance of the proposed MD-YOLO under conditions involving small objects, densely distributed objects, and cluttered backgrounds.
The UAVDT and NWPU VHR-10 datasets are further employed to analyze the robustness of the proposed model on vehicle detection and high-resolution remote-sensing detection tasks. UAVDT is primarily designed for UAV-based traffic surveillance scenarios and consists of vehicle targets captured under various flight altitudes, viewing angles, illumination conditions, and occlusion levels, which comprehensively reflect the detection performance of the model under conditions of densely distributed targets, small object scales, and cluttered backgrounds. We adopt the official train/test split of 30 training sequences (approximately 24,000 frames) and 20 testing sequences (approximately 16,000 frames), covering 3 vehicle categories (car, truck, bus). NWPU VHR-10, released by Northwestern Polytechnical University, is composed of high-resolution remote-sensing images and covers typical ground-object categories such as airplanes, ships, storage tanks, bridges, and vehicles. For NWPU VHR-10, which contains 650 positive images across 10 categories, we adopt a stratified 70%/15%/15% split (455/97/98 images) fixed by a random seed of 42. This dataset exhibits pronounced scale variations and strong background interference, and is therefore suitable for verifying the robustness of the proposed model on multi-category detection tasks in remote-sensing imagery.

4.2. Experimental Settings

All experiments in this study are implemented on the PyTorch deep learning framework. PyTorch is an open-source deep learning platform that has been widely adopted in fields such as computer vision and natural language processing owing to its flexibility and ease of use. All experiments are conducted on the Ubuntu 22.04 operating system. The hardware configuration comprises an Intel(R) Core(TM) i9-14900K CPU(Intel Corporation, Santa Clara, CA, USA), 64 GB of RAM, and an NVIDIA GeForce RTX 4090 GPU(NVIDIA Corporation, Santa Clara, CA, USA) with 24 GB of memory. The software environment consists of Python 3.8, PyTorch 2.4.1, and CUDA 11.8.
During training, all input images are uniformly resized to 640 × 640 pixels. The model is trained for 250 epochs with a batch size of 6. For optimization, the initial learning rate lr0 is set to 0.005, while the final learning rate is set to 0.01, and a cosine learning-rate schedule is enabled. To stabilize the early stage of training, a warm-up strategy with warmup_epochs = 3.0 is adopted. During the training process, no pretrained weights were used. The momentum and weight decay are set to 0.937 and 0.0005, respectively, in order to stabilize parameter updates and mitigate overfitting. Mosaic augmentation is disabled during the last 15 epochs to stabilize final convergence. Each experiment was independently repeated three times with different random initializations, and the reported results are the mean over the three runs; the maximum observed standard deviation across all metrics was below 0.3 percentage points. Through the appropriate configuration of these hyperparameters, the model achieves stable convergence, higher detection accuracy, and better generalization capability on the object detection tasks. A more detailed summary of the hyperparameter settings is provided in Table 1.

4.3. Evaluation Metrics

We adopt Precision, Recall, mAP50, and mAP50-95 as the primary accuracy metrics, and further employ the number of parameters (Params) and Giga Floating-Point Operations (GFLOPs) to analyze the complexity and computational cost of the model. Precision and Recall measure the correctness of the detection results and the coverage of ground-truth targets, respectively, and are formulated as in Equations (12) and (13):
P r e c i s i o n = T P T P + F P
R e c a l l = T P T P + F N
where T P denotes the number of correctly detected targets, F P denotes the number of background regions or objects from other categories that are incorrectly identified as the target, and F N denotes the number of ground-truth targets that are missed by the model. A higher Precision indicates fewer false positives produced by the model, whereas a higher Recall indicates fewer missed detections of ground-truth targets.
Average Precision (AP) measures the overall detection performance of the model on a single category. It is essentially defined as the area under the Precision–Recall curve computed across different confidence thresholds, and can be expressed as Equation (14). For multi-category object detection tasks, the mean Average Precision (mAP) is defined as the average AP across all categories, as formulated in Equations (14) and (15):
A P = ∫ 0 1 P R d R
m A P = 1 N ∑ i = 1 N A P i
where N denotes the total number of object categories, and A P ( i ) denotes the Average Precision of the i -th category. In this study, mAP50 and mAP50-95 are adopted as the two principal accuracy indicators. Specifically, mAP50 refers to the mean Average Precision computed at an IoU threshold of 0.5, which reflects the detection capability of the model under a relatively loose localization criterion. In contrast, mAP50-95 represents the average of the mAP values computed at IoU thresholds ranging from 0.50 to 0.95 with a step size of 0.05. It imposes stricter requirements on the localization accuracy of the predicted bounding boxes and therefore provides a more rigorous and comprehensive evaluation of the overall detection performance of the model.
In addition to detection accuracy, we further employ the number of parameters (Params) and Giga Floating-Point Operations (GFLOPs) to evaluate the complexity of the model. Params denotes the total number of model parameters, typically measured in millions (M), and is used to characterize the model’s storage overhead and structural scale. GFLOPs denotes the number of floating-point operations, in billions, required for one forward inference pass, and is used to reflect the computational complexity of the model. For UAV platforms, a detection model is expected not only to achieve high recognition accuracy but also to maintain a low parameter count and computational cost, so as to satisfy the practical requirements of edge-device deployment and real-time detection.

4.4. Comparative Experimental Results

To verify the effectiveness of the proposed MD-YOLO model, we compare it with several representative object detectors under identical experimental settings, including the YOLO-series detectors YOLOv8s, YOLOv10s, YOLOv11s, YOLOv12s, YOLOv13s, and YOLO26s, as well as the Transformer-based end-to-end detector RT-DETR-R18, as summarized in Table 2. Specifically, the YOLO-series detectors represent the mainstream single-stage anchor-free paradigm, while RT-DETR-R18 serves as a representative end-to-end Transformer-based detector, which is included to validate the performance advantage of MD-YOLO under a cross-paradigm comparison. All models are evaluated using a consistent set of metrics, including precision, recall, mAP50, mAP50–95, parameter count (Params), and inference speed (FPS). FPS is measured on a single RTX 4090 with batch size 1 at 640 × 640 input in FP32, averaged over 500 forward passes after 50 warmup runs.
As shown in Table 2, MD-YOLO achieves the best overall detection performance on the VisDrone-2019 test set. Specifically, MD-YOLO reaches 49.9% Precision, 39.3% Recall, 39.8% mAP50, and 23.1% mAP50-95, yielding notable gains of 3.9, 3.2, 2.9, and 3.1 percentage points over the YOLO26 baseline, respectively. Moreover, in the cross-paradigm comparison with the end-to-end Transformer-based detector RT-DETR-R18, MD-YOLO surpasses it by 3.6 and 2.4 percentage points in mAP50 and mAP50-95, respectively, while using only 46.5% of its parameters and 39.1% of its computational cost, fully demonstrating the dual advantage of the proposed method in both accuracy and efficiency. To verify that MD-YOLO improves small-object detection rather than merely raising overall accuracy, we further evaluate the method on VisDrone-2019 using the standard COCO-style size stratification, in which detections are grouped into small (area < 322 px2), medium, and large categories. As shown in Table 3, MD-YOLO achieves an AP_S of 16.8%, representing a 3.4 percentage-point improvement over the YOLO26 baseline. The gains on AP_M and AP_L are 2.7 and 3.6 percentage points, respectively, indicating that MD-YOLO improves detection performance across objects of different sizes. These results demonstrate that the proposed method not only enhances the classification and detection capability of the model but also maintains reliable localization accuracy under stricter IoU thresholds, thereby supporting the effectiveness of the lightweight architectural innovations introduced in this work. Compared with other YOLO variants and the recently proposed YOLO26, MD-YOLO outperforms all YOLO-series baselines and the RT-DETR-R18 detector across all accuracy-related metrics, further verifying its robustness in small-object detection on UAV aerial imagery.
To further analyze the detection performance of the model on different object categories, we adopt Precision–Recall (PR) curves to visually compare the per-class detection results of the baseline YOLO26 and the proposed MD-YOLO on the VisDrone-2019 dataset, as illustrated in Figure 6 and Figure 7. The PR curve jointly characterizes the variation of Precision and Recall across different confidence thresholds. The closer the curve approaches the upper-right corner, the higher the Precision that the model can simultaneously maintain at a high Recall level, and thus the better the detection performance for the corresponding category.
As shown in Figure 6 and Figure 7, YOLO26 achieves an overall mAP50 of 0.369 across all categories, whereas the proposed MD-YOLO improves this value to 0.398, corresponding to an overall gain of 2.9 percentage points. In terms of per-class performance, MD-YOLO achieves higher AP values on the majority of categories in the dataset: the AP of the bus category is improved from 0.511 to 0.584, the motor category from 0.420 to 0.463, and the tricycle category from 0.243 to 0.282. These results demonstrate that MD-YOLO achieves consistent improvements across most categories on VisDrone-2019. Meanwhile, the car category maintains a consistently high detection accuracy across both models, and MD-YOLO reaches an AP of 0.801, indicating that the model also exhibits stable and superior detection capability for categories with abundant training samples and relatively stable geometric characteristics.
From the overall trend of the PR curves, the averaged curve of MD-YOLO lies consistently above that of YOLO26, and particularly maintains a higher Precision level in the medium-to-high Recall regime. This indicates that the improved model effectively reduces false-positive risk while simultaneously recalling more ground-truth targets. The observed performance gains are primarily attributed to three factors: the joint modeling of local details and global contextual information provided by the IMO backbone, the dynamic aggregation of multi-scale contextual features achieved by DS-SPPF, and the effective mitigation of downsampling information loss offered by HPConv during feature fusion. In summary, the PR-curve results confirm that MD-YOLO delivers superior overall detection performance on the multi-category UAV object detection task of VisDrone-2019.

4.5. Ablation Experiments

We conduct ablation experiments on the VisDrone-2019 dataset with YOLO26 as the baseline model. The ablation study adopts a full configuration-matrix design: starting from the YOLO26 baseline, we enumerate all combinations of the three proposed modules (IMO, DS-SPPF, and HPConv), yielding eight configurations in total, while keeping the training strategy, input size, data augmentation pipeline, and evaluation metrics consistent across all of them. This controlled-variable approach enables a clear analysis of the contribution of each module to detection accuracy, model complexity, and computational cost. The experimental results are shown in Table 4.
As shown in Table 4, each module improves over the baseline when added individually (Exp2–Exp4), every pairwise combination (Exp5–Exp7) further outperforms the corresponding single-module variants, and the full model (Exp8) attains the best result on all accuracy metrics. Compared with the baseline YOLO26, the final MD-YOLO model improves Precision from 46.0% to 49.9%, Recall from 36.1% to 39.3%, mAP50 from 36.9% to 39.8%, and mAP50-95 from 20.0% to 23.1%, corresponding to gains of 2.9 and 3.1 percentage points in mAP50 and mAP50-95, respectively. Meanwhile, the number of model parameters decreases slightly from 9.4 M to 9.3 M, indicating that the proposed method enhances detection accuracy without noticeably increasing model complexity, thereby demonstrating the efficiency of the improved architecture. As shown in Table 5, the progressive ablation shows that MILA contributes +0.7 mAP50 and Local Enhancement contributes +0.5 mAP50 on top of the previous configuration, confirming that MILA and the local-enhancement branch contribute complementarily to IMO’s overall effectiveness.
Examining the single-module configurations, IMO alone (Exp2) provides the largest gain, raising mAP50/mAP50-95 to 38.0%/22.3%, which confirms that the improved backbone is the dominant contributor. DS-SPPF alone (Exp3) and HPConv alone (Exp4) each improve mAP50 over the baseline (37.2% and 37.6% vs. 36.9%), indicating that both modules are independently beneficial. Among the pairwise combinations, IMO+DS-SPPF (Exp5) further lifts Precision, Recall, and mAP50 to 49.2%, 38.0%, and 38.4%, while mAP50-95 shows a marginal change from 22.3% to 22.0%. Mechanistically, DS-SPPF mainly strengthens multi-scale context aggregation, which benefits recall and coarse-IoU detection but offers little additional gain for precise boundary regression under stricter IoU thresholds. Importantly, this marginal fluctuation is fully compensated once HPConv is added: the complete model (Exp8) attains the highest mAP50-95 of 23.1%, confirming that the three modules are complementary rather than conflicting. As shown in Table 6, HPConv achieves the best mAP50/mAP50-95 among the evaluated detail-preservation strategies without increasing channel dimensionality, confirming that its hybrid pooling design is well-matched to small-object downsampling in UAV imagery.
In summary, IMO, DS-SPPF, and HPConv improve model performance along three complementary dimensions: backbone feature extraction, multi-scale contextual aggregation, and downsampling detail preservation. Specifically, IMO enhances the capture of local details and the expressiveness of deep semantic representations; DS-SPPF improves the model’s adaptability to targets with drastic scale variations; and HPConv effectively mitigates the loss of small-object information during downsampling. When combined, the three modules enable the model to achieve optimal detection accuracy while maintaining a relatively low parameter count, thereby demonstrating the strong complementarity and effectiveness of the proposed architectural improvements.
Figure 8 illustrates the training dynamics of MD-YOLO, including the training/validation losses, Precision, Recall, mAP50, and mAP50-95 across epochs. Both the training and validation losses decrease steadily and remain in close alignment without divergence, indicating stable optimization and no divergence between training and validation behavior. Meanwhile, Precision, Recall, mAP50, and mAP50-95 rise rapidly and stabilize in the later epochs, demonstrating that the model achieves effective and stable convergence.

4.6. Visualization Analysis

To qualitatively illustrate the effectiveness of the proposed method, we select images featuring densely distributed targets and cluttered backgrounds under diverse UAV scenarios from the experimental dataset. Figure 9 presents a qualitative comparison of the detection results produced by YOLO26 and MD-YOLO under various UAV scenes. Specifically, Figure 9a,b illustrate a dense-pedestrian scenario in a plaza. The baseline YOLO26 (Figure 9a) exhibits noticeable missed detections in distant high-density crowds and completely fails to identify the bicycles interspersed among the pedestrians. In contrast, MD-YOLO (Figure 9b) not only detects a greater number of small-scale pedestrians in the far-field region but also successfully identifies multiple bicycles concealed within the crowd, demonstrating its stronger fine-grained discrimination capability under high-density and heavily overlapping target scenarios. Figure 9c,d present a nighttime low-illumination urban traffic scene captured from an aerial perspective. YOLO26 (Figure 9c) shows evident missed detections in building-shadowed regions and along the curved road edges, and it fails to identify the truck target in the scene. In contrast, MD-YOLO (Figure 9d) detects additional small-scale vehicles in shadow-covered road areas and successfully recognizes the occluded truck target, further verifying the advantages of the proposed method in detecting small and occluded objects under low-illumination and cluttered-background conditions.
Overall, MD-YOLO shows more stable detections across these examples in complex scenarios. In densely populated traffic scenes, YOLO26 frequently misses distant small-scale vehicles and occluded pedestrians, whereas MD-YOLO successfully detects a greater number of small targets, substantially reducing the incidence of missed detections. In occlusion scenarios, where targets retain only partially visible regions, YOLO26 often produces unstable bounding boxes and occasionally fails to localize the targets entirely. In the illustrated scenes, MD-YOLO produces additional correct detections that YOLO26 misses, including some partially occluded and distant small targets.
Figure 10 presents a comparison of feature-activation heatmaps between YOLO26 and MD-YOLO in traffic scenes. The heatmaps in Figure 10 are computed by aggregating Grad-CAM responses over the P3/P4/P5 multi-scale features fed into the detection head, with the sum of class scores and bounding-box coordinates of the top-2% confident detections used as the backward target. As shown in the figure, the response regions of YOLO26 are relatively scattered under cluttered backgrounds and are susceptible to interference from roads, building shadows, and background textures, indicating that the model’s attention is not sufficiently concentrated on the true target regions. In contrast, MD-YOLO exhibits more focused high-response regions around vehicles, pedestrians, and tricycles, particularly in target-dense regions, where the model produces more continuous and explicit activation responses across multiple targets. These heatmaps qualitatively suggest that MD-YOLO tends to focus more on target regions and is less affected by background clutter.

4.7. Evaluation on Additional UAV Datasets

To further assess the robustness and applicability of the MD-YOLO architecture across diverse UAV imaging scenarios, we conduct independent training and evaluation on two additional public benchmarks: UAVDT and NWPU VHR-10. For each dataset, both the baseline YOLO26 and MD-YOLO are trained from scratch using identical hyperparameter configurations as those adopted on VisDrone-2019 (Table 1), and evaluated on the corresponding official test split. The overall results are summarized in Table 7.
On the UAVDT dataset, MD-YOLO maintains its detection advantage under an independent training setting. Compared with YOLO26, MD-YOLO achieves gains of 5.2, 8.6, 8.3, and 3.7 percentage points in Precision, Recall, mAP50, and mAP50-95, respectively. The UAVDT dataset primarily consists of vehicle targets captured from UAV perspectives and exhibits challenges such as drastic scale variations, motion blur, illumination changes, and occlusion. The higher Recall achieved by MD-YOLO on this dataset indicates that the model successfully detects a greater number of ground-truth targets, thereby substantially alleviating the issue of missed detections. Meanwhile, the simultaneous improvements in mAP50 and mAP50-95 demonstrate that the model enhances both the classification and localization accuracy of vehicle targets.
On the NWPU VHR-10 dataset, MD-YOLO achieves higher values than YOLO26 for Recall, mAP50, and mAP50-95, further validating the effectiveness of the proposed method. These results, obtained through independent training on datasets with different imaging characteristics (low-altitude UAV surveillance in UAVDT and high-altitude high-resolution remote sensing in NWPU VHR-10), indicate that the improvements introduced by IMO, DS-SPPF, and HPConv are consistently reproducible rather than tied to a specific dataset distribution.

5. Discussion

To address the challenges of drastic scale variations, severe occlusion, and cluttered backgrounds in UAV aerial imagery, this paper proposes the MD-YOLO detection model. Experimental results demonstrate that MD-YOLO achieves notable improvements over the YOLO26 baseline across all primary detection metrics. Specifically, Precision increases from 46.0% to 49.9%, indicating a reduction in false positives and enhanced reliability of predictions. Recall increases from 36.1% to 39.3%, indicating that MD-YOLO retrieves more ground-truth targets across the evaluated UAV scenarios. Meanwhile, mAP50 improves from 36.9% to 39.8%, and mAP50-95 from 20.0% to 23.1%. These results confirm that MD-YOLO not only exhibits superior overall detection capability under a relatively loose IoU threshold but also achieves higher localization accuracy and stability under stricter multi-threshold evaluation criteria. The above findings indicate that the three proposed modules—IMO, DS-SPPF, and HPConv—positively impact detection performance from distinct perspectives. Unlike approaches that rely solely on deepening or widening the network, the proposed method optimizes the model along three complementary dimensions: backbone feature extraction, multi-scale contextual modeling, and downsampling detail preservation. Consequently, MD-YOLO achieves improved detection accuracy while maintaining a low parameter count, thereby striking a favorable balance between accuracy and complexity.
MD-YOLO enhances detection performance while preserving a low parameter count. Compared with YOLO26, the parameter count of MD-YOLO remains essentially unchanged (9.4 M → 9.3 M), and the GFLOPs rise from 20.5 G to 22.3 G—a moderate increase that remains within an acceptable range. This result demonstrates that the proposed method does not rely on a substantial expansion of network scale to achieve accuracy gains; rather, it improves feature representation quality through more targeted architectural design. For UAV object detection tasks, a model must not only deliver high detection accuracy but also maintain practical deployability. By achieving accuracy improvements under well-controlled parameter and computational budgets, MD-YOLO lays a solid foundation for its practical deployment on edge devices.
Nevertheless, the proposed method has certain limitations. First, in scenarios involving extremely small objects, severe occlusion, and high-density targets, object boundaries are often blurred and inter-instance distances are minimal, which may still lead to missed detections or inaccurate localization. Second, when images exhibit pronounced motion blur, low illumination at night, or strong glare interference, the texture and edge information of targets is attenuated, potentially limiting the effectiveness of the local-enhancement branch. While we report FPS on a desktop RTX 4090, we have not yet profiled inference latency and energy consumption on low-power UAV edge devices; on-board real-time performance thus requires further validation. We acknowledge that Params and GFLOPs are only proxies for computational complexity; actual on-board latency and energy consumption depend on hardware and quantization choices, which we plan to investigate in future work. Finally, while qualitative examples suggest improved detection on occluded and densely distributed targets, we did not perform attribute-stratified quantitative evaluation on occlusion level, truncation, or image-level density. A rigorous size- and attribute-stratified analysis is left as future work.
Although the experiments in this work focus on three daytime UAV benchmarks (VisDrone-2019, UAVDT, and NWPU VHR-10), the three proposed modules may also benefit other small-object domains; however, validating this generalization is beyond the scope of this work and left for future study. Under more challenging conditions such as heavy rain, fog, low-light conditions, or strong atmospheric scattering, however, small-object edges and textures degrade further and may exceed the compensation capacity of the local-enhancement branch; integrating our design with image restoration preprocessing or weather-adaptive attention constitutes an important direction for all-weather UAV deployment.
Future research can be pursued along the following directions. First, model compression techniques such as pruning, quantization, and knowledge distillation can be integrated to further reduce computational overhead and enhance edge-deployment efficiency. Second, finer-grained feature-enhancement mechanisms can be designed specifically for extremely small objects and densely occluded targets to improve the model’s discriminative capability for weakly responsive targets. Third, lightweight attention mechanisms or temporal modeling can be incorporated to extend single-frame detection to UAV video detection scenarios, thereby enhancing detection stability in complex dynamic environments. Overall, MD-YOLO demonstrates favorable effectiveness and application potential in UAV small-object detection tasks; however, there remains room for further improvement in handling extremely complex scenarios and achieving real-time deployment on edge devices.

6. Conclusions

To address the challenges of drastic scale variations, severe occlusion, and cluttered backgrounds in UAV aerial imagery, this paper proposes an improved object detection model, MD-YOLO. Built upon the YOLO26 baseline, the model introduces the IMO backbone to enhance the representation of edges, textures, and local structural features of small objects; designs the DS-SPPF module to improve the aggregation of multi-scale contextual information; and replaces the conventional downsampling convolutions in the Neck with HPConv to mitigate the loss of fine-grained small-object details during scale transitions. Experimental results demonstrate that MD-YOLO outperforms competing models on the VisDrone-2019 dataset. Compared with YOLO26, MD-YOLO exhibits notable improvements in Precision, Recall, mAP50, and mAP50-95. Ablation studies further validate the effectiveness of each proposed module, and per-class detection results indicate that the model delivers more stable detection performance on small-object categories, while the qualitative visualization further suggests improved robustness for occluded targets, which warrants a rigorous attribute-stratified evaluation in future work. Independent experiments on additional datasets also demonstrate that MD-YOLO maintains consistent effectiveness on the UAVDT and NWPU VHR-10 datasets. Overall, MD-YOLO enhances the accuracy of small-object detection in UAV scenarios while maintaining a manageable model scale, thereby achieving a favorable accuracy–efficiency trade-off. Future work will focus on further model compression, edge-deployment optimization, and the enhancement of generalization capability in extremely complex scenarios.

Author Contributions

Conceptualization, T.Z. and W.Z.; methodology, M.Y. and W.Z.; software, W.Z.; validation, S.S., Q.-E.W. and J.C.; formal analysis, W.Z.; investigation, T.Z. and W.Z.; resources, W.Z.; writing—original draft preparation, T.Z. and W.Z.; writing—review and editing, S.S., Q.-E.W., Q.D. and J.C.; visualization, Z.F., T.Z. and W.Z.; supervision, S.S.; project administration, Q.-E.W. and Q.D. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by the project “A Collaborative Perception System for Multi-Source Traffic Targets Based on a Trusted Data Space” (Grant No. A01GY26F012), Key Science and Technology Project of Henan Province University (26B413002) and Henan Provincial Key Laboratory of Livestock and Poultry Genetic Improvement and Healthy Breeding (C81JX260004).

Data Availability Statement

The datasets used in this study are publicly available, including VisDrone-2019, UAVDT, and NWPU VHR-10. These datasets can be obtained from their official release websites.

Conflicts of Interest

Author Qi Ding was employed by the company Ningbo Whale Flare Information Technology Co., Ltd., Ningbo, China. The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

References

  1. Mujtaba, G.; Liu, W.; Alshehri, M.; AlQahtani, Y.; Almujally, N.A.; Liu, H. Aerial Images for Intelligent Vehicle Detection and Classification via YOLOv11 and Deep Learner. Comput. Mater. Contin. 2026, 86, 1–19. [Google Scholar] [CrossRef] [Scilit]
  2. Al-Obeidat, N.; Li, Z. A Multimodal UAV-Based Pipeline for Precision Agriculture: Aerial Stress Detection with YOLO and High-Fidelity Disease Classification Using DeiT. Procedia Comput. Sci. 2025, 257, 921–926. [Google Scholar] [CrossRef] [Scilit]
  3. Ait el Haj, R.; Benelmostafa, B.-E.; Medromi, H. YOLOv8-ECCα: Enhancing Object Detection for Power Line Asset Inspection Under Real-World Visual Constraints. Algorithms 2026, 19, 66. [Google Scholar] [CrossRef] [Scilit]
  4. Huo, Z.; Fang, L.; Chu, Y.; Dang, S.; Yang, J.; Li, L.; Li, X.; Ren, S. Precise Urban Tree Species Identification and Biomass Estimation Using UAV-Handheld LiDAR Synergy and YOLOv11 Deep Learning. Int. J. Appl. Earth Obs. Geoinf. 2026, 146, 105049. [Google Scholar] [CrossRef] [Scilit]
  5. Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 39, 1137–1149. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  6. Dai, J.; Li, Y.; He, K.; Sun, J. R-FCN: Object Detection via Region-Based Fully Convolutional Networks. In Proceedings of the 30th International Conference on Neural Information Processing Systems (NIPS 2016), Barcelona, Spain, 5–10 December 2016; pp. 379–387. [Google Scholar]
  7. Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.-Y.; Berg, A.C. SSD: Single Shot Multibox Detector. In Computer Vision—ECCV 2016; Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2016; pp. 21–37. [Google Scholar] [CrossRef] [Scilit]
  8. Jocher, G.; Chaurasia, A.; Qiu, J. Ultralytics YOLOv8, Version 8.0.0. 2023. Available online: https://github.com/ultralytics/ultralytics/releases/tag/v8.0.0 (accessed on 18 July 2026).
  9. Wang, A.; Chen, H.; Liu, L.; Chen, K.; Lin, Z.; Han, J.; Ding, G. YOLOv10: Real-Time End-to-End Object Detection. arXiv 2024, arXiv:2405.14458. [Google Scholar]
  10. Jocher, G.; Qiu, J. Ultralytics YOLO11, Version 11.0.0. 2024. Available online: https://github.com/ultralytics/ultralytics (accessed on 18 July 2026).
  11. Tian, Y.; Ye, Q.; Doermann, D. YOLOv12: Attention-Centric Real-Time Object Detectors. arXiv 2025, arXiv:2502.12524. [Google Scholar] [CrossRef] [Scilit]
  12. Lei, M.; Li, S.; Wu, Y.; Hu, H.; Zhou, Y.; Zheng, X.; Ding, G.; Du, S.; Wu, Z.; Gao, Y. YOLOv13: Real-Time Object Detection with Hypergraph-Enhanced Adaptive Visual Perception. arXiv 2025, arXiv:2506.17733. [Google Scholar] [CrossRef] [Scilit]
  13. Jocher, G.; Qiu, J.; Liu, M.; Lyu, S.; Akyon, F.C.; Kalfaoglu, M.E. Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models. arXiv 2026, arXiv:2606.03748. [Google Scholar] [CrossRef] [Scilit]
  14. He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 770–778. [Google Scholar] [CrossRef] [Scilit]
  15. Howard, A.G.; Zhu, M.; Chen, B.; Kalenichenko, D.; Wang, W.; Weyand, T.; Andreetto, M.; Adam, H. MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. arXiv 2017, arXiv:1704.04861. [Google Scholar] [CrossRef] [Scilit]
  16. Li, C.; Zhao, R.; Wang, Z.; Xu, H.; Zhu, X. RemDet: Rethinking Efficient Model Design for UAV Object Detection. In Proceedings of the AAAI Conference on Artificial Intelligence, Philadelphia, PA, USA, 25 February–4 March 2025; Volume 39, pp. 4643–4651. [Google Scholar] [CrossRef] [Scilit]
  17. Yue, T.; Lu, X.; Cai, J.; Chen, Y.; Chu, S. YOLO-MST: Multiscale Deep Learning Method for Infrared Small Target Detection Based on Super-Resolution and YOLO. Opt. Laser Technol. 2025, 187, 112835. [Google Scholar] [CrossRef] [Scilit]
  18. Yu, W.; Wang, X. MambaOut: Do We Really Need Mamba for Vision? In Proceedings of the 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 11–15 June 2025; pp. 4484–4496. [Google Scholar] [CrossRef] [Scilit]
  19. Du, D.; Zhu, P.; Wen, L.; Bian, X.; Ling, H.; Hu, Q.; Peng, T.; Zheng, J.; Wang, X.; Zhang, Y.; et al. VisDrone-DET2019: The Vision Meets Drone Object Detection in Image Challenge Results. In Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW), Seoul, Republic of Korea, 27–28 October 2019; pp. 213–226. [Google Scholar] [CrossRef] [Scilit]
  20. Yu, H.; Li, G.; Zhang, W.; Huang, Q.; Du, D.; Tian, Q.; Sebe, N. The Unmanned Aerial Vehicle Benchmark: Object Detection, Tracking and Baseline. Int. J. Comput. Vis. 2020, 128, 1141–1159. [Google Scholar] [CrossRef] [Scilit]
  21. Cheng, G.; Zhou, P.; Han, J. Learning Rotation-Invariant Convolutional Neural Networks for Object Detection in VHR Optical Remote Sensing Images. IEEE Trans. Geosci. Remote Sens. 2016, 54, 7405–7415. [Google Scholar] [CrossRef] [Scilit]
  22. Zheng, Y.; Jing, Y.; Zhao, J.; Cui, G. LAM-YOLO: Drones-Based Small Object Detection on Lighting-Occlusion Attention Mechanism YOLO. Comput. Vis. Image Underst. 2025, 261, 104489. [Google Scholar] [CrossRef] [Scilit]
  23. Li, H.; Qu, H. DASSF: Dynamic-Attention Scale-Sequence Fusion for Aerial Object Detection. In Pattern Recognition and Computer Vision (PRCV 2024); Lecture Notes in Computer Science; Springer: Singapore, 2025; Volume 15037, pp. 212–227. [Google Scholar] [CrossRef] [Scilit]
  24. Wang, X.; Peng, Y.; Shen, C. Efficient Feature Fusion for UAV Object Detection. In Proceedings of the 2025 International Joint Conference on Neural Networks (IJCNN), Rome, Italy, 30 June–5 July 2025; pp. 1–8. [Google Scholar] [CrossRef] [Scilit]
  25. Moser, B.B.; Raue, F.; Frolov, S.; Palacio, S.; Hees, J.; Dengel, A. Context-Aware Aerial Object Detection: Leveraging Inter-Object and Background Relationships. arXiv 2024, arXiv:2404.04140. [Google Scholar] [CrossRef] [Scilit]
  26. Zang, Z.; Lin, C.; Tang, C.; Wang, T.; Lv, J. Zero-Shot Aerial Object Detection with Visual Description Regularization. Proc. AAAI Conf. Artif. Intell. 2024, 38, 6926–6934. [Google Scholar] [CrossRef] [Scilit]
  27. Bakirci, M. Performance evaluation of low-power and lightweight object detectors for real-time monitoring in resource-constrained drone systems. Eng. Appl. Artif. Intell. 2025, 159, 111775. [Google Scholar] [CrossRef] [Scilit]
  28. Rey, L.; Bernardos, A.M.; Dobrzycki, A.D.; Carramiñana, D.; Bergesio, L.; Besada, J.A.; Casar, J.R. A Performance Analysis of You Only Look Once Models for Deployment on Constrained Computational Edge Devices in Drone Applications. Electronics 2025, 14, 638. [Google Scholar] [CrossRef] [Scilit]
  29. Liu, Z.; Mao, H.; Wu, C.-Y.; Feichtenhofer, C.; Darrell, T.; Xie, S. A ConvNet for the 2020s. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 11966–11976. [Google Scholar] [CrossRef] [Scilit]
  30. Rao, Y.; Zhao, W.; Tang, Y.; Zhou, J.; Lim, S.N.; Lu, J. HorNet: Efficient High-Order Spatial Interactions with Recursive Gated Convolutions. Adv. Neural Inf. Process. Syst. 2022, 35, 10353–10366. [Google Scholar] [CrossRef] [Scilit]
  31. Katharopoulos, A.; Vyas, A.; Pappas, N.; Fleuret, F. Transformers Are RNNs: Fast Autoregressive Transformers with Linear Attention. In Proceedings of the International Conference on Machine Learning (ICML), Virtual, 12––18 July 2020; pp. 5156–5165. [Google Scholar]
  32. Su, J.; Ahmed, M.; Lu, Y.; Pan, S.; Bo, W.; Liu, Y. RoFormer: Enhanced Transformer with Rotary Position Embedding. Neurocomputing 2024, 568, 127063. [Google Scholar] [CrossRef] [Scilit]
  33. Dong, X.; Bao, J.; Chen, D.; Zhang, W.; Yu, N.; Yuan, L.; Chen, D.; Guo, B. CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 12124–12134. [Google Scholar] [CrossRef] [Scilit]
  34. Wu, S. Traffic Scene Small Target Detection Method Based on YOLOv8n-SPTS Model for Autonomous Driving. In Proceedings of the 2025 IEEE 3rd International Conference on Electrical, Automation and Computer Engineering (ICEACE), Chongqing, China, 4–6 July 2025; pp. 181–186. [Google Scholar] [CrossRef] [Scilit]
  35. Zhang, R. Making Convolutional Networks Shift-Invariant Again. In Proceedings of the 36th International Conference on Machine Learning (ICML), Long Beach, CA, USA, 9–15 June 2019; pp. 7324–7334. [Google Scholar]
  36. Gao, Z.; Wang, L.; Wu, G. LIP: Local Importance-Based Pooling. In Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 3355–3364. [Google Scholar] [CrossRef] [Scilit]
  37. Sunkara, R.; Luo, T. No More Strided Convolutions or Pooling: A New CNN Building Block for Low-Resolution Images and Small Objects. In Proceedings of the European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases, Grenoble, France, 19–23 September 2022; pp. 443–459. [Google Scholar]
  38. Han, D.; Wang, Z.; Xia, Z.; Han, Y.; Pu, Y.; Ge, C.; Song, J.; Song, S.; Zheng, B.; Huang, G. Demystify Mamba in Vision: A Linear Attention Perspective. Adv. Neural Inf. Process. Syst. 2024, 37, 127181–127203. [Google Scholar] [CrossRef] [Scilit]
  39. Zhao, Y.; Lv, W.; Xu, S.; Wei, J.; Wang, G.; Dang, Q.; Liu, Y.; Chen, J. DETRs Beat YOLOs on Real-Time Object Detection. arXiv 2023, arXiv:2304.08069. [Google Scholar]
  40. Zhang, Z. Drone-YOLO: An Efficient Neural Network Method for Target Detection in Drone Images. Drones 2023, 7, 526. [Google Scholar] [CrossRef] [Scilit]
  41. Zhong, H.; Zhang, Y.; Shi, Z.; Zhang, Y.; Zhao, L. PS-YOLO: A Lighter and Faster Network for UAV Object Detection. Remote Sens. 2025, 17, 1641. [Google Scholar] [CrossRef] [Scilit]
  42. Zhu, X.; Lyu, S.; Wang, X.; Zhao, Q. TPH-YOLOv5: Improved YOLOv5 based on transformer prediction head for object detection on drone-captured scenarios. In Proceedings of the 2021 IEEE/CVF International Conference ON Computer Vision, Montreal, BC, Canada, 1–17 October 2021; pp. 2778–2788. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Overall architecture of the proposed MD-YOLO network.
Figure 1. Overall architecture of the proposed MD-YOLO network.
Algorithms 19 00737 g001
Figure 2. Structure of GatedCNN Block.
Figure 2. Structure of GatedCNN Block.
Algorithms 19 00737 g002
Figure 3. Structure of Mamba-Inspired Linear Attention Adapter.
Figure 3. Structure of Mamba-Inspired Linear Attention Adapter.
Algorithms 19 00737 g003
Figure 4. Structure of the proposed Dynamic Scale SPPF.
Figure 4. Structure of the proposed Dynamic Scale SPPF.
Algorithms 19 00737 g004
Figure 5. Structure of the proposed Hybrid Pooling Conv Module.
Figure 5. Structure of the proposed Hybrid Pooling Conv Module.
Algorithms 19 00737 g005
Figure 6. Precision–recall curves of YOLO26s baseline model on VisDrone-2019 dataset.
Figure 6. Precision–recall curves of YOLO26s baseline model on VisDrone-2019 dataset.
Algorithms 19 00737 g006
Figure 7. Precision–recall curves of MD-YOLO model on VisDrone-2019 dataset.
Figure 7. Precision–recall curves of MD-YOLO model on VisDrone-2019 dataset.
Algorithms 19 00737 g007
Figure 8. The training dynamics of MD-YOLO.
Figure 8. The training dynamics of MD-YOLO.
Algorithms 19 00737 g008
Figure 9. Visualization results of the baseline model (a,c,e) and the proposed model (b,d,f) on the VisDrone-2019 and NWPU VHR-10 datasets.
Figure 9. Visualization results of the baseline model (a,c,e) and the proposed model (b,d,f) on the VisDrone-2019 and NWPU VHR-10 datasets.
Algorithms 19 00737 g009
Figure 10. Grad-CAM visualization comparison between YOLO26 and the proposed MD-YOLO. (a) Original image; (b) YOLO26 heatmap; (c) MD-YOLO heatmap.
Figure 10. Grad-CAM visualization comparison between YOLO26 and the proposed MD-YOLO. (a) Original image; (b) YOLO26 heatmap; (c) MD-YOLO heatmap.
Algorithms 19 00737 g010
Table 1. Training hyperparameter settings.
Table 1. Training hyperparameter settings.
ParametersValue
Epochs250
Image Size640
Batch Size6
OptimizerMuSGD
Initial Learning Rate0.005
Weight Decay0.0005
Momentum0.937
Warmup Epochs3
Mosaic1.0
Mixup0.0
Table 2. Performance comparison of algorithms on VisDrone-2019 dataset.
Table 2. Performance comparison of algorithms on VisDrone-2019 dataset.
ModelPrecision (%)Recall (%)mAP50 (%)mAP50-95 (%)Params (M)GFLOPsFPS (f/s)
YOLOv8s41.933.730.917.511.526.1178.3
YOLOv10s46.334.632.419.58.024.5195.3
YOLOv11s45.234.832.919.29.521.3148.3
YOLOv12s46.635.233.519.79.121.2156.2
YOLOv13s44.834.734.419.39.521.3106.7
YOLO26s46.036.136.920.09.420.5158.8
RT-DETR-R18 [39]46.838.636.220.720.057.0147.6
Drone-YOLO [40]--35.620.410.9--
PS-YOLO [41]44.93432.321.019.520.0150.0
TPH-YOLO [42]47.638.338.722.855.3150.0-
MD-YOLO (Ours)49.939.339.823.19.322.3124.7
Table 3. Size-stratified detection performance on VisDrone-2019 test set.
Table 3. Size-stratified detection performance on VisDrone-2019 test set.
ModelAP_S (%)AP_M (%)AP_L (%)
YOLOv8s11.832.640.5
YOLOv11s12.633.442.6
YOLO26s13.432.842.8
RT-DETR15.733.645.9
MD-YOLO16.835.546.4
Table 4. Ablation experiment results on VisDrone-2019 dataset.
Table 4. Ablation experiment results on VisDrone-2019 dataset.
ModelIMODS-SPPFHPConvp (%)R (%)mAP50 (%)mAP50-95 (%)Params (M)GFLOPs
Exp1 46.036.136.920.09.420.5
Exp2√ 48.737.938.022.39.221.3
Exp3 √ 47.136.637.220.89.620.7
Exp4 √46.836.437.620.79.320.3
Exp5√√ 49.238.038.422.09.322.5
Exp6√ √48.638.538.722.69.221.6
Exp7 √√47.437.838.221.49.520.4
Exp8 (Ours)√√√49.939.339.823.19.322.3
Table 5. Internal ablation of the IMO backbone on the VisDrone-2019 dataset.
Table 5. Internal ablation of the IMO backbone on the VisDrone-2019 dataset.
ModelPrecision (%)Recall (%)mAP50 (%)mAP50-95 (%)Params (M)GFLOPs
GatedCNN only46.736.436.820.49.220.3
+Local Enhancement47.236.837.321.09.221.0
+MILA48.737.938.022.39.221.3
Table 6. Comparison of downsampling strategies in the Neck of MD-YOLO on the VisDrone-2019 dataset.
Table 6. Comparison of downsampling strategies in the Neck of MD-YOLO on the VisDrone-2019 dataset.
ModelBackboneSPPFmAP50 (%)mAP50-95 (%)Params (M)GFLOPs
Standard ConvIMODS-SPPF38.422.09.322.5
Anti-aliased IMODS-SPPF38.221.89.422.6
SPD-ConvIMODS-SPPF38.722.410.622.6
HPConvIMODS-SPPF39.823.19.322.3
Table 7. Generalization experiment results.
Table 7. Generalization experiment results.
DatasetModelsPrecision (%)Recall (%)mAP50 (%)mAP50-95 (%)
UAVDTYOLO26s47.647.343.825.7
MD-YOLO52.855.952.129.4
NWPU VHR-10YOLO26s92.691.093.463.6
MD-YOLO93.192.394.865.0
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Zhang, T.; Zou, W.; Yu, M.; Song, S.; Wu, Q.-E.; Chen, J.; Fan, Z.; Ding, Q. MD-YOLO: An Improved YOLO26-Based Model for Small-Object Detection in UAV Aerial Imagery. Algorithms 2026, 19, 737. https://doi.org/10.3390/a19090737

AMA Style

Zhang T, Zou W, Yu M, Song S, Wu Q-E, Chen J, Fan Z, Ding Q. MD-YOLO: An Improved YOLO26-Based Model for Small-Object Detection in UAV Aerial Imagery. Algorithms. 2026; 19(9):737. https://doi.org/10.3390/a19090737

Chicago/Turabian Style

Zhang, Tongpo, Wangquan Zou, Moyan Yu, Shaojing Song, Qin-E Wu, Jian Chen, Ziyu Fan, and Qi Ding. 2026. "MD-YOLO: An Improved YOLO26-Based Model for Small-Object Detection in UAV Aerial Imagery" Algorithms 19, no. 9: 737. https://doi.org/10.3390/a19090737

APA Style

Zhang, T., Zou, W., Yu, M., Song, S., Wu, Q.-E., Chen, J., Fan, Z., & Ding, Q. (2026). MD-YOLO: An Improved YOLO26-Based Model for Small-Object Detection in UAV Aerial Imagery. Algorithms, 19(9), 737. https://doi.org/10.3390/a19090737

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop