Skip to Content
Applied SciencesApplied Sciences
  • Article
  • Open Access

12 March 2026

22 Pages

Accurate Detection of Large-Leaf Tea Buds in Mountainous Tea Plantations Based on an Improved YOLO Framework

,
,
,
,
and
1
College of Big Data and Intelligent Engineering, Southwest Forestry University, Kunming 650224, China
2
College of Forestry, Southwest Forestry University, Kunming 650224, China
3
College of Landscape Architecture and Horticulture, Southwest Forestry University, Kunming 650224, China
*
Authors to whom correspondence should be addressed.

Abstract

Tea buds are the key raw material for high-quality tea production, and their accurate perception is essential for intelligent harvesting and quality-oriented management. However, tea bud detection in mountainous large-leaf tea plantations remains challenging because small, densely distributed targets are embedded in complex field environments, significantly limiting the stability and accuracy of existing detection methods. To address these challenges, this study proposes an improved tea bud detection model, termed YOLO-LAR, for mountainous large-leaf tea plantations in Yunnan Province, China, which is developed as an enhanced framework based on the YOLOv11 baseline. YOLO-LAR improves feature representation through multi-scale feature fusion, enabling more effective detection of densely distributed small tea buds. In addition, an optimized downsampling strategy is employed to preserve critical spatial information, and a context-enhanced feature aggregation mechanism is introduced to strengthen robustness under complex backgrounds and illumination variations. The results demonstrate that YOLO-LAR achieves precision, recall, mAP@0.50, and mAP@0.50:0.95 of 0.959, 0.908, 0.961, and 0.814, respectively, outperforming mainstream YOLO-based models, including YOLOv11n, YOLOv10n, and YOLOv8n. These results indicate that YOLO-LAR provides an effective and practical solution for accurate tea bud detection, offering strong technical support for intelligent harvesting and precision management in mountainous tea plantation environments.

1. Introduction

Tea plant (Camellia sinensis) is a perennial agroforestry crop widely cultivated in tropical and subtropical regions, forming an essential agricultural foundation that supports rural livelihoods and the global tea supply chain [1,2]. As one of the most important cash crops in mountainous and hilly areas, tea production plays a critical role in sustaining smallholder farmers’ income and promoting regional economic development [3]. With increasing labor shortages and rising production costs, the tea industry is undergoing a gradual transition toward mechanization and intelligent agricultural systems [4]. In this background, improving harvesting efficiency while maintaining tea quality has become a core engineering concern in modern tea production systems [5]. Accurate perception of tea buds and young leaves serves as a fundamental prerequisite for intelligent harvesting, yield estimation, and precision management, especially in complex natural tea garden environments. Despite the increasing mechanization of tea harvesting, accurate detection of small, densely distributed tea buds under complex field conditions remains a critical technical challenge, motivating the development of specialized detection models.
Accurate detection of large-leaf tea buds in mountainous tea plantations is a challenging yet critical task for automated harvesting systems. In particular, the large-leaf tea plantations typical of Yunnan Province, China, are characterized by small, densely distributed buds that are frequently occluded by surrounding leaves [6]. In addition, irregular row spacing, uneven terrain, and strong variations in natural illumination further complicate visual perception [7,8]. If tea buds are not accurately detected, harvesting operations may suffer from missed targets or excessive leaf inclusion, leading to reduced efficiency, compromised raw material consistency, and lower tea quality [9]. Traditionally, harvesting is performed either manually or using mechanical tools. Manual plucking, though capable of preserving leaf integrity and ensuring premium quality, is labor-intensive and economically unsustainable for large-scale operations [10]. On the other hand, mechanical harvesting significantly improves efficiency but often leads to physical damage and reduced quality [11,12]. These limitations highlight the urgent need for automated, accurate tea bud detection systems that can support intelligent harvesting and modernization of the tea industry [13].
These limitations highlight the urgent need for automated, accurate tea bud detection systems that can support intelligent harvesting and modernization of the tea industry. Existing automated methods, including traditional image processing and standard YOLO-based detectors, struggle to detect small tea buds that are densely clustered and often occluded, particularly in mountainous large-leaf tea plantations with variable lighting and uneven terrain.
Previous studies have explored tea bud detection using conventional image processing techniques based on handcrafted features such as color, texture, and morphology [14,15,16,17]. While these traditional methods are capable of extracting certain low-level visual features under controlled conditions, they generally suffer from limited robustness and poor generalizability in complex and dynamic tea plantation environments. Their performance often degrades in the presence of variations in lighting, occlusion, background clutter, and shoot morphology, thus constraining their practical applicability in real-world tea plantations.
With the advancement of precision agriculture, deep learning-based object detection methods have increasingly been applied to crop detection tasks [18] and harvesting tasks [19], enabling automated, efficient, and precise crop management. Compared with traditional image processing methods based on handcrafted features, deep learning models demonstrate stronger feature representation capabilities and improved robustness to environmental variations such as citrus (Citrus sinensis) [20], strawberry (Fragaria ananassa) [21], and mango (Mangifera indica) [22]. Among them, one-stage detectors represented by the You Only Look Once (YOLO) series have attracted considerable attention due to their favorable balance between detection accuracy and inference efficiency, making them suitable for real-time deployment in mechanized harvesting systems. YOLO-based approaches have been successfully applied to various agricultural scenarios, including fruit detection, disease identification, and general crop monitoring [23,24,25].
However, standard YOLO models often exhibit limited effectiveness when applied to tea bud detection in mountainous environments. Tiny targets, dense spatial distribution, and complex foliage frequently lead to missed detections and imprecise localization, which can compromise both harvesting efficiency and tea quality [26]. More importantly, tea bud perception for intelligent harvesting in mountainous tea gardens imposes stricter requirements on detection precision and spatial consistency than those of general agricultural monitoring tasks. Effective perception must preserve fine-grained spatial details, maintain stable localization under severe occlusion, and remain robust to strong illumination variations and background complexity. Nevertheless, most existing YOLO-based methods are primarily designed for standard detection benchmarks or relatively structured agricultural fields [27,28]. They have not been explicitly optimized for the operational requirements of large-leaf tea bud harvesting in mountainous tea gardens. As a result, their performance can degrade under real harvesting conditions, where maintaining both accuracy and consistency is critical for efficient and high-quality harvesting.
YOLO11, released by the Ultralytics team in September 2024, introduces an improved backbone and neck architecture that enhances feature extraction, accelerates inference, and balances accuracy with efficiency while supporting edge deployment. However, its effectiveness in detecting tiny and densely clustered tea buds under complex plantation conditions has not yet been systematically evaluated.
To evaluate tea bud detection under realistic agricultural conditions, this study focuses on representative large-leaf tea plantations in central Yunnan Province, characterized by complex canopy structures, uneven terrain, and strong illumination variability. In such environments, small and densely distributed tea buds are frequently occluded by surrounding leaves, making accurate detection challenging. Existing automated methods, including traditional image processing techniques and standard YOLO-based detectors, often fail to provide precise and stable localization under these conditions.
A detection framework capable of handling dense buds under severe occlusion, variable lighting, and uneven terrain is therefore required to address this technical gap. Such a framework should integrate fine-grained feature refinement and multi-scale perception modules to enhance localization precision, detection stability, and robustness in complex field environments. Based on these design principles, we introduce YOLO-LAR, an enhanced variant of YOLOv11n specifically developed for large-leaf tea bud detection in mountainous tea gardens.
The main contributions of this study are summarized as follows:
  • A large-leaf tea bud detection dataset was constructed from representative tea plantations in Yunnan Province, providing realistic samples under complex field conditions.
  • An improved YOLOv11-based detection framework, termed YOLO-LAR, is proposed to enhance the detection accuracy of small and densely distributed tea buds in mountainous environments.
  • Extensive experiments demonstrate that YOLO-LAR achieves superior accuracy and robustness compared with mainstream YOLO-based models, supporting intelligent and non-invasive tea bud harvesting.

2. Study Area and Dataset

2.1. Study Area

Yunnan Province in southwestern China has a favorable ecological environment and diverse climate, making it one of the most important regions for large-leaf tea. This study focuses on two representative tea-producing areas in central Yunnan (24°41′–25°05′ N, 102°09′–102°44′ E): the Qizi Tea Plantation and the Yunyi Chun Tea Plantation. These tea plantations are situated on plateau terrain with large topographic relief. The region experiences a subtropical monsoon climate, with an annual mean temperature ranging from 15.4 to 24.2 °C, a historical maximum of 32.2 °C, and a minimum of −3 °C, and annual precipitation between 788 and 1000 mm, mainly from May to October. Examples of the two representative tea plantation environments are shown in Figure 1 as Case a and Case b. The complex canopy structure, variable illumination, and heterogeneous natural backgrounds in these plantations provide a suitable real-world setting for evaluating the performance and robustness of tea bud detection models.
Figure 1. Study area and representative tea garden environments. The two sampling sites are (a) the Qizi Tea Plantation and (b) the Yunyi Chun Tea Plantation.

2.2. Dataset and Preprocessing

The dataset and preprocessing procedures used in this study are fully described in this section, including image acquisition, annotation, and data augmentation, to ensure transparency and reproducibility.

2.2.1. Image Acquisition

RGB images of tea buds were collected from the two representative plantations introduced in Section 2.1 using the rear camera of an iPhone 15, which captures images at a native resolution of 4032 × 3024 pixels. A total of 1024 images were acquired under varying weather conditions, including sunny and cloudy days, and from different shooting distances and viewpoints (Figure 2), to capture the visual complexity of real tea garden environments.
Figure 2. Representative RGB images of tea buds: (a) cloudy day; (b) close distance; (c) upward view; (d) sunny day; (e) long distance; (f) top view.

2.2.2. Image Dataset Preprocessing

All collected images were then manually annotated using the Make Sense online labeling tool. To enhance visual diversity and improve model robustness, moderate spatial transformations were applied, including center cropping and random flipping. Additionally, photometric adjustments, such as brightness and contrast modifications, along with light Gaussian noise, were introduced to simulate variations in illumination and environmental disturbances under real-world imaging conditions. Representative examples of the augmented images are shown in Figure 3. All augmentations preserved the original annotations, resulting in a total of 9378 valid images. The final dataset was randomly divided into training, validation, and test sets at an 8:1:1 ratio, corresponding to 7502, 938, and 938 images, respectively.
Figure 3. Representative samples of data augmentation applied to tea bud images: (a) original image; (b) brightness image (50% size); (c) salt noise image; (d) horizontally flipped image; (e) pepper noise image; (f) randomly contrast-adjusted image.

3. Methods

The dataset and preprocessing procedures for training and evaluation are detailed in Section 2.2.

3.1. YOLOv11 and Environment Configuration

YOLOv11 exhibits improved accuracy, efficiency, and architectural design compared with previous YOLO versions [29]. Compared with YOLOv8, YOLOv11 incorporates the C3k2 multi-scale fusion module and employs depthwise convolutions in the detection head, enhancing feature representation while reducing computational cost. To ensure fair and reproducible evaluation, all models were trained and tested within a unified experimental environment. Ablation studies used the same hyperparameter configuration, whereas comparison experiments followed the default settings of their respective models. All experiments were conducted on a single workstation to maintain consistent hardware and software conditions (Table 1). The workstation was equipped with an Intel Xeon Silver 4214R CPU (Intel Corporation, Santa Clara, CA, USA) and an NVIDIA GeForce RTX 3080 Ti GPU (NVIDIA Corporation, Santa Clara, CA, USA). The software environment included Python 3.12, PyTorch 2.3.0, PyCharm 2025.1, CUDA 12.1, and cuDNN 8.9.1. Training was performed for 150 epochs using 640 × 640 input images and optimized with stochastic gradient descent (SGD) with momentum. Unless otherwise specified, all remaining hyperparameters followed the default configuration of the YOLOv11 framework.
Table 1. Hardware configuration and operating environment.

3.2. YOLO-LAR Model Overview

To provide a clear methodological overview and explicitly link the data preparation, model improvements, feature extraction, feature selection, and evaluation processes, Figure 4 presents the complete methodological framework of the proposed YOLO-LAR model. This flowchart illustrates the end-to-end pipeline from dataset construction to detection results, highlighting the key enhancements in feature representation for small and densely distributed tea buds in complex plantation environments.
Figure 4. Methodological workflow of the proposed YOLO-LAR, illustrating dataset construction, feature extraction, feature selection, and detection.
Figure 4 shows the methodological framework of the proposed YOLO-LAR model. The pipeline includes (1) dataset construction with image acquisition, augmentation, and splitting; (2) the improved network architecture based on YOLOv11n, with targeted enhancements for feature extraction (C3k2_RFCAConv), feature preservation (ADown), and feature selection (SPPF-LSKA); and (3) evaluation and visualization of detection results.
Building upon this framework, the proposed YOLO-LAR is an enhanced YOLOv11n model tailored for tea bud detection. Input images (640 × 640) are preprocessed using the augmentation strategies described in Section 2.2.2. The backbone network extracts multi-scale features, with three key improvements introduced to address the challenges of small, dense targets and complex backgrounds:
The integration of RFCAConv into the original C3k2 module results in the generation of C3k2_RFCAConv, replacing standard C3k2 in both the backbone and neck. This enables multi-scale receptive field attention and coordinate-aware weighting, enhancing feature extraction for fine-grained local details and long-range context. Conventional strided convolution downsampling is replaced with ADown in the backbone and neck paths. This adaptive downsampling preserves critical spatial information during resolution reduction, preventing the loss of tiny bud features. The original SPPF is upgraded to SPPF-LSKA at the backbone end. This adaptive large-kernel attention selectively aggregates multi-scale semantics, improving global-to-local context fusion and robustness against occlusion and weak saliency.

3.2.1. Multi-Scale Receptive Field Attention Module

Although the C3k2 structure in YOLOv11n improves feature aggregation through kernel decomposition, its fixed receptive field and limited positional encoding restrict its ability to capture fine-grained and directionally sensitive features. Multi-scale attention mechanisms have been shown to enhance both local detail preservation and global contextual perception [30]. To address these limitations, we adopt the Receptive Field Coordinate Attention Convolution (RFCAConv) [31], which combines receptive-field attention (RFA) and coordinate attention (CA) [32] to enhance discriminative feature representation (Figure 5).
Figure 5. RFCAConv module structure.
The feature transformation process of RFCAConv is formulated as follows: To capture multi-scale and direction-sensitive features, the RFCAConv module first generates high-dimensional receptive field features from the input feature map X ∈ R B × C × H × W .
This operation can be formulated as follows:
G = D W Conv k ( X ) , G ∈ R B × ( C k 2 ) × H × W
where each pixel in G encodes local context within its receptive field, and the channel dimension is expanded from C to C k 2 to capture multi-scale information.
Directional attention weights along the horizontal and vertical axes, A h and A w , are then computed using the coordinate attention (CA) module, which aggregates contextual information and encodes long-range positional dependencies. Finally, the input features are re-weighted using these attention maps and projected back to the original channel dimension through a convolutional layer, yielding the output feature map:
O = Conv ( G ⊙ A h ⊙ A w )
where ⊙ denotes element-wise multiplication. This operation emphasizes target-related features while suppressing irrelevant background, completing the RFCAConv module in a closed-loop and efficient manner. By embedding RFCAConv into the C3k2 bottleneck, we generate the C3k2_RFCAConv module (Figure 6), which realizes multi-scale receptive-field attention while preserving the lightweight design of C3k2. Input features are split into two branches: the main branch sequentially passes through two RFCAConv modules, while the shortcut branch preserves the original pathway. The outputs are concatenated and fused to produce the final feature map. This design enhances multi-scale feature fusion and channel selectivity, improving detection accuracy and generalization for small, densely packed, or partially occluded tea buds in complex plantation environments.
Figure 6. C3k2_RFCAConv module structure: (a) Structure of the C3k_RFCAConv block; (b) C3k2 bottleneck with RFCAConv (C3k=False); (c) Proposed C3k2_RFCAConv module with RFCAConv embedded in C3k2 (C3k=True).

3.2.2. ADown Downsampling Module

YOLOv11n downsamples feature maps using a standard 3 × 3 convolution with a stride of 2, which reduces spatial resolution while increasing the number of channels. This operation may cause feature entanglement, leading to higher computational cost and reduced model efficiency. To address this issue, we adopt the ADown module, adapted from the YOLOv9 project, which implements an adaptive downsampling strategy to preserve the fine-grained spatial information of small agricultural targets, such as tea buds, while reducing computational overhead (Figure 7).
Figure 7. Structure of the ADown module. Note: H and W denote the height and width of the feature map, respectively, and C denotes the number of channels.
The ADown module operates via a lightweight dual-branch mechanism. The input feature map X ∈ R B × C × H × W is first processed by a 2 × 2 average pooling operation to suppress background noise and reduce spatial redundancy:
X avg ( b , c , i , j ) = 1 2 × 2 ∑ m = 0 1   ∑ n = 0 1   X ( b , c , 2 i + m , 2 j + n )
The pooled features are then evenly split along the channel dimension into two branches. In the first branch, a 3 × 3 convolution enhances local texture and preserves fine structural details:
Y 1 = Conv 3 × 3 ( X avg )
The second branch retains complementary information, and the output of both branches is concatenated to produce the final downsampled feature map:
Y = Concat [ Y 1 , Y 2 ] ∈ R B × C ′ × H 2 × W 2
In the improved YOLOv11n backbone, the input image is first processed by two standard convolutions, followed by a C3k2_RFCAConv module for multi-scale feature extraction. All subsequent downsampling convolutions are replaced with ADown modules, preserving local and global features while reducing computational overhead.

3.2.3. SPPF-LSKA Module

The standard SPPF module performs cross-layer semantic fusion via bottom-up feedforward decoding but lacks explicit mechanisms to leverage prior-stage semantic information, which limits its ability to represent small and weakly salient targets, such as densely distributed tea buds. To overcome this limitation, we propose the SPPF-LSKA module, which combines the multi-scale feature aggregation of SPPF achieved by Large Selective Kernel Attention (LSKA) [33] (Figure 8). LSKA dynamically adjusts its receptive fields according to spatial context, allowing the module to capture both prior-stage semantic cues and discriminative local and global features. This integration enhances the network’s ability to represent small, weakly salient targets, which is critical for accurate tea bud detection under challenging environments.
Figure 8. Structural illustration of the SPPF-LSKA module.
To implement this adaptive receptive field modeling, the LSKA mechanism begins by decomposing a K × K convolutional kernel into three components: a (2d-1) × (2d-1) depthwise convolution, a (K/d) × (K/d) depthwise dilated convolution, and a 1 × 1 pointwise convolution, as illustrated in Figure 9a. Subsequently, both the 2D depthwise and dilated convolutional kernels are further factorized into sequential 1D horizontal and vertical convolutions, as shown in Figure 9b,c. These decomposed convolutional operations are then concatenated in series to construct a computationally efficient attention module. This hierarchical decomposition enables LSKA to capture multi-scale spatial dependencies while maintaining low computational overhead.
Figure 9. Decomposition and factorization of large convolutional kernels: (a) Decomposition of a K × K convolution into DW-Conv, DW-D-Conv, and a 1 × 1 convolution; (b) Factorization of DW-Conv into 1D convolutions; (c) Factorization of DW-D-Conv into 1D convolutions. Different colors indicate different positions involved in the convolution operation.

3.3. Evaluation Indicators

To comprehensively evaluate the proposed model, both detection accuracy and computational efficiency were considered. The evaluation metrics include precision (P), recall (R), average precision (AP), average precision (mAP), model parameters, and GFLOPs. These metrics jointly reflect the model’s detection performance and computational efficiency.
Precision and recall are fundamental metrics that quantify detection accuracy and sensitivity. Let true positives (TP) denote correctly detected positive samples, false positives (FP) denote negative samples incorrectly predicted as positive, and false negatives (FN) denote missed positive samples. Precision measures the proportion of correctly predicted positives among all predictions, while recall measures the proportion of actual positives correctly identified:
P = T P T P + F P
R = T P T P + F N
Average precision (AP) is computed as the area under the precision–recall curve for each class, and the mean average precision (mAP) is obtained by averaging AP over all classes. In this study, two tea bud categories are considered. Accordingly, mAP@0.50 (all) and mAP@0.50:0.95 (all) are reported to evaluate the overall detection performance at a single Intersection over Union (IoU) threshold (0.50) and across multiple IoU thresholds (0.50–0.95), respectively:
m A P @ 0.50 = 1 m ∑ i = 1 m   A P i
mAP @ 0.50 : 0.95 = 1 m ∑ i = 1 m   1 10 ∑ j = 1 10   A P i ( τ j )
To assess computational efficiency, GFLOPs and model parameters are measured. GFLOPs quantify the number of floating-point operations required for a single forward pass and are approximately calculated for standard convolutional layers as follows:
GFLOPs = 2 × ( K × K × C i n × C o u t × H × W + C o u t ) 10 9
where K denotes the kernel size, C i n and C o u t represent the numbers of input and output channels, and H and W are the height and width of the feature map. Model parameters indicate memory consumption and overall model size. Together, these metrics provide a balanced assessment of both detection accuracy and computational efficiency.

4. Results and Discussion

4.1. Performance of YOLO-LAR Network

To evaluate the performance differences between the two models, this section presents a comparison of their training curves in terms of precision, recall, mAP@0.50, and mAP@0.50:0.95 (Figure 10). Overall, both models exhibited continuous performance improvement during training, while YOLO-LAR consistently showed higher accuracy and faster convergence than YOLOv11n. The precision of YOLO-LAR rose quickly and reached about 0.50 by epoch 15, earlier than YOLOv11n, and remained higher throughout the training process (Figure 10a). This result revealed that YOLO-LAR achieved more reliable positive predictions throughout the training process. Furthermore, the recall curves showed a similar trend (Figure 10b). YOLO-LAR displayed a steeper increase during the first 30 epochs and remained consistently higher than YOLOv11n, achieving a high final recall value, indicating fewer missed detections. In Figure 10c, mAP@0.50 of both models exhibited an overall upward trend, while YOLO-LAR maintained a clear performance margin over YOLOv11n in most training periods. Notably, after epoch 60, YOLO-LAR reaches a higher final mAP@0.50, highlighting its improved detection accuracy under a moderate IoU threshold. Under stricter IoU conditions (Figure 10d), both YOLO-LAR and YOLOv11n exhibited similar trends, indicating that the relative performance differences between the two models are consistent even when more precise localization is required. The above results collectively indicated that YOLO-LAR consistently outperformed YOLOv11n in all evaluation metrics, achieving more reliable positive predictions, fewer missed detections, and higher detection accuracy under both moderate and strict IoU conditions.
Figure 10. Training process of YOLOv11n and YOLO-LAR: (a) precision; (b) recall; (c) mAP@0.50; (d) mAP@0.50:0.95.

4.2. Module Effectiveness

To assess the effect of the proposed C3k2_RFCAConv on the representational capability of the detection network, a comparative analysis is conducted with several attention-enhanced C3k2 variants under the same YOLOv11n framework, as summarized in Table 2. These variants represent typical attention paradigms, including Efficient Multi-Scale Attention (EMA [34]), ConvFormer [35], and Deformable Attention Transformer (DAttention [36]). The results indicate that C3k2_RFCAConv achieves the best performance among the five variants in key metrics, attaining the highest mAP@0.50 of 0.945 and recall of 0.889. Although its precision of 0.942 is slightly lower than that of C3k2_EMA, it maintains a highly competitive level. More notably, C3k2_RFCAConv significantly outperforms the other variants in mAP@0.50:0.95, reaching 0.787, which highlights its superior localization accuracy under stricter IoU thresholds. In terms of computational efficiency, it also performs well, with GFLOPs of 6.5, matching that of C3k2_ConvFormer.
Table 2. Comparison of different methods for improving C3k2.
Experimental results showed that incorporating advanced feature extraction strategies into the C3k2 bottleneck significantly enhances detection performance. The proposed C3k2_ RFCAConv module, by introducing a receptive field and a coordinate-aware attention mechanism, achieves leading performance across multiple metrics. This is consistent with previous studies showing that multi-scale receptive field modules can enhance detection precision [37]. Figure 11 provides qualitative evidence for the effectiveness of individual modules at different stages of the network, illustrating how feature extraction, downsampling, and contextual modeling are progressively enhanced. This improvement is further illustrated in Figure 11a: compared with the baseline network, which exhibits scattered and weak activations, C3k2_RFCAConv generates more compact and structurally consistent feature maps that align well with the spatial distribution of tender buds.
Figure 11. Qualitative visualization of feature activation maps for different modules in YOLO-LAR. All activation maps are generated from the same input image: (a) comparison between the baseline C3k2 and the proposed C3k2_RFCAConv; (b) comparison between standard convolution-based downsampling and the ADown module; (c) comparison between the baseline SPPF and the proposed SPPF-LSKA.
Beyond the preceding analysis of feature refinement modules, we next focus on the downsampling stage, which plays a key role in preserving fine-grained features essential for small-object representation. To analyze the effectiveness of the ADown module, this study replaced the original downsampling layer with five typical downsampling strategies, as summarized in Table 3, including the baseline Conv, DWConv [38], LDConv [39], GCConv [40], and the introduced ADown module. The remaining network architecture was kept consistent across all experiments. ADown achieved the highest overall performance with a precision of 0.941, a recall of 0.862, and an mAP@0.50:0.95 of 0.740. Its mAP@0.50 of 0.934 remained high but was slightly lower than that of DWConv. At the same time, ADown was the most efficient among the compared modules, with just 5.3 GFLOPs and 2.10 M parameters. Overall, ADown attains the best trade-off between precision and efficiency [25]. As shown in Figure 11b, ADown exhibits more compact and stable activation responses on tea bud regions, indicating that it effectively preserves discriminative features during downsampling.
Table 3. Comparison of different downsampling modules.
Although feature refinement and downsampling optimization effectively enhance local representation and efficiency, they remain insufficient for modeling long-range contextual dependencies in complex agricultural scenes. To rigorously assess the effectiveness of SPPF-LSKA, several representative attention mechanisms—EMA [34], CAA [41], SimAM [42], and MLCA [43]—were individually integrated into the SPPF layer in place of the original configuration, while keeping the YOLOv11n backbone unchanged. The performance of these methods in terms of accuracy and efficiency for tea bud detection is summarized in Table 4. SPPF-LSKA outperformed the alternatives across multiple evaluation metrics, achieving a precision of 0.944, a recall of 0.843, mAP@0.50 of 0.929, and mAP@0.50:0.95 of 0.727. These results indicate that SPPF-LSKA exhibits robust performance under complex backgrounds [44]. Among the alternatives, SPPF-CAA achieved relatively high precision and recall, reaching 0.927 and 0.865, respectively, but required higher computational resources of 6.7 GFLOPs. By contrast, SPPF-SimAM and SPPF-MLCA required lower computational resources of 6.3 GFLOPs but delivered lower precision and mAP scores. As shown in Figure 11c, SPPF-LSKA produces more spatially coherent and structurally consistent activation patterns over tea bud regions and their surrounding context, confirming its ability to capture broader semantic information in cluttered backgrounds. Overall, the experimental results indicate that performance gains in dense tea bud detection are primarily driven by improved feature preservation at early network stages and enhanced contextual modeling at deeper layers.
Table 4. Comparison of different methods for improving SPPF.

4.3. Ablation Experiments

To systematically evaluate the effectiveness of the C3k2_RFCAConv, ADown, and SPPF-LSKA modules, an ablation study was performed by progressively integrating these components into the YOLOv11n baseline model. The corresponding results are summarized in Table 5. All experiments were conducted under identical training configurations to ensure a fair comparison. Although the baseline YOLOv11n model already achieves relatively high overall performance, achieving a precision of 0.933, a recall of 0.839, an mAP@0.50 of 0.918, and an mAP@0.50:0.95 of 0.703, its capability remains limited when detecting small, dense, and highly overlapping tea buds.
Table 5. Results of ablation experiments.
After integrating the C3k2_RFCAConv module, the recall markedly increased from 0.839 to 0.889, and mAP@0.50:0.95 improved from 0.703 to 0.787, indicating fewer missed detections and more accurate localization. Precision and mAP@0.50 also rose from 0.933 to 0.942 and from 0.918 to 0.945, respectively. This indicates that the C3k2_RFCAConv module effectively enhances feature representation and discriminative capability. When only the ADown module was introduced, the model achieved a precision of 0.941, an mAP@0.5 of 0.934, a recall of 0.862, and an mAP@0.50:0.95 of 0.740. In addition, both the GFLOPs and parameter count were reduced by 15.9% and 18.6%, respectively. With SPPF-LSKA alone, the model reached a precision of 0.944, which is higher than that of the baseline as well as the configurations incorporating either ADown or C3k2_RFCAConv, with a slight increase in parameters. C3k2_RFCAConv combined with SPPF-LSKA achieved the highest precision of 0.961, with a corresponding recall of 0.890. Adown combined with SPPF-LSKA resulted in a precision of 0.950, a recall of 0.881, and an mAP@0.50 of 0.947.
Upon integrating all three modules, the model achieved its best performance. Compared with the baseline YOLOv11n, precision, recall, mAP@0.50, and mAP@0.50:0.95 improved by 2.79%, 8.22%, 4.68%, and 15.8%, respectively, while the number of parameters and GFLOPs were reduced. Notably, the substantial improvement in mAP@0.50:0.95 reflects enhanced localization performance under stricter evaluation criteria.
The ablation study confirms the synergistic effect of the C3k2_RFCAConv ADown and SPPF-LSKA modules. The incorporation of the C3k2_RFCAConv module strengthens fine-grained feature extraction and directional sensitivity, improving the detection of small and densely packed tea buds. The ADown module enhances model compactness and computational efficiency through an adaptive downsampling strategy, maintaining accuracy. The SPPF-LSKA module improves multi-scale feature representation and contextual information aggregation, allowing more accurate recognition of overlapping tea buds, with a slight increase in parameters and GFLOPs. In summary, C3k2_RFCAConv strengthens feature extraction and localization for small and dense tea buds; ADown compresses the model through an efficient downsampling strategy, maintaining a balance between accuracy and efficiency; and SPPF-LSKA enhances multi-scale feature representation and contextual information aggregation, improving the recognition of overlapping tea buds, with a slight increase in parameters and computational cost. Collectively, these three modules increase detection precision and overall performance while reducing computational complexity, demonstrating that the proposed architecture effectively balances accuracy and efficiency for dense agricultural object detection.

4.4. Comparative Experiments

In order to evaluate the improved performance of the YOLO-LAR model, this study compared it with five representative methods, including YOLOv5n [45], YOLOv8n [46], YOLOv9t [47], YOLOv10n [48], and YOLOv11n. The detailed results are summarized in Table 6. Among the YOLO series models, YOLOv11n and YOLOv9t showed relatively limited detection capability. YOLOv11n achieved a precision of 0.933, a recall of 0.839, an mAP@0.50 of 0.918, and an mAP@0.50:0.95 of 0.703, indicating constraints in detecting small and densely clustered tea buds. YOLOv9t, while being lightweight with fewer than 2 million parameters, reached a higher recall of 0.892 and mAP@0.50:0.95 of 0.779, resulting in slightly lower precision but higher mAP values compared with YOLOv11n. Models with intermediate performance, including YOLOv5n, YOLOv8n, and YOLOv10n, achieved higher precision and mAP@0.50 values, ranging from 0.942 to 0.958 and 0.943 to 0.949, respectively. These models exhibited more balanced detection accuracy, but they required larger computational resources, with GFLOPs ranging from 6.5 to 8.1 and parameters between 2.25M and 3.01M, which may limit real-time applicability for large-scale tea bud detection. In comparison, YOLO-LAR consistently outperformed all other models across all evaluation metrics. It achieved high precision and recall, with mAP@0.50 of 0.96 and mAP@0.50:0.95 of 0.81, while maintaining a moderate model size and the lowest GFLOPs among the tested models. The larger improvement under stricter IoU thresholds indicates enhanced localization precision rather than merely increased detection confidence.
Table 6. Performance comparison of different models.
The results suggested that the relatively limited detection capability of YOLOv11n and YOLOv9t was mainly due to their architectural constraints [49,50]. YOLOv11n, despite having a recent backbone, lacked sufficient fine-grained feature extraction for small and densely clustered tea buds, leading to moderate recall and localization performance. YOLOv9t, with a simplified backbone and fewer convolutional layers, sacrificed detailed feature representation for efficiency, which led to lower localization precision despite a slightly higher recall. The intermediate-performance models, including YOLOv5n, YOLOv8n, and YOLOv10n, employed deeper backbones and more feature aggregation layers, which enhanced detection accuracy compared to the lower-performance models. However, the larger number of convolutional layers and expanded feature maps increased computational requirements [51], which may limit real-time application in large-scale tea bud detection. In contrast, YOLO-LAR integrates advanced feature refinement, downsampling, and contextual modules on top of the YOLOv11n backbone, resulting in improved fine-grained feature representation and multi-scale perception. These findings indicate that YOLO-LAR is highly suitable for accurate and efficient tea bud detection in complex tea plantations. Going beyond the aggregate performance metrics summarized in Table 6, Figure 12 provides a more detailed analysis of the classification behavior of YOLO-LAR using the F1–confidence curve and the confusion matrix.
Figure 12. Classification performance of the proposed YOLO-LAR model: (a) F1–confidence curve; (b) normalized confusion matrix on the test set (Cohen’s kappa = 0.84).
To further verify the performance of the YOLO-LAR model in the tea bud detection task, images of tea buds under various environmental conditions were selected for visual comparison and analysis. As shown in Figure 13, under sunny, close-range conditions, both YOLOv11n and YOLO-LAR successfully detected all visible tea buds. In sunny, long-range scenes, YOLOv11n missed approximately 6 tea buds, whereas the improved YOLO-LAR reduced the number of missed detections to about 4. Under cloudy, close-range conditions, neither model exhibited significant missed detections; however, in the more challenging cloudy, long-range scenes, YOLOv11n missed approximately 7 tea buds, while YOLO-LAR missed only around 4. Overall, the improved YOLO-LAR model exhibits stronger robustness under complex lighting and long-distance conditions.
Figure 13. Tea bud detection results using YOLOv11n and YOLO-LAR under sunny and cloudy conditions.

4.5. Limitations and Perspectives

Although the dataset was collected from two representative tea plantations, the large-leaf morphology and dense canopy structure introduce substantial intra-class variability, partially mitigating dataset scale limitations. The present study proposes YOLO-LAR, a YOLOv11n-based model designed for accurate and real-time detection of tea buds in complex large-leaf tea plantations. A key contribution of this work lies in the construction of a dedicated large-leaf tea bud dataset, originating from mountainous tea plantations where cultivars exhibit pronounced morphological variability and dense foliage. Compared with previous tea bud datasets, which primarily focus on small-leaf cultivars in plains or gentle hills [7,52], the large-leaf tea bud dataset presented in this study complements these resources by better reflecting real-field complexities, supporting the development of robust detection models capable of handling high-density, fine-grained tea buds. It is worth noting that, to date, there is no unified publicly available benchmark dataset for tea bud detection [53]. As reported in previous studies, most existing methods rely on self-collected datasets under specific cultivation conditions, reflecting a common limitation in this research field. Therefore, dataset dependence remains an inherent challenge rather than a limitation unique to this study.
Based on this comprehensive dataset, YOLO-LAR demonstrates significant methodological advantages for small-target detection, effectively handling fine-grained and densely distributed tea buds under complex field conditions. Similar challenges in small-object detection have been addressed in other agricultural applications, such as weed identification [54] and citrus fruit detection [55], which provide useful insights and methodological references for developing robust tea bud detection methods.
Despite these advancements, current research often relies on combining multiple modules and techniques to boost model performance, which typically demands extensive trial-and-error and depends heavily on heuristic design choices. The proposed YOLO-LAR model is specifically designed for dense, small tea buds in mountainous large-leaf plantations. Its generalization to other regions, cultivars, or acquisition platforms has not yet been systematically validated and remains an open challenge. Nevertheless, the consistent performance improvements observed across multiple ablation settings and stricter IoU thresholds suggest that the gains are not solely attributable to overfitting to a specific dataset. This can lead to increasingly complex architectures that are difficult to interpret [9]. At the same time, detection accuracy remains sensitive to environmental factors such as variable illumination, leaf occlusion, and wind-induced motion. In addition, although the model performs well in terms of detection, it still lacks capabilities for counting or grading tea buds, which are important for yield estimation and management. While full iterative runs are not provided due to computational constraints, the experiments were strictly controlled for reproducibility, and more extensive statistical analysis will be conducted in future work.
In future studies, insights from emerging multi-modal research could be leveraged to enable more systematic integration of complementary information, fostering the development of efficient, interpretable, and generalizable detection models. Moreover, the generalization of YOLO-LAR will be systematically evaluated through multi-season data collection, cross-region validation, and testing on different tea cultivars and acquisition platforms, aiming to assess and improve robustness under domain shifts. In addition, multi-season data augmentation, temporal modeling, and multi-modal inputs such as RGB or infrared images should be incorporated to enhance robustness under dynamic field conditions. Finally, extending the model to enable tea bud counting and quality grading would provide critical information for yield estimation and intelligent plantation management.

5. Conclusions

This study proposes YOLO-LAR, a YOLOv11n-based framework designed for detecting small, densely distributed, and visually ambiguous tea buds. By integrating the C3k2_RFCAConv, ADown, and SPPF-LSKA modules, the model improves fine-grained feature extraction, downsampling efficiency, and multi-scale contextual representation, achieving high detection accuracy while maintaining computational efficiency in complex environments. Although YOLO-LAR outperforms existing YOLO variants, its robustness under multi-season conditions, varying illumination, and diverse canopy structures still requires further evaluation. Future work will explore multi-modal extensions and incorporate tea bud counting and quality assessment to support intelligent harvesting and sustainable agriculture.

Author Contributions

J.H.: writing—original draft, visualization, software, methodology, conceptualization. E.W.: validation, investigation. Y.L.: visualization, software. N.L.: methodology, conceptualization. W.X.: writing—review and editing, supervision, funding acquisition. L.W.: writing—review and editing, visualization, funding acquisition. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported in part by research grants from the National Natural Science Foundation of China (32360387, 32060320, 32160368); the“Ten Thousand Talents Program” Special Project for Young Top-notch Talents of Yunnan Province (YNWR-QNBJ-2020047); the Basic Research Program Key Project of Yunnan Province, China (202501AS070090). We thank the anonymous reviewers for their constructive comments on the earlier version of the manuscript.

Institutional Review Board Statement

Ethical review and approval were waived for this study because it did not involve human participants or animals.

Data Availability Statement

The data supporting the findings of this study are available from the corresponding author upon reasonable request. Due to the ongoing nature of this research project and to preserve the integrity of subsequent studies, the dataset is not publicly available at this stage.

Conflicts of Interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

References

  1. Kagira, E.K.; Kimani, S.W.; Githii, K.S. Sustainable methods of addressing challenges facing small holder tea sector in Kenya: A supply chain management approach. J. Mgmt. Sustain. 2012, 2, 75. [Google Scholar] [CrossRef] [Scilit]
  2. Thushara, S. Sri Lankan tea industry: Prospects and challenges. In Proceedings of the Proceedings of the Second Middle East Conference on Global Business, Economics, Finance and Banking, Dubai, United Arab Emirates, 22–24 May 2015. [Google Scholar]
  3. Koech, R.K. Tea Industry and Economic Development in Kenya: Export–Growth Linkages, Household Welfare, and Macroeconomic Challenges. EPRA Int. J. Econ. Bus. Manag. Stud. (EBMS) 2025, 12, 15–21. [Google Scholar]
  4. Hang, Z.; Tong, F.; Xianglei, X.; Yunxiang, Y.; Guohong, Y. Research status and prospect of tea mechanized picking technology. J. Chin. Agric. Mech. 2023, 44, 28. [Google Scholar]
  5. Zhang, W.; Chen, Y.; Wang, Q.; Chen, J. Researches on the tender leaf identification and mechanically perceptible plucking finger for high-quality green tea. J. Sci. Food Agric. 2025, 105, 2169–2178. [Google Scholar] [CrossRef] [Scilit]
  6. Wang, J.; Guo, Y.; Han, Y.; Zhu, Y.; Lan, Y. A top-view transformer-based framework for tiny tea bud detection with gradient-guided attention and position-sensitive loss. Comput. Electron. Agric. 2026, 240, 111123. [Google Scholar] [CrossRef] [Scilit]
  7. Yang, Q.; Gu, J.; Xiong, T.; Wang, Q.; Huang, J.; Xi, Y.; Shen, Z. RFA-YOLOv8: A Robust Tea Bud Detection Model with Adaptive Illumination Enhancement for Complex Orchard Environments. Agriculture 2025, 15, 1982. [Google Scholar] [CrossRef] [Scilit]
  8. Li, H.; Kong, M.; Shi, Y. Tea Bud Detection Model in a Real Picking Environment Based on an Improved YOLOv5. Biomimetics 2024, 9, 692. [Google Scholar] [CrossRef] [Scilit]
  9. Wu, Y.; Chen, J.; Wu, S.; Li, H.; He, L.; Zhao, R.; Wu, C. An improved YOLOv7 network using RGB-D multi-modal feature fusion for tea shoots detection. Comput. Electron. Agric. 2024, 216, 108541. [Google Scholar] [CrossRef] [Scilit]
  10. Abhiram, G. The evolution of tea harvesting: A comprehensive review of machinery and technological advancements. J. Biosyst. Eng. 2024, 49, 346–367. [Google Scholar] [CrossRef] [Scilit]
  11. Wang, Z.; Niu, H.; Zhang, J.; Zhang, W.; Mao, J.; Zhang, S.; Liu, J.; Zuo, G.; Zheng, Z.; Chi, Z. Automated tea shoot picking using the YOLO network and Mamba images segmentation for top-view detection with a monocular camera. J. Agric. Eng. 2025. [Google Scholar] [CrossRef] [Scilit]
  12. Zhou, J.; Gao, S.; Du, Z.; Xu, T.; Zheng, C.; Liu, Y. The Impact of Harvesting Mechanization on Oolong Tea Quality. Plants 2024, 13, 552. [Google Scholar] [CrossRef] [Scilit]
  13. Xie, S.; Sun, H. Tea-YOLOv8s: A tea bud detection model based on deep learning and computer vision. Sensors 2023, 23, 6576. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  14. Tang, Y.; Han, W.; Hu, A.; Wang, W. Design and experiment of intelligentized tea-plucking machine for human riding based on machine vision. Nongye Jixie Xuebao/Trans. Chin. Soc. Agric. Mach. 2016, 47, 15–20. [Google Scholar]
  15. Zhang, L.; Zhang, H.; Chen, Y.; Dai, S.; Li, X.; Kenji, I.; Liu, Z.; Li, M. Real-time monitoring of optimum timing for harvesting fresh tea leaves based on machine vision. Int. J. Agric. Biol. Eng. 2019, 12, 6–9. [Google Scholar] [CrossRef] [Scilit]
  16. Karunasena, G.; Priyankara, H. Tea bud leaf identification by using machine learning and image processing techniques. Int. J. Sci. Eng. Res. 2020, 11, 624–628. [Google Scholar] [CrossRef] [Scilit]
  17. Zhang, L.; Zou, L.; Wu, C.; Jia, J.; Chen, J. Method of famous tea sprout identification and segmentation based on improved watershed algorithm. Comput. Electron. Agric. 2021, 184, 106108. [Google Scholar] [CrossRef] [Scilit]
  18. Dalal, M.; Mittal, P. A Systematic Review of Deep Learning-Based Object Detection in Agriculture: Methods, Challenges, and Future Directions. Comput. Mater. Contin. 2025, 84, 57–91. [Google Scholar] [CrossRef] [Scilit]
  19. Xiao, F.; Wang, H.; Xu, Y.; Zhang, R. Fruit detection and recognition based on deep learning for automatic harvesting: An overview and review. Agronomy 2023, 13, 1625. [Google Scholar] [CrossRef] [Scilit]
  20. Dhiman, P.; Kukreja, V.; Manoharan, P.; Kaur, A.; Kamruzzaman, M.; Dhaou, I.B.; Iwendi, C. A novel deep learning model for detection of severity level of the disease in citrus fruits. Electronics 2022, 11, 495. [Google Scholar] [CrossRef] [Scilit]
  21. Shin, J.; Chang, Y.K.; Heung, B.; Nguyen-Quang, T.; Price, G.W.; Al-Mallahi, A. A deep learning approach for RGB image-based powdery mildew disease detection on strawberry leaves. Comput. Electron. Agric. 2021, 183, 106042. [Google Scholar] [CrossRef] [Scilit]
  22. Xiong, J.; Liu, Z.; Chen, S.; Liu, B.; Zheng, Z.; Zhong, Z.; Yang, Z.; Peng, H. Visual detection of green mangoes by an unmanned aerial vehicle in orchards based on a deep learning method. Biosyst. Eng. 2020, 194, 261–272. [Google Scholar] [CrossRef] [Scilit]
  23. Barboza, T.O.; Santos, A.F.d.; Bedwell, E.K.; Vellidis, G.; Lacerda, L.N. Corn Plant Detection Using YOLOv9 Across Different Soil Background Colors, Growth Stages, and UAV Flight Heights. Remote Sens. 2025, 18, 14. [Google Scholar] [CrossRef] [Scilit]
  24. Du, Y.; Han, Y.; Su, Y.; Wang, J. A lightweight model based on you only look once for pomegranate before fruit thinning in complex environment. Eng. Appl. Artif. Intell. 2024, 137, 109123. [Google Scholar] [CrossRef] [Scilit]
  25. Yao, J.; Li, Y.; Xia, Z.; Nie, P.; Li, X.; Li, Z. Wtad-yolo: A lightweight tomato leaf disease detection model based on yolo11. Smart Agric. Technol. 2025, 12, 101349. [Google Scholar] [CrossRef] [Scilit]
  26. Wang, M.; Li, Y.; Meng, H.; Chen, Z.; Gui, Z.; Li, Y.; Dong, C. Small target tea bud detection based on improved YOLOv5 in complex background. Front. Plant Sci. 2024, 15, 1393138. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  27. Guo, S.; Yoon, S.-C.; Li, L.; Wang, W.; Zhuang, H.; Wei, C.; Liu, Y.; Li, Y. Recognition and positioning of fresh tea buds using YOLOv4-lighted+ ICBAM model and RGB-D sensing. Agriculture 2023, 13, 518. [Google Scholar] [CrossRef] [Scilit]
  28. Chen, C.; Lu, J.; Zhou, M.; Yi, J.; Liao, M.; Gao, Z. A YOLOv3-based computer vision system for identification of tea buds and the picking point. Comput. Electron. Agric. 2022, 198, 107116. [Google Scholar] [CrossRef] [Scilit]
  29. He, L.-h.; Zhou, Y.-z.; Liu, L.; Cao, W.; Ma, J.-h. Research on object detection and recognition in remote sensing images based on YOLOv11. Sci. Rep. 2025, 15, 14032. [Google Scholar] [CrossRef] [Scilit]
  30. Niu, P.; Gu, J.; Zhang, Y.; Zhang, P.; Cai, T.; Xu, W.; Han, J. MDCGA-Net: Multiscale direction Context-Aware network with global attention for Building extraction from remote sensing images. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2024, 17, 8461–8476. [Google Scholar] [CrossRef] [Scilit]
  31. Zhang, X.; Liu, C.; Yang, D.; Song, T.; Ye, Y.; Li, K.; Song, Y. RFAConv: Innovating spatial attention and standard convolutional operation. arXiv 2023, arXiv:2304.03198. [Google Scholar]
  32. Hou, Q.; Zhou, D.; Feng, J. Coordinate attention for efficient mobile network design. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 19–25 June 2021; pp. 13713–13722. [Google Scholar]
  33. Lau, K.W.; Po, L.-M.; Rehman, Y.A.U. Large separable kernel attention: Rethinking the large kernel attention design in cnn. Expert Syst. Appl. 2024, 236, 121352. [Google Scholar] [CrossRef] [Scilit]
  34. Ouyang, D.; He, S.; Zhang, G.; Luo, M.; Guo, H.; Zhan, J.; Huang, Z. Efficient multi-scale attention module with cross-spatial learning. In Proceedings of the ICASSP 2023—2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, 4–10 June 2023; pp. 1–5. [Google Scholar]
  35. Yu, W.; Si, C.; Zhou, P.; Luo, M.; Zhou, Y.; Feng, J.; Yan, S.; Wang, X. Metaformer baselines for vision. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 46, 896–912. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  36. Xia, Z.; Pan, X.; Song, S.; Li, L.E.; Huang, G. Vision transformer with deformable attention. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 4794–4803. [Google Scholar]
  37. Qi, X.; Xu, C.; Liu, Y.; Ma, N.; Liu, H. TPGR-YOLO: Improving the Traffic Police Gesture Recognition Method of YOLOv11. Int. J. Adv. Comput. Sci. Appl. 2025, 16, 404–415. [Google Scholar] [CrossRef] [Scilit]
  38. Chollet, F. Xception: Deep learning with depthwise separable convolutions. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2017; pp. 1251–1258. [Google Scholar]
  39. Zhang, X.; Song, Y.; Song, T.; Yang, D.; Ye, Y.; Zhou, J.; Zhang, L. LDConv: Linear deformable convolution for improving convolutional neural networks. Image Vis. Comput. 2024, 149, 105190. [Google Scholar] [CrossRef] [Scilit]
  40. Yang, G.; Wang, Y.; Shi, D.; Wang, Y. Golden Cudgel Network for Real-Time Semantic Segmentation. In Proceedings of the Computer Vision and Pattern Recognition Conference, Nashville, TN, USA, 11–15 June 2025; pp. 25367–25376. [Google Scholar]
  41. Cai, X.; Lai, Q.; Wang, Y.; Wang, W.; Sun, Z.; Yao, Y. Poly kernel inception network for remote sensing detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 27706–27716. [Google Scholar]
  42. Yang, L.; Zhang, R.-Y.; Li, L.; Xie, X. Simam: A simple, parameter-free attention module for convolutional neural networks. In Proceedings of the International Conference on Machine Learning, Virtual, 18–24 July 2021; pp. 11863–11874. [Google Scholar]
  43. Wan, D.; Lu, R.; Shen, S.; Xu, T.; Lang, X.; Ren, Z. Mixed local channel attention for object detection. Eng. Appl. Artif. Intell. 2023, 123, 106442. [Google Scholar] [CrossRef] [Scilit]
  44. Zhao, L.; Liang, G.; Hu, Y.; Xi, Y.; Ning, F.; He, Z. YOLO-RLDW: An algorithm for object detection in aerial images under complex backgrounds. IEEE Access 2024, 12, 128677–128693. [Google Scholar] [CrossRef] [Scilit]
  45. Jocher, G.; Chaurasia, A.; Stoken, A.; Borovec, J.; Kwon, Y.; Michael, K.; Fang, J.; Wong, C.; Yifu, Z.; Montes, D. Ultralytics/Yolov5: V6. 2-Yolov5 Classification Models, Apple m1, Reproducibility, Clearml and Deci.ai Integrations; Zenodo: Geneva, Switzerland, 2022. [Google Scholar]
  46. Varghese, R.; Sambath, M. Yolov8: A novel object detection algorithm with enhanced performance and robustness. In Proceedings of the 2024 International Conference on Advances in Data Engineering and Intelligent Computing Systems (ADICS), Chennai, India, 18–19 April 2024; pp. 1–6. [Google Scholar]
  47. Wang, C.-Y.; Yeh, I.-H.; Mark Liao, H.-Y. Yolov9: Learning what you want to learn using programmable gradient information. In Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September–4 October 2024; pp. 1–21. [Google Scholar]
  48. Wang, A.; Chen, H.; Liu, L.; Chen, K.; Lin, Z.; Han, J. Yolov10: Real-time end-to-end object detection. Adv. Neural Inf. Process. Syst. 2024, 37, 107984–108011. [Google Scholar]
  49. Sapkota, R.; Flores-Calero, M.; Qureshi, R.; Badgujar, C.; Nepal, U.; Poulose, A.; Zeno, P.; Vaddevolu, U.B.P.; Khan, S.; Shoman, M. YOLO advances to its genesis: A decadal and comprehensive review of the You Only Look Once (YOLO) series. Artif. Intell. Rev. 2025, 58, 274. [Google Scholar] [CrossRef] [Scilit]
  50. Ma, Y.; Xi, C.; Ma, T.; Sun, H.; Lu, H.; Xu, X.; Xu, C. I-YOLOv11n: A lightweight and efficient small target detection framework for UAV aerial images. Sensors 2025, 25, 4857. [Google Scholar] [CrossRef] [Scilit]
  51. Ali, M.L.; Zhang, Z. The YOLO framework: A comprehensive review of evolution, applications, and benchmarks in object detection. Computers 2024, 13, 336. [Google Scholar] [CrossRef] [Scilit]
  52. Zhang, K.; Yuan, B.; Cui, J.; Liu, Y.; Zhao, L.; Zhao, H.; Chen, S. Lightweight tea bud detection method based on improved YOLOv5. Sci. Rep. 2024, 14, 31168. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  53. Liu, Z.; Zhuo, L.; Dong, C.; Li, J. YOLO-TBD: Tea bud detection with triple-branch attention mechanism and self-correction group convolution. Ind. Crops Prod. 2025, 226, 120607. [Google Scholar] [CrossRef] [Scilit]
  54. Chen, Z.; Chen, B.; Huang, Y.; Zhou, Z. GE-YOLO for Weed Detection in Rice Paddy Fields. Appl. Sci. 2025, 15, 2823. [Google Scholar] [CrossRef] [Scilit]
  55. Liao, Y.; Li, L.; Xiao, H.; Xu, F.; Shan, B.; Yin, H. YOLO-MECD: Citrus detection algorithm based on YOLOv11. Agronomy 2025, 15, 687. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.