1. Introduction
Citrus is an alternate-bearing crop, and its floral bud density is widely recognised as a critical determinant of annual fruit yield [
1]. Earlier estimation of citrus bud abundance enables more timely adjustment of orchard management and production planning [
2]. During the first three years after transplanting, citrus saplings should be managed to prioritize canopy development while minimizing or suppressing flowering. Once entering the mature fruiting stage, flowering intensity should be reduced for weakly growing trees and promoted for vigorous individuals to stabilize annual yield. Consequently, precise regulation of bud load is essential for economically sustainable citrus production. However, a single mature citrus tree can produce thousands of individual buds, posing substantial challenges to accurate quantification. Neither manual counting nor conventional machine vision approaches can achieve reliable and consistent estimation [
3]. Current practice still relies heavily on visual inspection and experiential estimation by growers. Such conventional methods are labor-intensive, time-consuming, costly, and highly subjective, resulting in low efficiency and poor scalability for large-scale orchard management [
4]. With the accelerating transition toward intelligent and automated citrus cultivation, automated and accurate citrus bud estimation techniques are urgently needed [
5]. The development of intelligent vision-based estimation systems to replace manual counting is therefore critically needed to reduce labor input and improve operational efficiency in modern citrus production.
As a core technique for intelligent estimation robots, accurate and efficient detection algorithms are essential to improve estimation reliability and provide a solid technical foundation for practical deployment. Current citrus yield estimation approaches, which rely on color features, ultrasonic sensing, and aerial imagery, predominantly focus on fruit counting during the mature fruiting stage. For instance, a machine vision system was developed for real-time early citrus yield estimation using pixel distribution in the HSI color space, representing one of the earliest applications of computer vision in this field [
6]. To further improve estimation accuracy, an increasing number of studies have incorporated deep learning techniques into yield estimation pipelines. For example, Faster R-CNN and LSTM models have been combined with UAV imagery to detect citrus fruits and achieve early yield estimation, yielding higher accuracy than traditional vision-based methods [
7]. Another recent study integrated deep learning-based fruit detection with the XGBoost regression model, using pruning intensity and multi-view image features as inputs to improve individual-tree early yield estimation accuracy [
8]. Nevertheless, nearly all existing studies concentrate on fruit detection during the fruiting stage and cannot be directly extended to citrus bud estimation in the early flowering period.
Deep learning has been successfully applied to flower bud detection and yield estimation for other horticultural crops using RGB images, such as grapes and tea, which validates the feasibility of computer vision-based bud estimation in agricultural scenarios [
9,
10,
11]. However, citrus flower buds present distinct and more challenging characteristics compared with grape and tea buds: they are smaller in size, more densely distributed, and suffer from more severe occlusion by leaves and branches in complex orchard environments. Mutual shading and overlapping between foliage and citrus buds further escalate detection difficulty, rendering existing bud detection pipelines designed for other crops inapplicable to direct citrus bud estimation.
Recently, the YOLOv8 model has garnered widespread attention in agricultural small object detection due to its exceptional detection accuracy and inference speed, with successful applications in addressing challenges like small target size, dense distribution, and occlusion [
12]. It can effectively process high-resolution images and detect diverse object categories. For example, Gai et al. adapted YOLOv8 for blueberry detection to address challenges such as small size, dense distribution, and leaf occlusion [
13]. Similarly, to address challenges including the high color similarity between buds and background foliage, as well as occlusion caused by overlapping leaves, Liu et al. extended YOLOv8s to develop MAE-YOLOv8, which improves detection efficiency and accuracy for green crisp plums [
14]. For passion fruit, Chen et al. developed the YOLOv8-MDN-Tiny model based on YOLOv8s to enhance the detection of small-scale diseases and overcome limitations of conventional detection models [
15]. Notably, Lu et al. incorporated the lightweight CARAFE operator and a multi-efficient channel attention mechanism into YOLOv8, realizing fast and accurate small-object detection; this capability is particularly critical for UAV-based citrus orchard monitoring and highly relevant to the task of citrus bud estimation [
16].
Accurate citrus bud detection is crucial for refining agricultural management practices and early yield estimation. Inspired by the aforementioned advances, we adopt YOLOv8 for the citrus bud detection task and further enhance its performance. Notably, YOLOv8 offers five variants (n, s, m, l, and x) with varying widths and depths. Given the limited computational resources in real orchard environments, we select the lightweight YOLOv8n as the baseline model, due to its minimal model size and lowest computational complexity. However, accurate detection of citrus buds still faces two key challenges. First, citrus buds exhibit clustered growth patterns and are frequently occluded by surrounding foliage, which readily gives rise to missed detections. Second, natural citrus orchards feature complex backgrounds accompanied by variable illumination, shadow interference, lime coatings, and diverse vegetation, all of which introduce substantial background noise and interferences, impeding accurate citrus bud detection.
To address the challenges involved in citrus bud detection within complex orchard scenarios, this study presents a dedicated contextual and frequency-aware detector named CFADet to achieve accurate and reliable bud estimation. Specifically, a diverse citrus bud dataset is constructed under varied illumination and occlusion conditions. Moreover, spatial-to-depth convolution (SPDConv) is employed to reconstruct the network backbone, which effectively enhances fine-grained feature extraction for these small and densely clustered targets. To enhance the robustness of information transmission for small targets, an enhanced feature fusion network (EFF) embedded with a novel dual-path weighted fusion strategy is proposed to adaptively fuse multi-scale features from different network depths while integrating intermediate backbone layers as well as skip connections to better focus on tiny-scale characteristics. Meanwhile, a contextual boundary enhancement module (CBEM) based on dimension interaction and max-pooling is designed to highlight edge and contour information of citrus buds to make them more distinguishable from cluttered backgrounds. Furthermore, a frequency-aware module (FAM) with a specially designed attention weight filter is developed to adaptively regulate frequency components and suppress complex background noise, which effectively reduces missed detections and localization errors caused by bud overlap and leaf occlusion while improving the model’s perception capability toward citrus buds of varying sizes.
The rest of this paper is structured as follows.
Section 2 describes the self-constructed dataset and the proposed citrus bud detection model.
Section 3 presents the experimental results with detailed analysis.
Section 4 discusses the limitations of our approach and outlines potential directions for future work. Finally,
Section 5 summarizes the research.
2. Materials and Methods
2.1. Image Acquisition
In this study, all citrus bud images were collected at the Citrus Planting Base of Guangxi Shanshui Nonggang on 26 February 2024, between 9:00 a.m. and 6:00 p.m., with approximately one frame captured per minute per operator. The geographical coordinates of the base are
E longitude and
N latitude (WGS84 datum), as shown in
Figure 1a. During the acquisition period, weather conditions varied from clear and cloudy to overcast with light rain, resulting in naturally fluctuating illumination levels that reflect real orchard environments. The sampled citrus cultivar was Orah Mandarin (
Citrus reticulata cv. Orah). Image data were acquired using two handheld devices: a smartphone (iPhone 13, Apple Inc., Cupertino, CA, USA) and a digital camera (IXUS 125 HS, Canon Inc., Tokyo, Japan). To ensure data consistency, all devices operated under automatic ISO settings. The collected images were captured at a resolution of 3024 × 4032 pixels and stored in JPG format. To enhance dataset diversity and simulate realistic field inspection scenarios, citrus buds were photographed under both frontlighting and backlighting conditions at distances ranging from 20 cm to 80 cm. Specifically, 326 images were captured at close range (20–40 cm), 418 at medium range (40–60 cm), and 223 at long range (60–80 cm). This stratified distance distribution ensures that the dataset covers typical viewing conditions encountered during orchard monitoring. Challenging scenarios, including partial occlusion, leaf shading, and densely clustered buds, were deliberately recorded to improve dataset robustness, as illustrated in
Figure 1b,c. After quality screening based on image clarity, completeness, and annotation feasibility, a total of 967 high-resolution images were retained as the final dataset. Images with severe blurring, overexposure, excessive occlusion, or incomplete bud structures were removed to ensure dataset reliability and annotation accuracy.
2.2. Data Preprocessing
To ensure high-quality training data for accurate citrus bud detection, a series of data preprocessing steps were conducted, including manual labeling, dataset splitting, and data augmentation.
During manual annotation, citrus buds in the images were manually annotated using LabelImg (Tzutalin, Redwood City, CA, USA; Version 1.8.6). Precise rectangular bounding boxes were carefully drawn around each individual citrus bud to tightly align with its actual contour. For targets occluded by branches, leaves, or adjacent buds, annotators delineated the bounding boxes based on the typical morphological characteristics of Orah Mandarin (Citrus reticulata cv. Orah) buds to accurately reflect the true physical size of each target. The target buds were labeled into two categories: “green” and “white”, excluding extremely tiny background points smaller than 10 × 10 pixels (approximately less than 0.001% of the total image area). Such ultra-small targets suffer from optical blurring and pixel aliasing, resulting in incomplete morphological contours and introducing unwanted noise into the dataset, which further increases the computational burden of the model. Furthermore, such miniature distant buds are predominantly immature latent buds or non-productive distal shoots with an extremely low fruit-set rate, contributing negligible value to current-season yield formation and thus being irrelevant for practical yield estimation. This exclusion only eliminates unidentifiable ultra-small background noise while retaining all valid foreground buds, thereby imposing minimal impact on the model’s generalization ability and ensuring high-quality annotation and efficient inference for real-world orchard deployment.
Following standardized bud classification criteria [
17], white-green flower buds with obvious sepals were labeled as green buds, and those with unclear sepals and completely white bodies as white buds (
Figure 1c). To minimize labeling errors caused by the small target size, the annotation process was divided into three sequential stages: initial labeling according to classification criteria, error verification, and final review by professional agricultural experts to ensure labeling consistency and standardization. All annotation files were exported and stored in PASCAL VOC format.
Subsequently, 97 samples with less than 30% occluded buds were chosen to form the easy test set A. Another 97 samples with more than 30% occluded buds and relatively dense distribution were selected to create the challenging test set B. In addition, a combined test set AB was generated by merging sets A and B, covering a full gradient from simple to complex detection environments, which is highly consistent with real-world citrus bud detection tasks (e.g., orchard monitoring under natural illumination and post-lime-spraying conditions). This design ensures data authenticity and relevance to practical applications. After creating the test sets, the remaining 773 images were randomly assigned to the training set, and 97 images to the validation set, resulting in a training:validation:test ratio of 7:1:2 (676:97:194).
Adequate sample diversity is critical for robust training of deep neural networks. Accordingly, an online data augmentation strategy integrated within the training framework was adopted to expand the effective training sample space without modifying the original training dataset. This strategy performs real-time augmentation operations during model training, with specific techniques including random blurring, median blurring, contrast-limited adaptive histogram equalization (CLAHE), random flipping, MixUp, and Mosaic augmentation. These transformations simulate common field imaging degradations and enrich sample diversity, effectively enhancing the model’s robustness to lighting variations, occlusions, and bud posture differences [
12].
Beyond these augmentation operations, uniform resizing of input images to 640 × 640 pixels is another key preprocessing step. For the majority of citrus buds in our dataset, this resolution retains sufficient pixel features to guarantee reliable detection performance. However, feature loss is inevitable for extremely tiny buds after downsampling, which constitutes the core challenge addressed by all the proposed methods in this study. In addition, the selection of 640 × 640 resolution is a deliberate trade-off between detection accuracy and inference speed, which meets the real-time deployment requirements of practical orchard scenarios.
2.3. Data Analysis
Target size was defined based on the relative ratio of the target to the image area in previous research [
18]. A target was categorized as a small target when the median ratio of the bounding box area to the image area fell within the range of 0.08% to 0.58%, as specified in Equation (
1):
where
represents the median value of the set
, which in this case is the set of all ratios
across all bounding boxes in the dataset,
represents the area of the
j-th bounding box within the
i-th image, and
represents the area of the
i-th image.
In this study, all annotation coordinates in the dataset were normalized to a unified coordinate system. As shown in
Figure 2, the dataset contains a substantial number of small targets. This high proportion aligns with the visual distribution in the figure, confirming that small targets dominate the dataset. For further quantitative characterization and reproducibility, the average number of citrus buds per image was calculated as 36.82 ± 6.94 (mean ± standard deviation), indicating that citrus buds are numerous and densely distributed throughout the dataset. The dense and overlapping spatial distribution of green and white buds may lead to model confusion between adjacent targets—that is, failure to distinguish whether two bounding boxes correspond to separate buds—thus hindering high-precision target localization. Moreover, the limited number of pixels of these small targets in the feature map makes it challenging to extract robust features and preserve adequate spatial information as the number of network layers increases. These characteristics—small target dominance, limited pixel information, and dense overlapping distribution—collectively highlight the uniqueness of citrus bud detection compared to conventional object detection tasks, necessitating specialized model design.
2.4. CFADet
Given that most citrus flower buds are defined as small targets (as specified in
Figure 2) and are often accompanied by complex foreground issues (e.g., overlapping with leaves, branches, or adjacent buds) in orchard scenarios, their feature representations tend to be weak. To address this issue, this study first improved YOLOv8 by introducing dual-path weighted feature fusion (DPF), cross-layer connections, and a small-target detection layer to enhance the robustness of small-target information flow, thus forming CFADet. However, it is important to note that CFADet still inherits and exhibits key limitations of the original YOLOv8 architecture, which hinder further performance gains. The employment of conventional strided convolutions in backbone networks may lead to the loss of fine-grained information, thereby affecting the model’s feature extraction capability, particularly for small objects. Secondly, the feature extraction via downsampling and pooling blurs the boundary details of small objects, weakening their distinguishability from complex backgrounds. Thirdly, the default feature fusion strategy in YOLOv8 assigns large targets to deep, semantically rich feature maps, while small targets are mapped to shallow, lower-level feature maps. In complex orchard backgrounds, deep layers may overemphasize global contextual information (e.g., dense leaves and branches) while potentially underweighting small and densely distributed buds. This could lead to the misidentification of small targets as background noise within deep features, due to the overwhelming effect of background noise. These challenges, arising from the potential limitations of YOLOv8 in fine-grained feature retention, edge preservation, and background suppression, highlight the need for targeted architectural improvements.
To address these limitations, we integrated the following modules into CFADet, as shown in
Figure 3: (1) We replaced conventional strided convolutions with SPDConv in the backbone, as its design avoids excessive downsampling loss of fine-grained information, thereby preserving small-target details. (2) The contextual boundary enhancement module (CBEM) enhances contextual features and boundaries of small targets through dimensional interaction and max-pooling operations. (3) The frequency-aware module (FAM) performs background suppression in the frequency domain by distinguishing texture frequency differences between small buds and cluttered backgrounds. This alleviates misclassifications and missed detections of overlapping or occluded targets in complex backgrounds.
2.4.1. Preliminary Improvements to YOLOv8 for Citrus Bud Detection
In real citrus orchards, citrus buds are mostly small targets relative to the overall image dimensions, and are easily occluded by foreground elements such as branches and leaves, while complex backgrounds can obscure or confuse their detection. In this case, the features of citrus buds are prone to being overwhelmed. Existing methods, specifically the Feature Pyramid Network (FPN), attempt to retain more small-target features during multi-scale feature fusion. This is achieved by constructing a top-down pyramid that fuses high-level semantic features with low-level spatial features. Similarly, the Path Aggregation Network (PANet) was proposed to address this issue by incorporating a bottom-up path into the FPN architecture to improve the flow and fusion of low-level spatial features. This structure was adopted as the baseline feature fusion module in the original YOLOv8 [
20]. To further preserve small-target information during multi-scale feature fusion, bidirectional feature pyramid network (BiFPN) was developed. It introduces weighted bidirectional cross-scale fusion on the P3–P7 feature layers to improve the flexibility of multi-scale feature integration. However, BiFPN only implements weighted fusion at the feature level, omitting fine-grained channel-wise feature recalibration and the high-resolution shallow layers essential for detecting extremely small citrus buds. This makes it insufficient to extract the discriminative information of small targets in complex orchard backgrounds.
To address these issues, we proposed an enhanced feature fusion (EFF) network to optimize network performance and reduce detail loss. As shown in the
Figure 4, we designed a dual-path fusion to replace the simple concatenation. Specifically, for
N input feature maps
we perform feature selection at both the feature (feature weight) and channel (channel weight) levels. In the feature weight pathway, dynamic weighted allocation is applied to multi-scale feature maps via learnable weight parameters (
), facilitating feature fusion across scales. In the channel weight pathway, the channel attention mechanism strengthens the critical channel information of the individual feature maps
while suppressing redundant interference. Finally, the outputs of these pathways are fused by summation to achieve complementary feature information at different levels. The operations are calculated as follows:
where
,
and
represent the channel weight path, feature weight path and dual-path fusion, respectively.
represents a stack operation.
represents the activation function.
represents the input feature.
represents the learnable weight.
serves as a stabilizing term to prevent division by zero.
Furthermore, a detection layer is introduced as a shallow high-resolution branch to preserve fine-grained spatial details. This enables the neck to acquire richer information on small targets and renders the feature flow more robust against noise and occlusion during fusion. Additionally, cross-layer connections between C2–C4 (the backbone layers) facilitate the bidirectional propagation of shallow detail features and deep semantic features, thereby integrating original feature maps with contextual information.
Overall, these enhancements collectively strengthen the representation of small citrus bud features, mitigate detail loss during fusion and improve the robustness of small-target information flow. The enhanced feature fusion network lays a solid foundation for CFADet to detect small and occluded targets in natural orchard scenes.
2.4.2. Contextual Boundary Enhancement Module
The visual characteristics of small objects are often weak, so it is advantageous to make more precise judgements based on environmental information when detecting them. In citrus bud scenes, the appearance information of small buds is extremely limited, and making judgments based solely on their appearance is challenging. Instead, such judgments can be improved by incorporating contextual cues like pedicels (structures with more distinctive visual patterns than the buds themselves), thereby enhancing discriminability. Therefore, it is essential to utilize contextual information to guide bud detection. Current studies mainly focus on RFB-like structures, which process spatial features through convolutional branches with different kernel sizes and then perform channel-wise feature fusion via 1 × 1 convolution [
21]. However, in the context of citrus bud detection, this architecture has two critical limitations that hinder performance: (1) spatial and channel operations are computed independently without considering their interdependencies, and (2) large-kernel convolutions expand receptive fields by blending features from extensive regions, resulting in blurred fine boundaries (e.g., bud edges that are 1–2 pixels wide). To address these limitations, we propose a contextual boundary enhancement module (CBEM), as shown in
Figure 5. CBEM establishes spatial-channel dependencies through dimension exchange while reinforcing boundary features, thereby significantly improving the discriminative power of small object representations.
Specifically, the input tensor is first rotated to obtain and , while remains unchanged. This rotation enables interaction between the spatial and channel features in different dimensions. No rotation is performed for the first branch, and local spatial features are extracted directly using standard convolutions. For the latter two branches, cascaded standard convolutional operations are performed on and to capture dependencies between channel and spatial dimensions, with kernel sizes of and , respectively. This design leverages the directional sensitivity of asymmetric kernels: () kernels excel at capturing horizontal spatial correlations (e.g., lateral edges of citrus buds), while () kernels are more effective at capturing vertical dependencies (e.g., longitudinal contours). Reversing the order of these asymmetric kernels across the two branches enables the module to better model multi-orientational features. This capability is essential for distinguishing small, irregularly shaped buds from cluttered backgrounds (e.g., leaves and branches) and reduces false negatives caused by orientation ambiguity.
Additionally, dilated convolutions are added to the last two branches to expand the receptive field without reducing spatial resolution, thereby capturing broader contextual information. After dimension interaction, the latter two branches are rotated back to their initial shape
. Finally, the outputs of the three branches are concatenated to maximize feature information retention and ensure the model captures as much detail as possible. The mathematical expressions of the dimension interaction structure can be written as follows:
where
,
, and
represent standard convolution operations with kernel sizes of
,
, and
, respectively.
denotes the atrous convolution operation with a dilation rate of 3.
is the feature map concatenation operation.
,
, and
represent different rotation forms of input feature maps.
,
, and
represent the output feature maps of the three branches after standard and atrous convolution.
is the output feature map of the dimension interaction structure.
To avoid boundary blurring, we perform boundary enhancement following the method in [
22]. The boundary information of
is enhanced from four directions. The key to enhancing boundaries is determining whether a position is a boundary point. Suppose we want to capture the left boundary of an object in the feature map
. We determine whether there is a drastic change between a point and its neighbour to the left. Enhancement is then performed by using the rightmost point to traverse to the left, as specified in:
where
denotes the feature map for left boundary enhancement,
represents the
c-th channel of feature map
, and
represents the value at position
of the
c-th channel of the feature map
.
denotes the value at position
in the
c-th channel of
(the enhanced feature map for left boundaries). Similarly, boundary enhancement can be applied to the feature map in four directions: up, down, left, right, as shown in the right part of
Figure 5. Finally, the output of four-directional boundary enhancement and the original feature
are concatenated along the channel dimension to form
, integrating enhanced boundary details with original semantic information.
By employing dimensional exchange and multi-directional enhancement strategies, CBEM effectively captures contextual information and the boundaries of small objects. To further improve the representation of features of small targets, CBEM is integrated into the layer, which retains the most fine-grained spatial details and contains the most relevant information for small targets. The enhanced features from the layer are then fused into the layer, which is responsible for detecting small targets. Due to the computational complexity of the layer, an upsampling approach is selected for the merging process to ensure efficient integration of these features.
2.4.3. Frequency-Aware Module
Following the CBEM and EFF modules, the feature maps already incorporate local contextual information and provide an accurate representation of the characteristics of small buds. However, in real orchard environments, citrus buds may be affected by other features, such as leaves and branches covered in limes, and occlusion issues could cause the model to misclassify the background as targets, potentially leading to false alarms.
Traditional background suppression methods are dominated by channel-wise attention mechanisms, represented by the SE and ECA modules. While these approaches recalibrate the importance of channels to filter out redundant information, they ignore the distribution of spatial features and are unable to distinguish between buds and lime-covered leaves, which have similar texture and grayscale traits. To address the limitations of single-dimensional attention, SCAM—a mainstream spatial-channel joint attention paradigm—integrates dual-dimensional feature modeling to deliver more comprehensive feature optimization [
23]. Despite this advancement, SCAM remains confined to the spatial domain and cannot eliminate interference from backgrounds with highly analogous spatial characteristics. This results in unavoidable false alarms in complex orchard scenarios.
Essentially, all of the above methods rely on spatial-domain attention to analyse feature importance for background suppression. However, objects such as lime-covered leaves, which have similar texture and grayscale characteristics in the spatial domain, are difficult to distinguish using this method. To address this issue, existing studies have explored the impact of different frequency components on camouflaged target detection [
24]. For example, a frequency perception network has been proposed that can automatically separate high-frequency texture and low-frequency contour features via octave convolution. This realises coarse localisation of camouflaged objects through frequency cues [
25]. A two-stage frequency-aware framework has been developed, where frequency-domain features assist in identifying target “breakthrough points” and enhance the discriminability of low-contrast camouflaged objects from backgrounds [
26]. This frequency-based discriminability can also be leveraged in citrus bud detection, where the texture frequencies of lime-covered leaves and buds are distinct. Adjusting the frequency components in the frequency domain enables clearer separation of the target and background. Based on the above analysis, we propose the FAM module to adaptively recalibrate the responses of different frequency components, as shown in
Figure 6.
First, we use Fast Fourier Transform (FFT) to transform the input feature
from the spatial domain to the frequency domain (
). Subsequently, we split
into amplitude
and phase spectrum
P, which are derived from the magnitude and argument of the complex Fourier coefficients, respectively, as illustrated by the following equations:
where
represents the FFT operation applied to the feature.
and
are the real and imaginary parts of
, respectively. The amplitude
directly represents the strength and energy of the frequency components.
Furthermore, the amplitude component (
) obtained by the FFT contains more critical information for object detection [
27]. Consequently, this study focuses on exploring the influence of different frequency components in the amplitude spectrum while maintaining the phase spectrum unchanged. We design three strategies for modifying the amplitude, as shown in
Figure 6a–c. The first strategy employs a high-pass filter to suppress low-frequency components and retain high-frequency details. While this strategy is indeed feasible, it must be noted that the high-pass filter may result in the loss of numerous low-frequency components, which could consequently lead to a further weakening of the details pertaining to the small targets. The second strategy utilises a channel attention mechanism to adaptively enhance the high-frequency components. The third strategy constructs a low-frequency mask by iterating over each pixel in the frequency domain and calculating its distance from the center. If this distance is less than a predefined low-frequency radius, the corresponding low-frequency components are suppressed. The specific calculations for the three strategies are represented by the following formulas:
where
represents the channel attention mechanism.
represents the coordinates in the frequency domain, and
is the center of coordinates. The denominator
is the maximum distance value, used for normalization to ensure the filter values range between
.
d is the predefined low-frequency radius.
The results of these strategies are shown in
Table 1. All strategies are able to boost detection performance, but the difference between the first and third strategies is not significant. Therefore, we choose the second strategy for amplitude modification as shown in
Figure 6b. Specifically, we utilize
,
, and
depthwise (DW) convolutions to extract local features
F, thereby enhancing the perception of different receptive fields. Subsequently, adaptive average pooling captures global information of the entire feature map. A fully connected layer (FC) then reduces the dimension of the pooled features and finally generates a Sigmoid activated output. The output from the fully connected layer is then processed by an exponential function, expanding its value range from
to
. This exponential normalization makes the results more tolerant of positional errors. The process of the attention weight filter can be outlined as follows:
where
w denotes the attention weights generated to recalibrate amplitude components.
represents depthwise convolution operations with kernel sizes of
.
performs adaptive average pooling.
and
represent the fully connected layer and sigmoid activation function, respectively.
performs exponential normalization.
After enhancing the amplitude to obtain , we stack it with the original phase P to reconstruct the modified frequency-domain feature . is then converted back to the spatial domain by applying the inverse FFT. Finally, the spatial domain features and the frequency-domain enhanced features are fused to form the final output .
This approach, distinct from traditional attention mechanisms that operate solely in the RGB domain, offers a novel strategy for background suppression, particularly in challenging scenarios with occlusions and similar feature backgrounds.
2.4.4. SPDConv
In the shallow layers of the YOLOv8 structure, strided convolutions are widely utilised for the purpose of feature extraction. In most scenarios involving high-resolution images and reasonably large objects, there is redundant pixel information that allows strided convolution to skip details without significant impact. In the case of small objects, such as citrus buds, there is an absence of redundant information. This absence leads to loss of fine-grained details and poor feature learning.
To address this issue, we introduce SPDConv to reconstruct the backbone [
28]. As shown in
Figure 7, the stage comprises a spatial-to-depth layer and a feature extraction layer (consisting of
and
convolutions). This stage involves the downsampling of image features with the objective of preserving the essential information in the channels. It is followed by the execution of multiple
convolutions, which serves to further enrich the feature map. The SPDConv takes the feature map
as input and performs the spatial-to-depth downsampling. Specifically, SPDConv performs equidistant sampling on the spatial dimensions of the input features and then concatenates each sampling result along the channel dimension to obtain
. To avoid unbalanced sampling caused by strided convolutions, a
convolution is employed to decrease the channel dimension. This spatial-to-depth layer enables downsampling without losing fine-grained information—unlike strided convolutions, which skip pixels and discard critical details of small buds. Finally, a
convolution reduces the channel dimension to lower computational costs, followed by further feature extraction via
convolutions.
Overall, by integrating SPDConv with C2f in the backbone enhancement, the proposed design achieves improved feature extraction capability while simultaneously reducing model parameters. This reconstructed backbone effectively alleviates the loss of fine-grained details and enhances the learning of discriminative features, leading to a more efficient and robust feature representation.
Building on the foundational improvements from EFF, and combined with the subsequent enhancements of CBEM and FAM, these modules collectively provide CFADet with a comprehensive solution for accurate small object detection in challenging orchard environments.
2.5. Network Training and Performance Evaluation
The experimental setup comprises a 13th Gen Intel (R) Core (TM) i7-13700KF CPU, an NVIDIA GeForce RTX 4090 D (24 GB memory) GPU, and 62 GB of RAM. The operating system is Ubuntu 24.04, with Visual Studio Code (v1.86.2) as the development environment. The programming language used is Python 3.8.19, and the deep learning framework is PyTorch 1.12.0. CUDA version 12.4 and cuDNN version 8.9.5 are employed to accelerate computations. For the proposed model, the input dimensions are set to 640 × 640 × 3, and the batch size is 16. Adam optimizer is utilized with an initial learning rate of 0.01, which is dynamically adjusted using a cosine annealing schedule during training. To ensure comprehensive training, the model is trained for a maximum of 1000 epochs, with an early stopping mechanism (patience = 50 epochs) to prevent overfitting.
This study employed precision (P), recall (R), F1-score, and average precision (AP) for individual categories as the key metrics to evaluate the model’s performance. In binary classification tasks, samples are classified into four categories according to whether their actual labels match the model’s predictions: true positives (TP), false positives (FP), true negatives (TN), and false negatives (FN). The TP and FP values are sensitive to different Intersection over Union (IOU) thresholds, and their variations directly impact precision and recall, thereby affecting the AP metric. An AP value closer to 1 signifies better recognition performance. Mean average precision (mAP) is another important metric, which aggregates the AP values across various object classes to provide an overall performance measure. The formulas for calculating P, R, F1, AP and mAP are as follows:
Additionally, we evaluated the model complexity using the number of parameters, model size and GFLOPs. The number of parameters reflects the number of weights that the model’s convolutional layers need to learn. Meanwhile, the model size denotes the amount of memory space the model occupies on the hardware platform. GFLOPs represent the computational complexity of the model. The formulas are as follows:
where
denotes a constant order,
K represents the convolution kernel size,
C is the number of channels,
M is the input image size, and
i is the layer index.
4. Discussion
We developed a citrus bud detection model named CFADet specifically for early bud estimation. Comprehensive experiments in
Section 3 (comparative experiments, ablation studies, and on-device deployment validation) demonstrate that our model outperforms state-of-the-art methods in both accuracy and computational efficiency, making it highly suitable for precise early yield estimation. In this section, we provide an analysis of the experimental results, underscore the contributions of the proposed enhancements, and highlight potential future research directions to further improve its performance.
Ablation experiments validated the superior accuracy and efficacy of the introduced enhancements in citrus bud detection for early yield estimation. EFF adaptively fuses multi-scale features from different layers to provide complementary information. CBEM enhances the model’s discriminative ability for small targets through multi-branch convolution and pooling operations, making the targets more distinguishable in complex backgrounds. Furthermore, FAM offers a novel strategy for background suppression, leading to improved detection accuracy, especially in challenging scenarios involving occlusions and feature-similar backgrounds. Finally, SPDConv is introduced to reconstruct the backbone, thus enhancing the model’s ability to retain fine-grained information. These enhancements collectively render the model highly versatile for early yield estimation in agricultural applications. Future research could explore extending these improvements to other object detection tasks within environments of comparable complexity, such as vision-based small target drone detection in complex outdoor scenes and attention-based object detection for intricate traffic scenes [
44,
45].
Beyond validating individual module contributions, comparisons with other object detection models further highlight CFADet’s strengths in detecting small citrus buds under dense foliage and varying lighting conditions. This advantage is critical for early agricultural yield estimation, as accurate small-bud detection directly affects yield prediction precision. This accuracy advantage is reflected in its 87.8% mAP, outperforming mainstream models like YOLO11s (82.1%) and Cascade R-CNN (80.7%) in complex orchard scenarios. Despite these significant improvements, there is still room for further enhancement, particularly in handling extreme occlusion and device-specific optimization. Notably, CFADet may still fail to detect buds blocked at extreme angles or occluded by branches and other buds, which could affect the accuracy of early yield estimation. Addressing this issue could involve integrating more robust occlusion-handling mechanisms, such as attention-guided feature recovery, which has been proven effective for adaptive feature interaction and enhancement in complex vision tasks [
46]. In terms of deployment, we verified the feasibility of CFADet in real-world applications by deploying the converted NCNN model on mobile devices via Android Studio. Nevertheless, there remains room for further optimization across diverse devices, a step that could further boost inference speed.
In addition to the model-specific limitations discussed, the scope of this study is further constrained by the types of citrus data currently available. The limitation to these two data types restricts the application scope of this study in citrus to yield estimation and flower thinning alone. Expanding coverage to all citrus types would extend the research scope to citrus disease management and growth status monitoring [
47]. However, attaining full coverage across all citrus types still presents a challenge [
48]. Considering the limited scale and single-orchard source of the experimental dataset, two prominent issues deserve in-depth discussion: on the one hand, the constrained sample diversity and single-scenario source introduce potential overfitting risks, as the model may overly adapt to the specific training scenes instead of learning universal bud features; on the other hand, the lack of cross-regional and cross-environmental data directly limits the model’s generalization ability, leading to possible performance degradation when deployed to unseen orchards with different growth conditions, bud morphologies or climatic backgrounds. Moreover, the insufficient environmental and sample diversity cannot be fully resolved by preliminary training strategies alone, representing a core limitation rooted in the dataset. To overcome these interconnected limitations caused by dataset constraints, we will collaborate with agricultural institutions to collect multi-stage citrus data across diverse regions, ensuring the dataset covers various growth stages, environmental conditions and citrus varieties. We will also further optimize the training regime to suppress overfitting risks and enhance the model’s adaptability to unseen scenes. Beyond the current two categories, we plan to incorporate the morphological characteristics of citrus flowers and fruits throughout their growth cycles to develop a comprehensive detection algorithm. This approach is anticipated to enhance the model’s robustness and broaden its applicability to a wide range of agricultural settings, thereby facilitating more comprehensive citrus cultivation management.
5. Conclusions
In this study, a contextual and frequency-aware citrus bud detection framework, CFADet, is proposed to achieve accurate and efficient citrus bud detection in complex orchard environments and support reliable early yield estimation. The proposed framework integrates four key enhancements to address the challenges of small object size, dense distribution, and severe background interference in real-world orchard scenarios. Specifically, an enhanced feature fusion network (EFF) with dual-path adaptive weighting is designed to strengthen multi-scale feature aggregation and improve spatial information transmission for small targets. A contextual boundary enhancement module (CBEM) is introduced to capture surrounding contextual cues and refine boundary representations, thereby improving the discriminative ability of citrus buds in cluttered environments. Furthermore, a frequency-aware module (FAM) is developed to suppress complex background noise by adaptively regulating frequency-domain components, which enhances feature robustness under varying illumination and occlusion conditions. Spatial-to-depth convolution (SPDConv) is employed to reconstruct the backbone to preserve fine-grained spatial details while reducing model parameters and improving computational efficiency. Experimental results on the self-constructed citrus bud dataset demonstrate that CFADet achieves 81.1% precision, 80.9% recall, 81.0% F1-score, and 87.8% mAP, showing competitive performance compared with existing detection methods. CFADet also achieves real-time performance on mobile devices with 29 FPS, validating its applicability in resource-constrained orchard environments. As a preliminary study, future work will focus on expanding the dataset across more regions and growth stages and further improving model robustness under complex orchard environments, providing a stronger foundation for large-scale intelligent orchard monitoring and early yield estimation.