Skip to Content
InsectsInsects
  • Article
  • Open Access

25 September 2026

22 Pages

An Explainable Hybrid Machine Learning Approach Based on MPSO-XGBoost for Pest Density Prediction in Almond Orchards

,
,
,
,
and
1
Department of Continuing Education Center, Firat University, Elazig 23119, Türkiye
2
Department of Computer Technologies, Bingöl University, Bingöl 12000, Türkiye
3
Department of Software Engineering, Faculty of Engineering and Naturel Sciences, Malatya Turgut Özal University, Malatya 44200, Türkiye
4
Department of Software Engineering, Faculty of Engineering, Firat University, Elazig 23119, Türkiye
This article belongs to the Section Insect Pest and Vector Management

Simple Summary

Predicting pest density in almond orchards is essential for preventing crop damage and reducing unnecessary pesticide use. This study introduces an explainable hybrid machine learning approach that integrates MPSO-based hyperparameter optimization, XGBoost classification, and rule extraction to classify pest density levels (low, medium, or high) by evaluating weather data, seasonal timing, and tree growth characteristics. The proposed approach reached approximately 80% accuracy in comparisons with 16 machine learning algorithms tested under the same conditions. Crucially, the model translates complex predictions into clear “IF–THEN” rules based on thresholds such as tree age and temperature, allowing agricultural experts to examine the conditions underlying model decisions.

Abstract

Accurate prediction of pest density and population levels in almond production is of great importance for crop yield, the prevention of quality losses, and the effectiveness of integrated pest management (IPM) strategies. Complex feature interactions in agricultural field data can affect both the predictive performance and interpretability of machine learning models. In this study, an explainable hybrid machine learning framework combining Mutated Particle Swarm Optimization (MPSO), XGBoost classification, and rule extraction was developed to classify pest insect density levels. The proposed approach was evaluated using two complementary validation settings. First, it was compared with 16 classification algorithms, including ensemble, deep learning, and rule-based methods, under stratified 5-fold cross-validation, achieving a mean accuracy of 79.63% and a mean Weighted F1-Score of 79.79%. In addition, nested stratified 5-fold cross-validation was employed to separate MPSO-based hyperparameter optimization from outer-fold performance evaluation. Under this more rigorous evaluation, the proposed approach achieved a mean accuracy of 79.27 ± 1.66% and a mean Weighted F1-Score of 79.29 ± 1.38%. Furthermore, interpretable IF–THEN decision rules were derived from the trained XGBoost decision structures to provide a transparent representation of the feature–threshold combinations associated with the classification decisions.

1. Introduction

The almond is a type of hard-shelled fruit with high economic value that is grown in many regions of the world, particularly in the Mediterranean climate zone [1,2]. To sustain high yields and product quality in almond production, growing conditions must be properly managed, and diseases and pests must be effectively controlled. In this context, insect pests can cause damage to leaves, shoots, and fruits during different stages of the tree’s development, leading to a reduction in yield and quality [1,3]. The extent of this damage varies depending on the density of insect pests in orchards [1]. Numerous insect species have been recorded as pests in almond orchards worldwide, and the level of damage they cause to leaves, shoots, and fruit varies depending on the species, phenological stage, and population density [1,4]. It has been reported that pest species capable of developing significant populations in almond orchards are also present in Türkiye’s almond production areas, and the timing of control measures should be based on population monitoring [5]. Pest density varies depending on environmental conditions, seasonal changes, and the structural characteristics of the trees [6,7]. Failure to identify these changes in pest density promptly can lead to delays in control measures and, consequently, increased crop losses. Therefore, it is necessary not only to determine the current density of pests but also to predict high-risk periods by evaluating the variables that affect density. Consequently, the accurate and timely prediction of pest populations is of great importance for implementing control measures at the appropriate time, reducing crop losses, and using production resources more efficiently.
The development and distribution of insect pests in orchards are closely related not only to their biological characteristics but also to the environmental conditions in which they live [6,7]. In particular, temperature and relative humidity can influence the insects’ development period, survival, and reproductive activities, leading to seasonal fluctuations in population density [6,7]. In addition, wind speed, elevation, and habitat characteristics can affect the movement and distribution of pests within orchards [8,9,10]. In fact, it has been shown that the increase in temperature directly changes the development rate, generation number and first appearance time of harmful insects in the season, which increases pest pressure in hard-shelled fruit species such as almonds [11]. In a machine learning-based study conducted under different microclimate conditions, it was determined that temperature is the dominant determinant, while humidity and wind have a secondary effect [8]. Changes in pest density are not limited solely to environmental conditions. Structural characteristics such as tree age, height, and canopy width can alter pests’ feeding and sheltering conditions, leading to variations in density levels among trees. This situation leads to a divergence in density values in samples taken from different parts of the tree within the same orchard; it has been reported that sampling carried out in different layers of the crown structure yields significantly different results [12]. Studies conducted in Mediterranean garden systems have shown that changes in tree density, age, and size determine the abundance of canopy arthropods [13]. In addition, it has been observed that the species diversity detected in almond orchards can vary greatly between orchards, which is associated with structural and environmental differences at the orchard level [2]. Since the effects of these variables do not occur at the same level during different stages of the growing season, temporal characteristics must also be considered when assessing pest density. Indeed, previous studies have shown that the combined evaluation of meteorological variables and plant phenology contributes to the prediction of pest populations [14,15]. Therefore, the reliable prediction of pest density depends on evaluating environmental, temporal, and tree-structure-related variables not in isolation but in conjunction, while preserving the relationships among them.
Field observations, trap monitoring, and expert assessments provide important information for monitoring insect pests. However, since these methods require regular monitoring across large production areas, they can pose significant challenges in terms of time and labor [7]. In order to reduce this limitation, deep learning-based monitoring systems that automatically detect pests from trap images have been developed in recent years [16]. Although these systems, integrated with light and sticky traps, have been shown to provide high accuracy in garden conditions [17,18], these approaches focus on counting the pest from the image; they do not include the evaluation of environmental and structural variables that determine the density. Furthermore, evaluating the relationships between environmental conditions, tree characteristics, and temporal changes using traditional methods is not always easy [16]. This situation can make it difficult to detect changes in pest density promptly and to accurately identify high-risk periods. Therefore, supplementing field observations with prediction methods capable of analyzing multidimensional data is crucial for determining pest density more reliably and promptly [7,14].
In line with this need, machine learning methods offer an effective approach for identifying the complex relationships among the numerous variables that influence pest density [8,19]. These methods can predict the presence or density levels of pests by leveraging patterns learned from past observations. Indeed, previous studies that evaluated temperature, relative humidity, plant phenology, and other environmental variables together have demonstrated that the emergence times and population levels of insect pests can be predicted using machine learning methods [7,14,15]. These studies have reported that using lagged meteorological variables improves forecasting accuracy; for example, 20–35 day lagged temperature and humidity data in orchards show a high correlation with population density [20,21]. XGBoost, one of the methods that can be used for this purpose, stands out in the analysis of multidimensional agricultural data because it can model nonlinear relationships and feature interactions among variables [22]. However, obtaining reliable results from a powerful classification model such as XGBoost depends on properly addressing the class distribution in the data and the hyperparameters that determine the model’s learning process.
Random cross-validation can yield overly optimistic performance estimates in datasets where observations are not independent of each other, particularly when repeated observations from the same unit are distributed across training and validation folds [23,24]. Additionally, XGBoost prediction performance depends on selecting appropriate values for the hyperparameters that determine the model structure and learning process. Particle Swarm Optimization, used to find effective solutions in large and complex hyperparameter spaces, guides the search process by leveraging the individual and collective experiences of candidate solutions [25,26,27]. The mutation mechanism added to this method aims to reduce the likelihood of premature convergence by preserving particle diversity [28].
PSO variants enhanced with mutations and similar diversity conservation mechanisms have been shown in various application areas to yield better results in determining XGBoost hyperparameters compared to grid and random search [26,27,29]. However, some comparisons conducted on agricultural data have also reported that Bayesian optimization approaches provide higher performance than PSO and genetic algorithms, indicating that the choice of optimization method depends on the problem [30]. While high classification performance is important in pest density prediction, it is also necessary to understand which variables and thresholds the model uses to make decisions. Particularly in agricultural decision-support applications, the ability of domain experts to interpret predictions is crucial for the reliability and practical applicability of model results [31,32]. Therefore, translating complex model decisions into understandable IF–THEN rules can help ensure both prediction performance and interpretability [33]. In methods for deriving rules from tree communities, it is emphasized that the obtained rule set should be evaluated not only in terms of its accuracy but also in terms of criteria such as its coverage/support, fidelity to the original model, and the number and length of rules [34,35,36]. The studies summarized above demonstrate that pest populations can be estimated using environmental data. However, the reviewed literature indicates that most previous studies focus on individual methodological components, while limited attention has been given to integrating population-based hyperparameter optimization, robust model evaluation, and the extraction of interpretable rules from model decisions within a unified framework. To our knowledge, an interpretable pest density classification that considers environmental, temporal, and tree structure variables together in almond orchards has not been addressed previously. The contribution of this study is to combine these components within a unified validation framework designed to minimize information leakage during hyperparameter optimization and performance evaluation, while transforming the resulting model decisions into rules that can be examined by field experts. In this context, the contribution is at the level of application and integration, not methodology.
Accordingly, this study developed an explainable hybrid machine learning approach to predict pest insect density in almond orchards based on environmental, temporal, and tree structure-related variables. The proposed approach integrates Mutated Particle Swarm Optimization to determine XGBoost hyperparameters and rule extraction to make model decisions understandable within a unified framework. The key contributions of this study to the literature and practice are summarized below:
  • Hyperparameter optimization with MPSO: The key hyperparameters of the XGBoost model were determined using Mutated Particle Swarm Optimization (MPSO), which aims to preserve particle diversity during the search process and reduce the likelihood of premature convergence. This enabled an effective search across the extensive hyperparameter space.
  • Integrated evaluation of multidimensional data: In estimating pest density, environmental and meteorological variables were evaluated alongside temporal characteristics and structural features such as tree age, height, and canopy width. This ensured a comprehensive modeling of the relationships among the various variables affecting pest density.
  • Combining prediction performance and interpretability: The proposed approach was compared with 16 different classification algorithms under the same evaluation conditions and achieved an average accuracy of 79.63% and an average Weighted F1-Score of 79.79%. Furthermore, by deriving IF–THEN rules from the XGBoost model’s decision structure—which include thresholds related to environmental, temporal, and tree structure factors—the model’s predictions were made examinable and interpretable by domain experts.

3. Materials and Methods

The flowchart presented in Figure 1 illustrates the proposed explainable hybrid machine learning approach, which integrates nested stratified 5-fold cross-validation, Mutated Particle Swarm Optimization (MPSO)-based XGBoost hyperparameter optimization, performance evaluation on unseen outer validation folds, and explainable rule extraction. Hereafter, the term “proposed approach” refers to the overall framework, whereas the term “model” specifically refers to the trained XGBoost classifier. In the first stage, the dataset is loaded and preprocessed, including missing-value handling, numerical conversion of temporal variables, and encoding of the categorical target variable. The dataset is then divided using an outer stratified 5-fold cross-validation scheme, in which four folds are used for training and the remaining fold is retained as unseen validation data in each iteration. Within each outer training set, an inner stratified 5-fold cross-validation procedure is employed for MPSO-based hyperparameter optimization. A population of 100 particles explores the XGBoost hyperparameter space, including the number of estimators, maximum tree depth, learning rate, subsampling ratio, feature subsampling ratio, and gamma. For each candidate parameter set, the mean weighted F1-Score across the five inner validation folds is used as the fitness value. During the optimization process, the personal best (pBest) and global best (gBest) positions are updated, followed by adaptive inertia adjustment, velocity and position updates, and mutation. The optimization continues until the maximum number of iterations is reached or the early-stopping criterion is satisfied. The best hyperparameters obtained for each outer fold are then used to train an XGBoost classifier on the corresponding outer training set, and its performance is evaluated on the unseen outer validation fold. Finally, the results from all five outer folds are aggregated to calculate the mean and standard deviation of the performance metrics, and explainable IF–THEN decision rules are derived from the trained XGBoost decision structure to provide transparent and interpretable predictions (Figure 1).
Figure 1. General workflow of the proposed explainable hybrid machine learning approach, including nested stratified 5-fold cross-validation (CV), Mutated Particle Swarm Optimization (MPSO)-based Extreme Gradient Boosting (XGBoost) hyperparameter optimization, outer-fold performance evaluation, and explainable rule extraction.

3.1. Dataset Description

  • The dataset used in this study contains a total of 1645 observations and 16 variables, comprising 15 predictor features and one target class variable. The predictor variables represent environmental and meteorological conditions, tree morphological characteristics, and temporal information. The target variable, Class, has a three-class structure consisting of Low (L), Medium (M), and High (H) pest density levels. The predictor variables are grouped into three main categories:
  • Environmental and Meteorological Attributes: Altitude, Temperature, Humidity, Wind Speed, and Habitat.
  • Tree Morphological Features: TreeAge, TreeHeight, and TreeCanopyWidth.
  • Temporal Attributes: Year (2020–2022), Month, Day, Day_of_Year, Week, Day_of_Week, and Quarter.
Detailed definitions of all attributes in the dataset, together with their data types and descriptive statistical summaries (minimum, maximum, mean, and median values), are presented in Table 2.
Table 2. Definitions and statistical summary of the features in the dataset.
The field data were collected between 2020 and 2022 from almond orchards located in Keban District and Nimri Village of Elazığ Province and Eğil District of Diyarbakır Province, Türkiye. Field observations were conducted through direct in situ visits to the selected orchards. Depending on pest activity and field conditions, the orchards were monitored through daily, weekly, and monthly inspections during the study period. Pest density was determined based on the abundance of Cimbex quadrimaculata larvae observed on different almond trees. Overall, the dataset comprises observations collected across three study locations during the 2020–2022 sampling period.
For the classification task, the target variable was defined according to the observed number of C. quadrimaculata larvae. Observations with 0–2 larvae were assigned to the Low (L) class, those with 3–5 larvae to the Medium (M) class, and those with 6 or more larvae to the High (H) class. These categories represent relative pest density levels within the present dataset and should not be interpreted as universally applicable economic injury thresholds. The resulting class distribution is presented in Table 3.
Table 3. Target Class Distribution.
As shown in Table 3, Class L accounts for 39.51% of the observations, while Classes H and M represent 31.31% and 29.18%, respectively. The three classes are therefore represented at moderately different frequencies in the dataset.

3.2. Data Preprocessing and Nested Stratified 5-Fold Cross-Validation

To improve model generalizability and provide an unbiased estimate of predictive performance, the raw dataset was first cleaned, temporal features were converted to numerical formats, and the categorical target variable was numerically encoded. Model development and evaluation were then conducted using nested stratified 5-fold cross-validation. The outer 5-fold cross-validation was used for performance evaluation, while the inner 5-fold cross-validation was performed exclusively within each outer training fold for MPSO-based hyperparameter optimization. Stratification was applied to maintain approximately the original class distribution across the folds.

3.3. Implementation Details and Reproducibility

The analytical pipeline was implemented as a nested stratified cross-validation framework to ensure a clear separation between hyperparameter optimization and performance evaluation. First, the dataset was loaded and preprocessed, and the categorical target variable was numerically encoded using a LabelEncoder. The predictor matrix (X) and target vector (y) were then constructed. Model development and evaluation were performed using an outer Stratified 5-Fold Cross-Validation (CV) procedure with shuffling enabled and a random state of 42. In each outer iteration, four folds were used as the outer training set, while the remaining fold was retained as unseen validation data and was not involved in hyperparameter optimization.
Hyperparameter optimization was performed exclusively on the outer training set using Mutated Particle Swarm Optimization (MPSO). Within each outer training fold, an inner Stratified 5-Fold CV procedure was employed to evaluate candidate XGBoost hyperparameter configurations. The optimization search space comprised six XGBoost hyperparameters: n_estimators (50–300), max_depth (2–8), learning_rate (0.01–0.30), subsample (0.50–1.00), colsample_bytree (0.50–1.00), and gamma (0.00–5.00). The n_estimators and max_depth parameters were rounded to integer values when converting particle positions into XGBoost parameter configurations. For each candidate configuration, the mean Weighted F1-score obtained across the five inner validation folds was used as the fitness function. Thus, the outer validation observations were completely excluded from the MPSO and hyperparameter-selection process.
The MPSO procedure used a population of 100 particles and a maximum of 50 iterations. The inertia weight was adaptively decreased from 0.90 to 0.40 during optimization, while both the cognitive (c1) and social (c2) acceleration coefficients were set to 2.0. Mutation was applied with a probability of 0.20 and a mutation strength of 0.10 to increase population diversity and reduce premature convergence. The personal best (pBest) and global best (gBest) solutions were updated according to the fitness values obtained during each iteration. An early-stopping mechanism with a patience of five iterations and a minimum improvement threshold (min_delta) of 1 × 10−4 was used to terminate optimization when no meaningful improvement in the global best fitness was observed.
After MPSO was completed for an outer fold, the hyperparameter configuration associated with the best inner-CV fitness was selected. An XGBoost classifier was then trained using the complete outer training set with these optimized parameters and subsequently evaluated on the corresponding unseen outer validation fold. This procedure was independently repeated for all five outer folds. Model performance was assessed using accuracy, weighted precision, weighted recall, Weighted F1-Score, and log loss. Predictions from the outer validation folds were also consolidated to construct an out-of-fold confusion matrix and classification report. The mean and standard deviation of the performance metrics across the five outer folds were reported to characterize both predictive performance and its variability.
The complete sequence of data preprocessing, outer-fold separation, inner-fold hyperparameter optimization, XGBoost training, unseen outer-fold evaluation, aggregation of performance metrics, and subsequent explainable rule extraction is summarized in Figure 1. This nested design ensures that hyperparameter selection is performed independently of the observations used for final performance estimation, thereby reducing optimistic bias and providing a more rigorous assessment of model generalization.
Computational Environment: All analyses were implemented in Python 3.11.11 using the Spyder integrated development environment (IDE) on a 64-bit Windows 11 Pro operating system. The experiments were conducted on a computer equipped with an Intel Core i5-10400 CPU running at 2.90 GHz and 8 GB of RAM. The system included integrated Intel UHD Graphics 630; however, GPU acceleration and CUDA were not used, and all analyses were executed on the CPU. The principal Python packages used in the analysis were NumPy 1.26.4, pandas 3.0.1, scikit-learn 1.8.0, and XGBoost 3.1.2.

3.4. Extraction of Explainable Decision Rules

In line with the Explainable Artificial Intelligence (XAI) approach, an explainable rule-extraction procedure was applied to make the decision structure of the XGBoost classifier more transparent and interpretable. The decision rules were derived directly from the tree structures generated by the trained XGBoost model. Each decision tree consists of a sequence of internal decision nodes defined by feature-specific split thresholds and terminal leaf nodes representing the resulting model decisions. To obtain interpretable rules, the decision paths within the trees were examined from the root node to the corresponding terminal leaf nodes. The successive feature–threshold conditions encountered along each decision path were combined to construct the antecedent (IF) part of a rule, while the decision associated with the corresponding terminal path was represented as the consequent (THEN) part.
When a feature appeared more than once along the same decision path, its successive threshold conditions were combined to obtain a bounded feature interval where applicable. In this way, multiple tree-splitting conditions along a path could be represented in a compact interval-based form. The extracted rules therefore describe combinations of feature intervals associated with the model’s classification decisions and follow the general form:
I F L 1 ≤ X 1 ≤ U 1 A N D L 2 ≤ X 2 ≤ U 2 … L k ≤ X k ≤ U k T H E N   C l a s s = C
where X k represents a predictor involved in the decision path; L k and U k denote its lower and upper threshold bounds, respectively; and C represents the corresponding predicted class. This procedure converts the tree-based decision structure learned by XGBoost into explicit symbolic IF–THEN rules, allowing the combinations of feature conditions underlying the model’s predictions to be examined in a more transparent and human-interpretable form.
Rule accuracy was calculated by applying each extracted rule to the dataset and identifying all observations that satisfied the rule conditions. For a given rule, accuracy was defined as the proportion of observations satisfying the rule conditions whose actual class label matched the class predicted by the rule. Specifically, if N r denotes the total number of observations satisfying a rule and N c denotes the number of those observations whose actual class corresponds to the rule-predicted class, rule accuracy was calculated as follows:
R u l e   A c c u r a c y = N c N r
where N r is the total number of observations satisfying the rule conditions, and N c is the number of those observations whose actual class label matches the class predicted by the rule.

4. Results

The proposed MPSO-optimized XGBoost classification approach (the proposed method) was compared with 16 classification algorithms under a Stratified 5-Fold Cross-Validation protocol. The mean accuracy and mean Weighted F1-Score values, together with their standard deviations across the five folds, are presented in Figure 2.
Figure 2. Comparison of the mean accuracy (%) and mean Weighted F1-Score (%) of the proposed MPSO-XGBoost approach and other classification models under stratified 5-fold cross-validation. Error bars represent ±1 standard deviation (SD) across the five cross-validation folds.
As shown in Figure 2, the proposed method achieved a mean accuracy of 79.63% and a mean Weighted F1-Score of 79.79%. Standard XGBoost achieved corresponding values of 79.21% and 79.23%, respectively. Thus, the proposed method showed numerically higher mean values than standard XGBoost, with differences of 0.42 percentage points in accuracy and 0.56 percentage points in Weighted F1-Score. However, these differences were relatively small, and variability was observed across the cross-validation folds. Therefore, in the absence of a formal statistical significance test, the observed differences should be interpreted as modest numerical improvements rather than evidence of statistically significant superiority.
Other ensemble and machine learning methods also showed comparable classification performance, including CatBoost (78.91%), ExtraTrees (77.39%), Random Forest (77.20%), MLP (77.33%), and TabNet (76.72%). In comparison, the evaluated rule-based algorithms, including RIPPER, C4.5 Rules, IREP, and RuleFit, produced accuracy values ranging from 70.46% to 74.41%. Beyond predictive performance, the proposed approach incorporates an explainability component in which the decision structures learned by the XGBoost classifier are represented as interpretable IF–THEN rules. As described in the rule-extraction methodology, feature–threshold conditions along the decision paths of the XGBoost trees were converted into interval-based symbolic rules. This representation provides an interpretable view of the feature conditions underlying the model’s classification decisions while maintaining classification performance comparable to the best-performing methods evaluated in this study.
The performance of the proposed MPSO-optimized XGBoost approach was further evaluated using the consolidated confusion matrix (Figure 3a), class-specific performance metrics (Figure 3b), and performance variation across the five outer cross-validation folds (Figure 3c). The consolidated confusion matrix, generated by aggregating the out-of-fold predictions from the five outer validation folds, provides an overall representation of the classification results. The proposed approach achieved a mean accuracy of 79.27 ± 1.66% and a mean Weighted F1-Score of 79.29 ± 1.38% across the five outer folds.
Figure 3. Detailed performance analysis of the proposed method: (a) Consolidated 5-Fold Cross-Validation Confusion Matrix, (b) Class-based performance metrics (Precision, Recall, F1-Score), and (c) the variation in cross-foldperformance metrics (accuracy, Weighted F1-Score, log loss).
The class-specific results presented in Figure 3b show that classification performance varied among the three classes. For the H (High) class, precision, recall, and F1-score were 84.78%, 85.44%, and 85.11%, respectively. The L (Low) class achieved a precision of 85.44%, recall of 83.08%, and F1-score of 84.24%. Comparatively lower performance was observed for the M (Medium) class, with a precision of 65.59%, recall of 67.50%, and F1-score of 66.53%, indicating that this intermediate category was more difficult to distinguish than the H and L classes.
Figure 3c further illustrates the variation in accuracy, Weighted F1-Score, and log loss across the five outer cross-validation folds. Accuracy and Weighted F1-Score followed relatively similar patterns across the folds, with mean values of 79.27 ± 1.66% and 79.29 ± 1.38%, respectively, while the mean log loss was 0.5360 ± 0.0913. These fold-level results indicate relatively limited variation in accuracy and Weighted F1-Score across the outer validation subsets, although greater variability was observed in log loss. Overall, the results demonstrate relatively consistent predictive performance across the five outer folds while also revealing clear class-specific differences, particularly the comparatively lower classification performance observed for the M class.
In line with the Explainable Artificial Intelligence (XAI) approach, the symbolic decision rules derived from the optimized XGBoost model, together with their accuracy values, are presented in Table 4. These rules provide an interpretable representation of the feature–threshold combinations underlying the model’s classification decisions. Examination of the extracted rules showed that combinations of temporal (Day_of_Year, Month, Day), meteorological (Temperature, Humidity, Wind Speed, Altitude), and morphological (TreeAge, TreeHeight, TreeCanopyWidth) features were associated with the classification of pest density levels through specific threshold ranges.
Table 4. Explainable Decision Rules and Accuracy Rates Derived by the Proposed Method.
The decision rules representing Class L (Low Density) have high accuracy rates ranging from 0.96 to 0.99 (Table 4). These results indicate that the extracted rules identify relatively distinct feature–threshold combinations associated with Class L. In particular, during the later part of the season (185 ≤ Day_of_Year ≤ 206, i.e., mid- to late July) and at higher temperature values (28.5 °C ≤ Temperature ≤ 36.7 °C), observations satisfying specific tree height (2.0–4.5 m) or canopy width (1.4–2.8 m) conditions were classified as Class L with rule accuracies ranging from 0.96 to 0.99. These patterns indicate that combinations of seasonal timing, temperature, and morphological characteristics were associated with the model’s classification of observations as Class L. These relationships should be interpreted as model-derived predictive associations rather than biologically established thresholds or causal relationships.
The rules derived for Class H (High Density) exhibit accuracy rates ranging from 0.84 to 0.90 (Table 4). Tree age (8 ≤ TreeAge ≤ 20), seasonal timing (139 ≤ Day_of_Year ≤ 178, i.e., late May–late June), and temperature (22.2 °C ≤ Temperature ≤ 36.7 °C) frequently appeared in the extracted rules associated with Class H. The co-occurrence of these feature ranges within the rules indicates that combinations of tree age, seasonal timing, and temperature contributed to the model-derived predictive patterns associated with Class H. As with Class L, these thresholds represent the decision structure learned by the XGBoost model and should not be interpreted as causal, mechanistic, or universally applicable agronomic thresholds.
The accuracy rates of the rules defining Class M (Medium Density) were lower, ranging from 0.60 to 0.67 (Table 4). This finding is consistent with the comparatively lower class-specific performance observed in the confusion matrix (Figure 3a) and performance metrics (Figure 3b). The extracted rules for Class M involve combinations of morphological, temporal, and environmental variables whose ranges may overlap with patterns associated with Classes L and H. Therefore, these rules should be interpreted as model-derived predictive associations for the intermediate density category rather than as evidence of a biological or temporal transition stage or an operational early-warning threshold. Nested Cross-Validation Analysis
To further assess the generalization performance of the proposed MPSO-XGBoost approach while separating hyperparameter optimization from model evaluation, an additional nested stratified 5-fold cross-validation analysis was performed. In each outer fold, MPSO-based hyperparameter optimization was conducted exclusively within the corresponding outer training set using an inner stratified 5-fold cross-validation procedure, with Weighted F1-score used as the optimization fitness criterion. The optimized model was subsequently trained on the complete outer training set and evaluated on the unseen outer validation fold.
As shown in Table 5, across the five outer folds, the proposed approach achieved a mean accuracy of 0.7927 ± 0.0166 and a mean Weighted F1-Score of 0.7929 ± 0.0138. The mean precision and recall were 0.7967 ± 0.0147 and 0.7927 ± 0.0166, respectively, while the mean log loss was 0.5360 ± 0.0913. The relatively limited variation in accuracy and Weighted F1-Score across the outer folds indicates that the predictive performance was reasonably consistent across the five unseen validation subsets.
Table 5. Performance of MPSO-XGBoost across the outer folds of nested stratified 5-fold cross-validation.
The MPSO search space was defined as follows: n_estimators = 50–300, max_depth = 2–8, learning_rate = 0.01–0.30, subsample = 0.50–1.00, colsample_bytree = 0.50–1.00, and gamma = 0.00–5.00. The MPSO algorithm was configured with a population of 100 particles and a maximum of 50 iterations. The inertia weight was adaptively decreased from 0.90 to 0.40 during optimization, while the cognitive and social acceleration coefficients were set to c1 = 2.0 and c2 = 2.0, respectively. To improve exploration of the search space, the mutation rate (probability) was set to 0.20, with a mutation strength of 0.10. Early stopping was applied when no improvement greater than 1 × 10−4 was observed for five consecutive iterations. The optimal XGBoost hyperparameters obtained independently for each outer fold are presented in Table 6.
Table 6. Optimal XGBoost hyperparameters identified by MPSO within each outer training fold.
The number of estimators selected across the five outer folds ranged from 139 to 300, while the maximum tree depth ranged from 3 to 8. The optimized learning rate ranged from 0.0123 to 0.2614, the subsample ratio from 0.5000 to 1.0000, the colsample_bytree ratio from 0.5000 to 0.9013, and gamma from 0.0000 to 0.4415. The corresponding inner-CV Weighted F1-scores ranged from 0.7795 to 0.7911. The variation in the selected hyperparameters across the outer folds reflects the independent optimization performed within each outer training set. The corresponding outer validation fold was not used during hyperparameter selection.

5. Discussion

In this study, an explainable hybrid machine learning framework combining Mutated Particle Swarm Optimization (MPSO)-based hyperparameter optimization, XGBoost classification, and explainable rule extraction was evaluated for the classification of pest density levels based on environmental, morphological, and temporal variables. In the stratified 5-fold cross-validation experiment used for comparison with the 16 benchmark classification algorithms, the proposed approach achieved a mean accuracy of 79.63% and a mean Weighted F1-Score of 79.79%. These results were compared with those obtained from 16 classification algorithms commonly used in machine learning studies.
Standard XGBoost achieved a mean accuracy of 79.21% and a mean Weighted F1-Score of 79.23%. Thus, the proposed MPSO-XGBoost approach produced numerically higher mean values, corresponding to differences of 0.42 percentage points in accuracy and 0.56 percentage points in Weighted F1-Score. However, these differences were relatively small, and variability was observed across the cross-validation folds. Because no formal statistical significance test was performed for these comparisons, the observed differences should be interpreted as modest numerical differences rather than evidence of statistically significant superiority. Other ensemble methods, including CatBoost (78.91% accuracy) and Random Forest (77.20% accuracy), also demonstrated comparable classification performance.
In addition to the benchmark comparison, nested stratified 5-fold cross-validation was used to provide a more rigorous assessment in which hyperparameter optimization was performed within the inner folds and predictive performance was evaluated on unseen outer validation folds. Across the five outer folds, MPSO-XGBoost achieved a mean accuracy of 79.27 ± 1.66% and a mean Weighted F1-Score of 79.29 ± 1.38%. The similarity between these nested cross-validation results and those obtained in the standard stratified 5-fold comparison suggests that the observed predictive performance was reasonably stable under the more stringent nested evaluation procedure.

5.1. MPSO-Based Hyperparameter Optimization of XGBoost

In the proposed framework, MPSO was used to optimize the XGBoost hyperparameters, including the number of estimators, maximum tree depth, learning rate, subsampling ratio, feature subsampling ratio, and gamma. The mutation mechanism incorporated into the particle swarm optimization process was used to increase exploration of the predefined hyperparameter search space. The mean Weighted F1-Score across the inner validation folds was used as the fitness criterion during MPSO. The resulting optimized parameter configurations were subsequently used to train the XGBoost classifier.
The relatively small numerical differences between MPSO-XGBoost and standard XGBoost indicate that the contribution of MPSO should be interpreted cautiously. Under the applied evaluation protocol, MPSO-XGBoost achieved slightly higher mean accuracy and Weighted F1-Score values than standard XGBoost. However, because no formal statistical significance test was performed, these differences cannot be interpreted as evidence that MPSO consistently improves predictive performance.

5.2. Class-Based Performance and Feature Complexity Analysis

The class-specific metrics presented in Figure 3b indicate differences in classification performance among the three pest density categories. Based on the consolidated out-of-fold predictions from the nested cross-validation analysis, the H (High) class achieved an F1-score of 85.11%, while the L (Low) class achieved an F1-score of 84.24%. Comparatively lower performance was observed for the M (Medium) class, with a precision of 65.59%, a recall of 67.50%, and an F1-score of 66.53%. These results indicate that the intermediate M category was more difficult for the model to distinguish from the other two classes.
The confusion matrix (Figure 3a) further shows that a considerable proportion of the classification errors involved the M class, including 86 M observations classified as L and 70 M observations classified as H. Although the extracted rules for the M category include feature ranges that may overlap with those observed for the L and H categories, these patterns should be interpreted as model-derived predictive associations rather than evidence of a biological or temporal transition between pest density classes. Therefore, Class M is considered an intermediate classification category rather than a transitional stage. Future studies could investigate whether ordinal classification or fuzzy logic approaches provide a more suitable representation of the boundaries among these pest density categories.
The comparatively lower rule accuracy observed for the M class should also be considered when interpreting the explainable rule set. The extracted Class M rules achieved rule accuracy values ranging from approximately 0.60 to 0.67, indicating lower predictive accuracy for the extracted M-class rules than for the rules associated with the H and L categories. This finding represents a limitation of the current rule-based interpretation and suggests that the M category may contain less distinct or more overlapping predictive patterns. Further studies using additional data and validation across different seasons and locations are required to determine whether these patterns are reproducible.

5.3. Comparison with Previous Studies

Comparison with previous studies should be interpreted cautiously because the available studies differ in crop species, target pests, datasets, predictor variables, and prediction tasks. Marković et al. [7] demonstrated that temperature and relative humidity can support machine learning-based prediction of pest occurrence, while Ibrahim et al. [14] showed that combining meteorological variables with plant growth stages provides useful information for predicting pest population dynamics. Similarly, Pimentel et al. [20] reported that lagged meteorological variables contributed to the prediction of fruit fly populations in apple orchards, emphasizing the potential relevance of temporal information in addition to current environmental conditions.
In the present study, temperature and intra-seasonal temporal variables appeared in the extracted decision rules, indicating predictive associations between these variables and pest density categories. These findings are broadly consistent with previous studies emphasizing the relevance of meteorological and temporal information in pest prediction, although they should not be interpreted as evidence of causal relationships.
From a modeling perspective, Singh et al. [42] showed that combining multiple machine learning components within a dynamic ensemble improved pest population prediction compared with individual models, illustrating the potential value of integrated learning strategies for complex pest–environment relationships. In the context of almond orchards, Broadhead et al. [40] reported that pest abundance measurements alone did not improve end-of-season damage prediction and that cultivar- and orchard-related characteristics were stronger predictors. In the present study, tree age and tree height also appeared in the extracted decision rules, indicating that structural tree characteristics contributed to the model’s classification decisions.
These comparisons suggest that jointly considering meteorological, temporal, and crop structural information may provide useful predictive information for pest density classification. Nevertheless, the relative importance and generalizability of these associations require further evaluation across independent datasets, different growing seasons, and different geographical regions. The present study contributes to this research direction by integrating heterogeneous predictors with MPSO-optimized XGBoost classification and explainable IF–THEN rule extraction within a unified framework.

5.4. Balance Between Explainability and Predictive Performance

Traditional rule-based methods such as RIPPER, C4.5 Rules, IREP, and RuleFit provide transparent IF–THEN decision structures by design and achieved accuracy values ranging from 70.46% to 74.41% in the present experiments. Other machine learning approaches, including TabNet (76.72%) and MLP (77.33%), achieved different levels of predictive performance but generally provide less directly interpretable decision structures than explicit rule-based models.
The proposed framework addresses this interpretability challenge by deriving explicit IF–THEN rules from the decision structures learned by the XGBoost classifier. The extracted rules provide feature–threshold combinations derived from the XGBoost decision structures that can be examined by researchers and domain experts. For example, some extracted rules associated combinations of tree age and intra-seasonal temporal conditions with the H (High) category, whereas other rules associated combinations of temporal and temperature conditions with the L (Low) category. These relationships should be interpreted as model-derived predictive associations rather than causal or mechanistic relationships.

6. Conclusions

This study presented an explainable hybrid machine learning framework integrating Mutated Particle Swarm Optimization (MPSO)-based hyperparameter optimization, XGBoost classification, and explainable IF–THEN rule extraction for the classification of pest density levels in almond orchards. The framework jointly considered environmental, meteorological, temporal, and tree morphological variables to characterize Low (L), Medium (M), and High (H) pest density categories.
Under the applied evaluation protocol, the proposed MPSO-XGBoost approach achieved a mean accuracy of 79.63% and a mean Weighted F1-Score of 79.79%. Standard XGBoost achieved corresponding values of 79.21% and 79.23%, indicating modest numerical differences of 0.42 and 0.56 percentage points, respectively. Because no formal statistical significance test was performed, these differences should not be interpreted as evidence of statistically significant superiority. Instead, the results indicate that MPSO-XGBoost provided classification performance comparable to the strongest machine learning approaches evaluated in this study while enabling an additional rule-based interpretation of the model’s decision structure.
The explainability component converted feature–threshold combinations learned by XGBoost into explicit IF–THEN rules, providing an interpretable representation of the predictive associations underlying the classifications. The H and L categories showed stronger class-specific classification performance, whereas the intermediate M category was more difficult to distinguish. In particular, the lower rule accuracy observed for Class M (0.60–0.67) highlights a limitation of the current rule-based interpretation and indicates that the predictive patterns associated with this category were less distinct than those associated with the H and L categories. The extracted relationships should therefore be interpreted as model-derived predictive associations rather than causal or mechanistic relationships.
Overall, the findings demonstrate the potential of combining hyperparameter-optimized tree-based classification with interpretable rule extraction for pest density assessment in almond orchards. Nevertheless, the generalizability of the framework should be evaluated using independent datasets collected across additional growing seasons, geographical regions, and orchard conditions. Future studies could also investigate alternative approaches, including ordinal classification and fuzzy logic, to better characterize the intermediate M category and further improve the robustness and interpretability of pest density classification.

Supplementary Materials

The following supporting information can be downloaded at https://www.mdpi.com/article/10.3390/insects17100996/s1, Dataset S1: Dataset used for pest density classification in almond orchards.

Author Contributions

Conceptualization, İ.Ö., C.B. and B.A.; methodology, İ.Ö., D.D. and B.A.; software, C.B. and B.A.; validation, İ.Ö., H.B. and H.Y.; formal analysis, İ.Ö. and H.B.; investigation, İ.Ö. and C.B.; resources, C.B.; data curation, İ.Ö. and H.B.; writing—original draft preparation, İ.Ö. and D.D.; writing—review and editing, C.B., İ.Ö. and B.A.; visualization, İ.Ö. and D.D.; supervision, C.B.; project administration, İ.Ö. and C.B. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the Fırat University Scientific Research Projects Coordination Unit (FUBAP), grant number MF.26.64.

Data Availability Statement

The data presented in this study are available in the Supplementary Materials (Dataset S1) accompanying this article.

Acknowledgments

The authors gratefully acknowledge the Scientific and Technological Research Council of Türkiye (TÜBİTAK) for supporting the field studies and data collection conducted under Project No. 118O124. During the preparation of this manuscript, the authors used ChatGPT 5.6 (OpenAI) for the purposes of assisting with manuscript drafting, language editing, and improving the clarity and organization of the text. The authors have reviewed and edited the output and take full responsibility for the content of this publication.

Conflicts of Interest

The authors declare no conflicts of interest. The funder had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results.

Abbreviations

The following abbreviations are used in this manuscript:
MPSOMutated Particle Swarm Optimization
XAIExplainable Artificial Intelligence
IPMIntegrated Pest Management

References

  1. Rijal, J.P.; Joyce, A.L.; Gyawaly, S. Biology, Ecology, and Management of Hemipteran Pests in Almond Orchards in the United States. J. Integr. Pest Manag. 2021, 12, 24. [Google Scholar] [CrossRef] [Scilit]
  2. Callahan, P.; Rijal, J.; Gyawaly, S.; Gonzalez, F.; Joyce, A. Hemiptera Diversity Using the mtDNA COI Barcode in Almond Orchards in the Northern San Joaquin Valley of California. J. Econ. Entomol. 2026, 119, 2708–2720. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  3. Higbee, B.S.; Burks, C.S. Individual and Additive Effects of Insecticide and Mating Disruption in Integrated Management of Navel Orangeworm in Almonds. Insects 2021, 12, 188. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  4. Daane, K.M.; Yokota, G.Y.; Wilson, H. Seasonal Dynamics of the Leaffooted Bug Leptoglossus zonatus and Its Implications for Control in Almonds and Pistachios. Insects 2019, 10, 255. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Mamay, M.; Demir, S.; Sönmez, C.; Mutlu, Ç.; Wainwright, M.; Alharbi, S.A. Population Dynamics of Cabbage Looper [Trichoplusia ni (Hübner, 1803) (Lepidoptera: Noctuidae)] in Almond Orchards. J. King Saud Univ.—Sci. 2023, 35, 102384. [Google Scholar] [CrossRef] [Scilit]
  6. Fisher, J.J.; Rijal, J.P.; Zalom, F.G. Temperature and Humidity Interact to Influence Brown Marmorated Stink Bug (Hemiptera: Pentatomidae), Survival. Environ. Entomol. 2021, 50, 390–398. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  7. Marković, D.; Vujičić, D.; Tanasković, S.; Đorđević, B.; Ranđić, S.; Stamenković, Z. Prediction of Pest Insect Appearance Using Sensors and Machine Learning. Sensors 2021, 21, 4846. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  8. Arora, A.K.; Anderson, N.; Gadhave, K.R. Machine Learning Reveals Microclimate-Specific Drivers of a Cosmopolitan Supervector’s Population Dynamics. Ecol. Inform. 2026, 94, 103690. [Google Scholar] [CrossRef] [Scilit]
  9. Poggetti, L.; Raranciuc, S.; Chiabà, C.; Vischi, M.; Zandigiacomo, P. Altitude Affects the Distribution and Abundance of Two Non-native Insect Pests of the Common Walnut. J. Appl. Entomol. 2019, 143, 527–534. [Google Scholar] [CrossRef] [Scilit]
  10. Bergh, J.C.; Morrison, W.R.; Stallrich, J.W.; Short, B.D.; Cullum, J.P.; Leskey, T.C. Border Habitat Effects on Captures of Halyomorpha halys (Hemiptera: Pentatomidae) in Pheromone Traps and Fruit Injury at Harvest in Apple and Peach Orchards in the Mid-Atlantic, USA. Insects 2021, 12, 419. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  11. Jha, P.K.; Zhang, N.; Rijal, J.P.; Parker, L.E.; Ostoja, S.; Pathak, T.B. Climate Change Impacts on Insect Pests for High Value Specialty Crops in California. Sci. Total Environ. 2024, 906, 167605. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  12. Achhami, B.B.; Rijal, J.P. Comparing Brown Mite, Bryobia rubrioculus Scheuten (Acari: Tetranychidae) Sampling Methods in Almond Orchards in California. Exp. Appl. Acarol. 2026, 96, 24. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  13. Vasconcelos, S.; Pina, S.; Herrera, J.M.; Silva, B.; Sousa, P.; Porto, M.; Melguizo-Ruiz, N.; Jiménez-Navarro, G.; Ferreira, S.; Moreira, F.; et al. Canopy Arthropod Declines along a Gradient of Olive Farming Intensification. Sci. Rep. 2022, 12, 17273. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  14. Ibrahim, E.A.; Salifu, D.; Mwalili, S.; Dubois, T.; Collins, R.; Tonnang, H.E.Z. An Expert System for Insect Pest Population Dynamics Prediction. Comput. Electron. Agric. 2022, 198, 107124. [Google Scholar] [CrossRef] [Scilit]
  15. Skawsang, S.; Nagai, M.; Tripathi, N.K.; Soni, P. Predicting Rice Pest Population Occurrence with Satellite-Derived Crop Phenology, Ground Meteorological Observation, and Machine Learning: A Case Study for the Central Plain of Thailand. Appl. Sci. 2019, 9, 4846. [Google Scholar] [CrossRef] [Scilit]
  16. Teixeira, A.C.; Ribeiro, J.; Morais, R.; Sousa, J.J.; Cunha, A. A Systematic Review on Automatic Insect Detection Using Deep Learning. Agriculture 2023, 13, 713. [Google Scholar] [CrossRef] [Scilit]
  17. Zhang, W.; Huang, H.; Sun, Y.; Wu, X. AgriPest-YOLO: A Rapid Light-Trap Agricultural Pest Detection Method Based on Deep Learning. Front. Plant Sci. 2022, 13, 1079384. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  18. Čirjak, D.; Aleksi, I.; Lemic, D.; Pajač Živković, I. EfficientDet-4 Deep Neural Network-Based Remote Monitoring of Codling Moth Population for Early Damage Detection in Apple Orchard. Agriculture 2023, 13, 961. [Google Scholar] [CrossRef] [Scilit]
  19. Chen, C.-J.; Li, Y.-S.; Tai, C.-Y.; Chen, Y.-C.; Huang, Y.-M. Pest Incidence Forecasting Based on Internet of Things and Long Short-Term Memory Network. Appl. Soft Comput. 2022, 124, 108895. [Google Scholar] [CrossRef] [Scilit]
  20. De Souza Pimentel, E.; Da Silva Paes, J.; Ramos, Y.J.; Soares, J.M.; Guedes, A.G.; Da Silva Sant’Ana, L.C.; Da Silva, R.S.; Picanço, M.C. Predicting the Seasonal Dynamics of Fruit Fly Anastrepha fraterculus Populations in Apple Orchards Using Artificial Neural Networks. Neotrop. Entomol. 2025, 54, 94. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  21. Oliveira, A.A.S.; Bastos, C.S.; Da Silva Paes, J.; Da Silva Sant’Ana, L.C.; De Araújo, T.A.; De Araújo, A.C.A.; Barman, A.; Picanço, M.C. Artificial Neural Network Model for Predicting Population Dynamics of Anthonomus grandis grandis (Coleoptera: Curculionidae) in Cotton Fields, as a Function of Climatic Elements. Theor. Appl. Climatol. 2026, 157, 303. [Google Scholar] [CrossRef] [Scilit]
  22. Chen, T.; Guestrin, C. XGBoost: A Scalable Tree Boosting System. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining; ACM: San Francisco, CA, USA, 2016; pp. 785–794. [Google Scholar]
  23. Just, A.C.; Arfer, K.B.; Rush, J.; Dorman, M.; Shtein, A.; Lyapustin, A.; Kloog, I. Advancing Methodologies for Applying Machine Learning and Evaluating Spatiotemporal Models of Fine Particulate Matter (PM2.5) Using Satellite Data over Large Regions. Atmos. Environ. 2020, 239, 117649. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  24. John, K.; Saurette, D.D.; Heung, B. The Problematic Case of Data Leakage: A Case for Leave-Profile-out Cross—Validation in 3-Dimensional Digital Soil Mapping. Geoderma 2025, 455, 117223. [Google Scholar] [CrossRef] [Scilit]
  25. Kennedy, J.; Eberhart, R. Particle Swarm Optimization. In Proceedings of the ICNN’95—International Conference on Neural Networks; IEEE: Perth, WA, Australia, 1995; Volume 4, pp. 1942–1948. [Google Scholar]
  26. Qin, C.; Zhang, Y.; Bao, F.; Zhang, C.; Liu, P.; Liu, P. XGBoost Optimized by Adaptive Particle Swarm Optimization for Credit Scoring. Math. Probl. Eng. 2021, 2021, 6655510. [Google Scholar] [CrossRef] [Scilit]
  27. Vikas Lakra, A.; Jena, S.; Mishra, K. Optimizing XGBoost Hyperparameters for Credit Scoring Classification Using Weighted Cognitive Avoidance Particle Swarm. IEEE Access 2025, 13, 127292–127306. [Google Scholar] [CrossRef] [Scilit]
  28. Lu, Z.; Hou, Z.; Du, J. Particle Swarm Optimization with Adaptive Mutation. Front. Electr. Electron. Eng. China 2006, 1, 99–104. [Google Scholar] [CrossRef] [Scilit]
  29. Demir, S.; Sahin, E.K. Predicting Occurrence of Liquefaction-Induced Lateral Spreading Using Gradient Boosting Algorithms Integrated with Particle Swarm Optimization: PSO-XGBoost, PSO-LightGBM, and PSO-CatBoost. Acta Geotech. 2023, 18, 3403–3419. [Google Scholar] [CrossRef] [Scilit]
  30. Dada, B.A.; Nwulu, N.I.; Olukanmi, S.O. Bayesian Optimization with Optuna for Enhanced Soil Nutrient Prediction: A Comparative Study with Genetic Algorithm and Particle Swarm Optimization. Smart Agric. Technol. 2025, 12, 101136. [Google Scholar] [CrossRef] [Scilit]
  31. Lundberg, S.M.; Lee, S.-I. A Unified Approach to Interpreting Model Predictions. In Proceedings of the 31st International Conference on Neural Information Processing Systems; Curran Associates Inc.: Red Hook, NY, USA, 2017; pp. 4768–4777. [Google Scholar]
  32. Pai, D.G.; Balachandra, M.; Kamath, R. Explainable AI in Agriculture: Review of Applications, Methodologies, and Future Directions. Eng. Res. Express 2025, 7, 032202. [Google Scholar] [CrossRef] [Scilit]
  33. Friedman, J.H.; Popescu, B.E. Predictive Learning via Rule Ensembles. Ann. Appl. Stat. 2008, 2, 916–954. [Google Scholar] [CrossRef] [Scilit]
  34. Obregon, J.; Jung, J.-Y. RuleCOSI+: Rule Extraction for Interpreting Classification Tree Ensembles. Inf. Fusion 2023, 89, 355–381. [Google Scholar] [CrossRef] [Scilit]
  35. Bologna, G. A Rule Extraction Technique Applied to Ensembles of Neural Networks, Random Forests, and Gradient-Boosted Trees. Algorithms 2021, 14, 339. [Google Scholar] [CrossRef] [Scilit]
  36. Bonasera, L.; Carrizosa, E. A Unified Approach to Extract Interpretable Rules from Tree Ensembles via Integer Programming. Comput. Oper. Res. 2026, 185, 107283. [Google Scholar] [CrossRef] [Scilit]
  37. Rathod, S.; Yerram, S.; Arya, P.; Katti, G.; Rani, J.; Padmakumari, A.P.; Somasekhar, N.; Padmavathi, C.; Ondrasek, G.; Amudan, S.; et al. Climate-Based Modeling and Prediction of Rice Gall Midge Populations Using Count Time Series and Machine Learning Approaches. Agronomy 2021, 12, 22. [Google Scholar] [CrossRef] [Scilit]
  38. Reddy, B.N.K.; Rathod, S.; Sridhar, Y.; Kallakuri, S.; Pandit, P.; Jyostna, B.; Malathi, S.; Kumar, R.S.; Sharma, S.; Karthikeyan, K.; et al. An Innovative Two-Stage, Zero-Inflated, Hybrid Count Time Series Model for Predicting Rice Yellow Stem Borer Using Weather Parameters in Hotspot Locations of India. Smart Agric. Technol. 2025, 12, 101381. [Google Scholar] [CrossRef] [Scilit]
  39. Minruhi, P.; Kumari, P.L.; Rathod, S.; Murthy, B.R.; Devaki, K.; Choudhary, K.; Shankar, S.V.; Kumar, P. Machine Learning–Based Hybrid Framework for Weather-Driven Forewarning of Major Groundnut Pests. Int. J. Biometeorol. 2026, 70, 146. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  40. Broadhead, G.T.; Higbee, B.S.; Beck, J.J. Evaluating the Use of In-Season Measures of Pest Abundance to Predict End-of-Season Damage: A Study in Commercial Almond (Prunus dulcis). Pest Manag. Sci. 2025, 81, 1324–1332. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  41. Golan, K.; Kot, I.; Kmieć, K.; Górska-Drabik, E. Approaches to Integrated Pest Management in Orchards: Comstockaspis perniciosa (Comstock) Case Study. Agriculture 2023, 13, 131. [Google Scholar] [CrossRef] [Scilit]
  42. Singh, A.K.; Yeasin, M.; Paul, R.K.; Paul, A.K.; Sarkar, A. Dynamic Ensemble-Based Machine Learning Models for Predicting Pest Populations. Front. Appl. Math. Stat. 2024, 10, 1435517. [Google Scholar] [CrossRef] [Scilit]
  43. Shwartz-Ziv, R.; Armon, A. Tabular Data: Deep Learning Is Not All You Need. Inf. Fusion 2022, 81, 84–90. [Google Scholar] [CrossRef] [Scilit]
  44. Yu, J.; Zheng, W.; Xu, L.; Zhangzhong, L.; Zhang, G.; Shan, F. A PSO-XGBoost Model for Estimating Daily Reference Evapotranspiration in the Solar Greenhouse. Intell. Autom. Soft Comput. 2020, 26, 989–1003. [Google Scholar] [CrossRef] [Scilit]
  45. Ryo, M. Explainable Artificial Intelligence and Interpretable Machine Learning for Agricultural Data Analysis. Artif. Intell. Agric. 2022, 6, 257–265. [Google Scholar] [CrossRef] [Scilit]
  46. Rajbongshi, A.; Johora, F.T.; Hossain, A.; Sarker, M.S.; Rahman, M.H.; Rahman, M.W.; Alotaibi, F.T.; Moni, M.A. Leveraging Explainable AI for Sustainable Agriculture: A Comprehensive Review of Recent Advances. Artif. Intell. Rev. 2026, 59, 105. [Google Scholar] [CrossRef] [Scilit]
  47. Mo, J.; Stevens, M.M. A Degree-Day Model for Predicting Adult Emergence of the Citrus Gall Wasp, Bruchophagus fellis (Hymenoptera: Eurytomidae), in Southern Australia. Crop Prot. 2021, 143, 105553. [Google Scholar] [CrossRef] [Scilit]
  48. Barker, B.S.; Coop, L.; Wepprich, T.; Grevstad, F.; Cook, G. DDRP: Real-Time Phenology and Climatic Suitability Modeling of Invasive Insects. PLoS ONE 2020, 15, e0244005. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  49. Rincon, D.F.; Esch, E.D.; Gutierrez-Illan, J.; Tesche, M.; Crowder, D.W. Predicting Insect Population Dynamics by Linking Phenology Models and Monitoring Data. Ecol. Model. 2024, 493, 110763. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.