Next Article in Journal
The lncRNA011760/miR-Novel-91/NIPA2 ceRNA Network Regulates Salinity Stress Response in Sea Cucumber (Apostichopus japonicus)
Next Article in Special Issue
Acoustic Features of Rainbow Trout (Oncorhynchus mykiss) Ingesting Pelletized Feed
Previous Article in Journal
Protective Effects of Carvacrol Against Vibrio harveyi Infection in Sebastes schlegelii and Its Underlying Mechanisms
Previous Article in Special Issue
Fish Resource Assessment in the Huoyanshan Waters of Poyang Lake Using DIDSON and Deep Learning Models
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Prediction of High-Abundance Fishing Grounds for Chub Mackerel (Scomber japonicus) in the Northwest Pacific Ocean and Its Environmental Drivers Based on Interpretable Machine Learning Model

1
College of Information and Technology, Shanghai Ocean University, Shanghai 200090, China
2
Key Laboratory of Fisheries Remote Sensing, Ministry of Agriculture and Rural Affairs, East China Sea Fisheries Research Institute, Chinese Academy of Fishery Sciences, Shanghai 200090, China
*
Authors to whom correspondence should be addressed.
Fishes 2026, 11(5), 274; https://doi.org/10.3390/fishes11050274
Submission received: 25 February 2026 / Revised: 23 April 2026 / Accepted: 2 May 2026 / Published: 6 May 2026
(This article belongs to the Special Issue Technology for Fish and Fishery Monitoring—2nd Edition)

Abstract

Accurate prediction of fishing grounds plays a crucial role in supporting the efficient operation of ocean-going fishing vessels. Based on catch data of Chub Mackerel (Scomber japonicus) and multiple concomitant oceanographic variables from 2014 to 2022 in the Northwest Pacific Ocean, we employed four machine learning methods, including Random Forest (RF; scikit-learn v1.7.2), Extreme Gradient Boosting (XGBoost; xgboost v3.1.3), Light Gradient Boosting Machine (LightGBM; lightgbm v4.6.0) and Categorical Boosting (CatBoost; catboost v1.2.8), to construct a prediction model for high-abundance fishing grounds of Chub Mackerel. After selecting the optimal model through evaluation metrics, we applied the SHapley Additive exPlanations (SHAP; shap v0.44.1) method to visualize and interpret the optimal model, quantifying the importance of environmental factors on high-abundance fishing grounds, thus enhancing the interpretability and credibility of the machine learning model. The results indicated that the catch exhibited significant fluctuations at both interannual and intramonthly scales (p < 0.05). The annual catch showed a phased increasing trend, peaking in 2017 and 2018. Monthly catches were highest in September and October. Evaluated against established performance metrics, the RF model demonstrated the highest predictive performance with the highest values of accuracy and F1-score, 76.33% and 77.73%, Precision 72.81%, Recall 83.36%, ROC-AUC 0.8393, respectively, and was therefore selected as the most suitable for predicting Chub Mackerel fishing grounds. SHAP analysis identified the temporal variables year and month as the most influential predictors, followed by chlorophyll-a concentration (Chl-a), sea surface salinity (SSS), and sea surface temperature (SST). SHAP analysis can comprehensively reveal the degree and direction of influence of each variable at both global and local levels. These findings indicate that integrating machine learning with explainability techniques can enhance the scientific robustness and transparency of fishing ground forecasts, providing data-driven support for ecosystem-based fishery management.
Key Contribution: We present an interpretable machine learning approach that enhances fishing ground prediction and reveals the dominant drivers of Chub Mackerel distribution.

1. Introduction

The Northwest Pacific Ocean is characterized by a highly dynamic oceanographic system, where the confluence of the Kuroshio Current, pro-tidal currents, and the warm coastal currents of China generate nutrient-rich and diverse salinity conditions. This complex hydrographic environment fosters a wide range of habitats, making the region an important spawning ground, feeding area, and migratory corridor for numerous fish species. Consequently, it ranks among the most significant marine fishing grounds globally [1]. Chub Mackerel (Scomber japonicus) is a typical warm-temperate, oceanic migratory species in the Northwest Pacific, with significant ecological and economic value, supporting offshore fisheries in countries such as China, Japan, and South Korea [2]. From an economic perspective, the catch volume of this species in the region is substantial, with the average annual catch over the past five years ranging from 100,000 to 200,000 tons [3]. This scale of fishing not only provides significant economic benefits to the fishing industries of these countries but also supplies a large amount of marine products to global markets. From an ecological perspective, Chub Mackerel plays a crucial role in the aquatic food web. As an important component of energy transfer and trophic connections, Chub Mackerel plays a key role in the functioning of marine ecosystems. As a highly migratory species, the lifecycle of Chub Mackerel includes spawning in warmer seas, juvenile growth in higher-temperature areas, and adult feeding and reproduction in colder waters [4]. This migratory behavior is strongly influenced by environmental factors leading to significant seasonal and spatial distribution changes. As an important target species in fisheries, Chub Mackerel plays a vital role in the sustainable use of fishery resources, fisheries management, and the balance of marine ecosystems.
In previous studies, Chen et al. [5] used the Habitat Suitability Index (HSI) model to analyze the relationship between Chub Mackerel catch in the East China Sea in summer and environmental factors such as SST, SSS, Chl-a, and sea surface height anomaly. Xu et al. [6] applied a Generalized Additive Model (GAM) to examine the relationship between catch per unit effort (CPUE) of Chub Mackerel in the North Pacific high-seas, purse-seine fishery and environmental variables, including SST and Chl-a. Their research demonstrated that SST and Chl-a had significant effects on CPUE, highlighting the crucial role of environmental factors in the formation of fishing grounds. Chl-a reflects primary productivity and is commonly used as an indicator of phytoplankton biomass. Higher Chl-a levels generally indicate increased food availability for higher trophic levels, thereby indirectly influencing the distribution of Chub Mackerel.
However, the dramatic fluctuations in the marine environment driven by global climate change in recent years have further intensified the uncertainty in the assessment of mackerel resources and the prediction of fisheries in Japan. Traditional fisheries prediction methods, such as the Generalized Additive Model (GAM) [7], exhibit significant limitations in applications involving high-dimensional heterogeneous data, such as fishing vessel trajectories [8] and multi-source remote sensing data [9]. More specifically, they often fail to effectively capture the complex nonlinear interactions and dynamic associations between marine environmental factors and fish distribution. At the same time, issues of data quality—such as data gaps, noise, and inconsistent resolutions—along with potential biases introduced during preprocessing, also limit the effectiveness of these traditional methods in handling high-dimensional heterogeneous data.
In recent years, with the rapid advancement of artificial intelligence, machine learning techniques have been increasingly applied to fishery resource assessment and prediction. Models such as Random Forest (RF) [10] and Light Gradient Boosting Machine (LightGBM) [11] have demonstrated promising performance in fisheries-related prediction. For example, Han et al. [12] examined the prediction of Chub Mackerel fisheries in the Northwest Pacific Ocean and found that incorporating operational characteristics of light fishing vessels, such as the sensitivity of catches to lunar illumination, significantly enhanced the predictive performance of the LightGBM model. Although machine learning (ML) models often achieve superior accuracy, their “black-box” nature (i.e., the internal decision-making process of the model is opaque, making it difficult to easily interpret its predictions) constrains interpretability, thereby limiting their widespread adoption in domains where transparency and decision traceability are crucial.
To address this challenge, SHAP (SHapley Additive exPlanations) has emerged as a powerful interpretability tool [13]. SHAP, grounded in Shapley values from game theory, quantifies the marginal contribution of each input variable to the model output, thereby enabling both local and global interpretability of complex ML models.
At present, SHAP has been widely employed across multiple domains, including climate change modeling [14], and fishery resource prediction [15]. For instance, Lu et al. [16] integrated SHAP with diverse ML algorithms to analyze the influence of environmental variables on classification outcomes. Their results showed that SHAP not only enhanced model interpretability but also offered refined insights into variable importance ranking and sensitivity analysis, thereby deepening understanding of model-based decision processes.
Despite numerous studies investigating the relationship between Chub Mackerel (Scomber japonicus) and marine environmental factors [17,18], and the increasing application of machine learning methods in fishing ground prediction, several limitations remain. On the one hand, most studies mainly focus on improving predictive accuracy, while lacking an interpretation of the internal mechanisms of the models. On the other hand, differences in datasets and evaluation metrics across studies make it difficult to conduct systematic comparisons under a unified framework. To address the above limitations, this study develops an interpretable machine learning framework to predict high-abundance fishing grounds of Chub Mackerel in the Northwest Pacific and to explore the relationships between environmental factors and fish distribution. Using a long-term dataset (2014–2022) that integrates multiple marine environmental variables across a broad spatial extent, four mainstream ensemble learning models (Extreme Gradient Boosting (XGBoost), LightGBM, RF, and Categorical Boosting (CatBoost)) are systematically compared under a unified dataset and evaluation framework, ensuring robust and comparable performance assessment. High-abundance fishing grounds are defined as areas with daily catch greater than or equal to the median. Compared with directly predicting continuous catch values, identifying high-abundance fishing grounds is more effective for highlighting areas of resource aggregation and is more consistent with practical needs in fishing ground forecasting and fishery operation decision-making. After identifying the optimal model, SHAP (SHapley Additive exPlanations) is applied to interpret model outputs from both global and local perspectives, quantifying the contributions of environmental variables and revealing their relative importance and directional effects. This study enhances both predictive performance and interpretability, providing a more comprehensive understanding of the factors associated with the distribution of Chub Mackerel fishing grounds and offering scientific support for spatial fishery management.

2. Materials and Methods

2.1. Data Sources

The fishery data for Chub Mackerel used in this study were obtained from the fishing logbooks of Chinese lighted purse seine vessels, covering the period from April to December during 2014–2022. The data are based on raw catch records, specifically the actual weight of Chub Mackerel caught by fishing vessels during their operations. The study area extends approximately between 145 and 163° E longitude and 35–45° N latitude, with fishing activities conducted in the high seas of the Northwest Pacific Ocean, as illustrated in Figure 1. The dataset comprises 70,147 fishing records. These records include the catch amount for each fishing operation, the operation time, latitude and longitude information, as well as the corresponding environmental variables. The temporal resolution is daily, and the spatial resolution is 0.25° × 0.25°. Figure 2 illustrates the variation in the number of Chinese lighted purse seine vessels from 2014 to 2022. Overall, the number of vessels exhibits a pattern of initial increase followed by fluctuations. Vessel numbers increased rapidly during 2014–2017, declined markedly in 2018–2019, and then rose again from 2020 onward. This temporal pattern directly reflects the stage-dependent fluctuations in fishing intensity over the study period.
Marine environmental variables were obtained from the Copernicus Marine Service Platform (https://resources.marine.copernicus.eu/products, accessed on 3 May 2025). These data were extracted at the same temporal and spatial resolutions as the fishing logbooks and included sea surface temperature (SST), sea surface salinity (SSS), chlorophyll-a concentration (Chl-a), current velocity (CV), dissolved oxygen (DO), sea level anomaly (SLA), and mixed layer depth (MLD).
In particular, C V is derived from the zonal ( U g o s ) and meridional ( V g o s ) geostrophic current velocities using the following formula. These two velocity components represent the flow speed along the latitude and longitude directions, respectively. By combining the velocity components from both directions, CV provides a simplified parameter that describes the overall intensity of water flow, suitable for ocean circulation studies and fishing ground distribution models. In research conducted on larger spatial scales, this combined approach helps simplify the model and enhances its practical applicability, while avoiding excessive refinement of directional flow velocities. This makes the results more concise and easier to interpret.
C V = U g o s 2 + V g o s 2
This study adopts a binary classification approach by calculating the median value of daily catch. Fishing areas with daily catches equal to or exceeding this median are defined as “high-abundance fishing grounds,” while the remaining areas are classified as “low-abundance fishing grounds.” In our dataset, the number of high-abundance samples is 35,075, and the number of low-abundance samples is 35,072, indicating a nearly balanced distribution between the two categories. To comprehensively evaluate model performance, multiple commonly used metrics were calculated, among which the F1-score, integrating both precision and recall, was selected as an important evaluation metric.

2.2. Data Types

For spatiotemporal variables, year, month, and latitude/longitude were selected as the basic indicators to analyze changes in the distribution of fishing grounds, thereby reflecting the temporal evolution and spatial characteristics of Chub Mackerel fishing grounds. For environmental variables, a set of physical and biological factors closely associated with fishery formation was considered, including SST (°C), SSS (‰), Chl-a (mg/m3), CV (m/s), DO (mmol/m3), MLD (m), and SLA (m).
The results indicated that SST, Chl-a concentration, and SSS were the core environmental factors influencing the distribution of Chub Mackerel fisheries in the Northwest Pacific Ocean. Specifically, SST strongly affects feeding activity and migration pathways [19]; Chl-a, as an indicator of phytoplankton biomass, is associated with primary productivity and influences the availability of zooplankton and small pelagic prey, thereby indirectly affecting the distribution of Chub Mackerel [20,21,22] and SSS represents water mass structure, exerting a regulatory effect on the stability of fish habitats [23].
Although relatively few studies have examined the impacts of CV, DO, and MLD on Chub Mackerel fisheries, the literature suggests that these variables may play a significant role in regulating fish behavior and ecosystem processes.
DO [24] concentration directly influences respiratory metabolism and vertical habitat selection; hypoxic zones may constrain the spatial range available to fish. CV [25] not only provides the physical basis for migratory movements but also significantly affects the spatial distribution and aggregation of zooplankton, may influence fish feeding behavior. Interannual fluctuations in MLD [26] regulate the efficiency of nutrient upwelling, which in turn impacts primary productivity in the upper water column and the survival conditions and recruitment success of migratory fish larvae.
Figure 3 presents the histograms of potential marine environmental factors, while Table 1 summarizes their basic statistics (mean, minimum, maximum, and skewness).

2.3. Normalization Process

To eliminate the influence of different measurement scales among environmental variables, all continuous variables were normalized prior to model construction. In this study, Min–Max normalization was applied to rescale the original data to the range of 0–1. The transformation can be expressed as:
X = X X min   X max   X min  
where X   represents the normalized value, and X   denotes the original value of the variables, including fishing catch and environmental variables. X min   and X max   denote the minimum and maximum values of the corresponding variable, respectively. The environmental variables employed in this study exhibit considerable variation in their numerical ranges. Applying standardization is therefore crucial to alleviate potential biases introduced by these scale differences and to enhance the stability and overall performance of the machine learning models.

2.4. Machine Learning Algorithms

All machine learning models were implemented in Python version 3.10. In this study, four widely used machine learning models were selected: XGBoost (xgboost v3.1.3), LightGBM (lightgbm v4.6.0), RF (scikit-learn v1.7.2), and CatBoost (catboost v1.2.8).These models represent the most commonly applied ensemble learning approaches in structured data modeling. RF, proposed by Breiman in 2001, is based on bagging and random feature selection, offering stability and ease of use [27]. XGBoost, officially released in 2016, has gained popularity in data science competitions due to its sparsity awareness, efficient regularization, and parallelization capabilities [28]. LightGBM, also introduced in 2016 by Microsoft, is a gradient boosting decision tree (GBDT) framework that employs a leaf-wise growth strategy along with techniques such as Gradient-based One-Side Sampling (GOSS) and Exclusive Feature Bundling (EFB), making it particularly suitable for large-scale data processing. CatBoost, developed by Yandex in 2017, emphasizes efficient encoding of categorical features and mitigates the risk of information leakage [29]. Collectively, these models are widely applied in domains such as financial risk control, recommender systems, click-through rate (CTR) estimation, and healthcare, and are considered indispensable benchmark methods in contemporary machine learning practice.

2.5. SHAP Interpretation Methods

SHapley Additive exPlanations (SHAP) has emerged as an important method for interpreting machine learning models [30], particularly tree-based models such as XGBoost, LightGBM, RF, and CatBoost, since its introduction in 2017. In this study, SHAP analysis was implemented using the Python package shap v0.44.1. It is grounded in the Shapley value from game theory and is characterized by theoretical fairness and local accuracy. Through the TreeSHAP and FastTreeSHAP algorithms, SHAP enables efficient interpretation of complex tree models, providing clear explanations of single-sample prediction processes while also revealing global feature importance and feature interactions. In practice, SHAP not only enhances model interpretability but is also widely applied in feature selection and model optimization, making it a cornerstone of explainable AI. The core formulation of SHAP is derived from the Shapley value equation in game theory, which quantifies the “average marginal contribution” of each feature to the model output. For a model f, the SHAP value of feature i is defined as follows.
ϕ i f , x = S N { i } S ! N S 1 ! N ! f S { i } x S { i } f S x S
where N denotes the set of all features; S denotes any subset that does not contain feature i; S denotes the number of features in subset S; f S x S represents the model prediction using only the feature subset S. ϕ i denotes the SHAP value of feature i, its average marginal contribution to the model output.

2.6. Variable Selection and Predictive Performance Evaluation

Feature selection is a critical step in machine learning, as it helps reduce training time, decrease model complexity, and enhance prediction performance. A Pearson correlation coefficient close to ±1 suggests potential redundancy between variables, which may diminish the effective contribution of certain variables during model training [31]. In this study, Pearson correlation coefficients were calculated among all variables, as shown in Figure 4. The results indicate generally weak correlations between most variables, with a maximum absolute value not exceeding 0.87, which is below the common collinearity exclusion threshold of 0.9.
The variance inflation factor (VIF) is a key method for detecting multicollinearity among variables, quantifying the extent to which the variance of a variable is inflated due to its linear dependence on other predictors, and is regarded as a robust metric for assessing collinearity in regression models. The VIF values of the variables are shown in Figure 5. Specifically, the VIF values for latitude, longitude, and month are 9.3, 5.9, and 6.2, respectively.
Referring to the study [32], which advises caution regarding rigid rules of thumb for VIF thresholds, although the VIF values for month, longitude, and latitude are relatively high, they are all below 10, indicating no severe multicollinearity. Moreover, these three variables represent important spatiotemporal factors influencing the distribution of Chub Mackerel (Scomber japonicus) fishing grounds and are therefore retained in the modeling process. The VIF values for other environmental variables are all relatively low, confirming their suitability for inclusion in subsequent model construction.
The dataset was divided into a training set (70%) and a testing set (30%) using a stratified random splitting approach, ensuring that the proportions of the “high-abundance” and “low-abundance” classes in both subsets remained consistent with those in the original dataset, thereby avoiding potential biases in model training and evaluation caused by class imbalance.
In binary classification tasks such as fishing ground prediction, model performance is commonly assessed using four basic statistics derived from the confusion matrix: true positives (TP), false positives (FP), true negatives (TN), and false negatives (FN) [33]. These values are then used to calculate key performance metrics, including accuracy, precision, recall, and the F1-score, which are essential for model tuning and result interpretation.
A c c u r a c y = T P + T N T P + T N + F P + F N
P r e c i s i o n = T P T P + F P
R e c a l l = T P T P + F N
F 1 - score = 2 × Recall × Precision Recall + Precision
Among them, TP represents the number of samples correctly predicted by the model to be in the positive category; TN represents the number of samples correctly predicted by the model to be in the negative category; FP represents the number of samples incorrectly predicted by the model to be in the positive category; and FN represents the number of samples incorrectly predicted by the model to be in the negative category. Accuracy represents the proportion of overall correct predictions made by the model; precision reflects how many of the classes predicted as positive are truly positive; recall measures the ability of the model to identify all samples in the positive class; and the F1-score is the harmonic mean of precision and recall, which is used to comprehensively evaluate the performance of the model on imbalanced datasets. Additionally, to examine the trend variations in annual and monthly production, we conducted the Mann–Kendall trend test. This method is a non-parametric statistical approach used to detect monotonic trends (either increasing or decreasing) in time series data. The test was performed using the “Kendall” package in R software (version 4.0.3).

3. Results

3.1. Chub Mackerel Catch from 2014 to 2022

Spatiotemporal variations in these environmental factors are further reflected in the dynamic characteristics of catch amounts. Figure 6 illustrates the interannual and intermonthly variations in Chub Mackerel catch from 2014 to 2022. Interannually (Figure 6a), the catch fluctuated considerably (p < 0.05), with 2017 and 2018 representing peak years, each exceeding 120,000 tons, while 2014 recorded the lowest catch at less than 30,000 tons. Overall, the catch exhibited a phased upward trend.
Figure 6b shows the cumulative monthly catch aggregated across all years (2014–2022). Based on the distribution of monthly catch values, relatively higher catches are observed from June to November, with September and October representing the peak months. Catch levels during this period are consistently higher than those in earlier months (e.g., April–May), indicating a clear seasonal concentration of fishing activity.
This variation pattern is likely associated with the combined effects of seasonal and interannual fluctuations in the aforementioned environmental factors, the migratory behavior of Chub Mackerel, and corresponding fisheries management measures.

3.2. Comparison of Model Performance

In this study, four machine learning models—XGBoost, LightGBM, CatBoost, and RF—were selected for comparative analysis. The key hyperparameters of each model were optimized using a grid search strategy, with parameter selection performed via three-fold cross-validation on the training set. Overall, the performance of the four models across multiple evaluation metrics was relatively similar, with accuracy mainly ranging between 75% and 85%. Figure 7 presents a comparison of the four models across several evaluation metrics. In terms of metric structure, all models exhibited higher recall than precision, indicating a tendency to identify positive samples more readily during prediction. Among all models, RF demonstrated a relatively balanced trade-off between precision, recall, and F1-score, with a slightly higher F1-score, reflecting more stable predictive performance.
Figure 8 displays the ROC curves of each model. It can be observed that the ROC curves of the four models are highly overlapping, with only minor differences in AUC values, suggesting essentially consistent overall discriminative ability. Among them, RF achieved a slightly higher AUC value. Table 2 further provides the optimal hyperparameter configurations for each model determined via grid search and their corresponding evaluation results. All performance metrics were calculated based on an independent 30% hold-out test set, while three-fold cross-validation was used solely for hyperparameter optimization.

3.3. Model Interpretability

3.3.1. Global Feature Interpretation

To evaluate the relative contributions of different environmental factors in predicting the target variable, the optimized RF model was employed in this study. The input variables included temporal characteristics (year, month), environmental factors (SST, SSS, DO, MLD, Chl-a, CV, SLA), and geographic location (longitude, latitude), resulting in a total of 12 candidate variables. The RF model was tuned using grid search with cross-validation (GridSearchCV) to determine the optimal tree depth, split criterion, and maximum number of features. After training, model interpretability was analyzed using the TreeSHAP method to compute SHAP values for each feature. SHAP values quantify both the magnitude and direction of each feature’s contribution to the model output for individual predictions, thereby providing a robust approach to interpreting machine learning models.
The SHAP global interpretation plot, based on Shapley value theory, serves as an important tool for visualizing and quantifying the dependence of the model on each feature across the entire dataset. By aggregating SHAP values over all samples, this plot reveals the overall contributions of features, their relative importance rankings, and the distribution patterns of their influence on model predictions.
In the swarm plot (Figure 9), the x-axis represents SHAP values, which indicate the positive or negative contribution of each feature for individual samples. Dot colors correspond to the magnitude of feature values (blue = low, red = high), while the horizontal spread of the dots reflects the direction and intensity of feature influence. Results show that year and month contribute most significantly to the model output, as evidenced by their widest distribution of SHAP values and most pronounced impact. This may be associated with the scale of industrial development and the implementation of conservation management measures by regional fisheries management organizations. Environmental variables such as Chl-a, SSS, DO, and SST followed, all showing stable positive contributions.
In the bar chart (Figure 10), mean SHAP values provide a direct quantification of each feature’s average contribution across the dataset. Consistent with the swarm plot, year had the highest average contribution, followed by month, Chl-a, SSS, DO, and SST. In contrast, longitude, latitude, MLD, SLA, and CV had relatively smaller mean contributions, indicating limited global influence on model predictions.
The global decision plot (Figure 11) illustrates the cumulative contribution path from the baseline output to the final predicted values for each sample. Across all samples, year and month consistently played the leading roles, rapidly driving the predictions toward their outcomes. Environmental variables such as Chl-a and SSS provided supplementary contributions in the mid-to-late stages, while the remaining features contributed little to most predictions.
In summary, the three plots collectively demonstrate that the model is highly dependent on temporal features (year and month), while environmental features (Chl-a, SSS, DO, SST) provide complementary support. In contrast, spatial variables and other environmental factors play only minor roles, indicating that the model’s predictive process is primarily driven by temporal dynamics in combination with a few key environmental indicators.

3.3.2. Localized Feature Interpretation

The SHAP local interpretation plot is used to reveal the contribution paths of individual features in the model prediction process. For a specific sample, it illustrates the cumulative contribution of each feature relative to the baseline model output, thereby explaining the generation mechanism of the prediction for that sample. The SHAP force plot (Figure 12) visualizes both the direction and magnitude of each environmental variable’s influence on the model output using color coding (red = positive contribution, blue = negative contribution). To intuitively illustrate the local decision-making mechanisms of the model, this study employs SHAP force plots to interpret two representative test samples rather than randomly selected instances. Sample 1 corresponds to a case with a high predicted probability of belonging to the high-abundance class, whereas Sample 0 represents a case with a low predicted probability. For each sample, the SHAP values and the model baseline output are extracted, and the cumulative feature contribution path from the baseline to the final prediction is computed. By comparing these two samples located at opposite ends of the prediction spectrum, the direction and magnitude of the contributions of environmental variables under high- and low-abundance prediction scenarios can be more clearly visualized, thereby elucidating the model’s decision rationale.
Figure 12a presents an example predicted as high mackerel abundance (Sample 1), whereas Figure 12b shows an example predicted as low abundance (Sample 0). In Figure 12a, variables such as DO, Chl-a, SLA, and MLD exerted significant positive effects on the prediction, pushing the model output value close to 1. This indicates that these factors jointly contributed to classifying the sample into the high-abundance fishing zone. By contrast, in Figure 12b, variables such as latitude, Chl-a, and salinity showed strong negative effects, driving the final prediction well below the baseline value and suggesting that the sample was more likely to belong to the low-abundance zone.
The SHAP waterfall plot is used to visualize the stepwise formation of a model’s predicted values. Figure 13 illustrates how the model prediction for Sample 1 is constructed through the combined effects of the baseline output and the additive contributions of individual features. The baseline value corresponds to the model’s expected predicted probability for the high-abundance class in the training data (E[f(X)] ≈ 0.50), rather than a fixed probability of 0.5. The results indicate that year and month are the most important positive contributing factors for this sample, increasing the predicted probability relative to the baseline by approximately +0.13 and +0.05, respectively, ultimately raising the model output to 0.788. Environmental variables such as Chl-a, SSS, and SST also provide smaller but stable positive contributions, whereas other features have a negligible influence on this specific prediction.

4. Discussion

4.1. Analysis of Changes in Chub Mackerel Catch from 2014 to 2022

The interannual fluctuations in Chub Mackerel catches between 2014 and 2022 exhibit complex, multifactorial trends. Resource dynamics during this period were reflected not only in quantitative increases and decreases but also in deeper changes in ecosystem structure and fishery practices. Several factors may explain these variations: (1) shifts in the marine environment under global warming, particularly rising seawater temperatures and changes in the position of the Kuroshio Extension, significantly influenced habitat selection and spawning ground distribution, leading to regional redistribution and aggregation of stocks [34]; (2) Food competition and density effects between Japanese sardine and Chub Mackerel substantially impact the population dynamics of Chub Mackerel. An increase in sardine abundance may reduce food availability for Chub Mackerel, thereby affecting their growth and catch levels [35]. (3) The high catch levels in the high seas of the Northwest Pacific are closely linked to high fishing intensity, which may be one of the factors contributing to the sharp decline in catches in 2019 [36].
Regarding intermonthly fluctuations, the study confirmed that Chub Mackerel aggregated in the southwestern Sea of Japan from October to December before migrating southward to the East China Sea and adjacent offshore waters for wintering. The SST ranges and optimal intervals for Chub Mackerel fishing grounds differ between spring and summer. In spring, fishing grounds are found within an SST range of 7–19 °C, with an optimal SST of 11–15 °C. In summer, the SST range is 8–24 °C, with an optimal SST of 8–12 °C. This demonstrates a clear seasonal preference for specific temperature conditions [37].
In other regions, Atlantic mackerel (Scomber scombrus) also exhibits a significant preference for the thermocline and specific temperature ranges (8–12 °C), particularly during the colder seasons, when they adjust their habitat distribution in response to changes in the thermocline [38]. This behavior is similar to the monthly fluctuations observed in Chub Mackerel (Scomber japonicus) in the Northwest Pacific. The fishing grounds of Chub Mackerel in the Northwest Pacific Ocean are extensive, mainly concentrated between 39 and 43° N latitude and 149–154° E longitude. Key fishing areas include the high seas east of Honshu, the eastern coast of Japan, the northern East China Sea, the eastern edge of the Yellow Sea, and the Russian Far East [39], consistent with the findings of this study.
In summary, the spatial and temporal dynamics of Chub Mackerel catches reveal pronounced interannual and intermonthly variability. These findings provide an ecological basis for stock assessment and fishery management, emphasizing that fishery strategies should be more closely aligned with the natural migratory rhythms of the stock and the broader context of climate change.

4.2. Prediction Performance Analysis of Four Machine Learning Models

The applications of XGBoost, LightGBM, CatBoost, and RF in environmental prediction have grown steadily, demonstrating excellent modeling performance, particularly in nonlinear, multivariate systems where they show clear advantages. In fisheries, Han et al. [12] applied RF and LightGBM to operational and environmental data from 2014 to 2022 and achieved a prediction accuracy of approximately 75%. They further noted that variable importance changed significantly under different light conditions, suggesting that environmental factors such as sea surface height anomaly, salinity, and dissolved oxygen play key roles in determining fishing ground distribution.
Further research has also confirmed this, several studies have consistently shown that Chl-a, SSS, DO, and SST are the primary factors influencing fishery distribution in machine learning-based prediction models. Zhang et al. [40], for example, found that Chl-a was the most critical factor in January, while latitude exerted the strongest influence in year-round predictions of albacore tuna fisheries in the South Pacific Ocean using a RF model for feature importance analysis. Fang et al. [41] in a study of squid fisheries of Chile using the HSI model, reported that the combination of SST and SSS outperformed other variable configurations in predictive accuracy. Similarly, Zhang et al. [42] employed LightGBM and CatBoost in combination with SHAP interpretation to evaluate environmental variable importance in predicting Atlantic albacore tuna fisheries. Their results indicated that temperature and dissolved oxygen at a depth of 100 m were the most influential variables across multiple spatial scales, with SSS playing a secondary role.
Overall, while species-specific ecological traits and regional characteristics affect the ranking of variable importance, Chl-a, SSS, DO, and SST consistently emerge as the core environmental parameters for constructing high-performance fishery prediction models. Our results are broadly consistent with previous findings, and we emphasize that careful selection of optimal hyperparameters and appropriate normalization of environmental variables are essential when building such models.
The results of this study indicate that the importance of the “year” and “month” variables is higher than that of other marine environmental factors. This is primarily because interannual variations are mainly driven by the stages of fishery development and management policies. The light purse seine fishery in the Northwest Pacific of China began to develop in 2014, with limited fishing vessel scale and immature technology during the early stages, resulting in lower production. As the scale of operations expanded, production increased annually. By 2019, production experienced a temporary downturn due to industrial restructuring; however, following adaptive adjustments in the industry and the resource management measures implemented by the North Pacific Fisheries Management Organization (NPFC), production gradually increased and stabilized at around 100,000 tons. On a monthly scale, Chub Mackerel, as a typical seasonal migratory species, has a significantly concentrated fishing season, primarily from June to November, resulting in noticeable fluctuations in monthly production. These factors led to significant variations in Chub Mackerel production across years and months, directly influencing its SHAP values in the model.
However, this study has several limitations: (1) the catch data used were derived from existing fishery statistics, which did not cover all potential fishing areas and thus introduced spatial gaps; (2) although the input variables encompassed a broad range of marine environmental parameters, they did not include dynamic elements such as currents, cyclonic activity, ocean front positions, or extreme meteorological events, which may strongly influence fishing ground distribution at specific spatial and temporal scales; (3) socio-economic factors—including changes in fishing behavior, policy adjustments, and the intensity of fishing vessel operations—were not sufficiently considered, which may introduce prediction bias. (4) While this study has employed the SHAP method for interpretation, it has not explicitly captured the potential interactions among environmental factors (such as the synergistic effect between SST and Chl-a), nor has it quantitatively assessed prediction uncertainty, which to some extent limits the depth of ecological interpretation. Future research should integrate multi-source environmental dynamics and socioeconomic drivers, incorporating interaction features, uncertainty quantification methods, and additional dynamic environmental elements to enhance the accuracy and practical applicability of fishery prediction models.

4.3. Recommendations for the Management of the Chub Mackerel Fishery in the Northwest Pacific Ocean

Chub Mackerel is one of the most economically valuable migratory pelagic fish species in the Northwest Pacific, and its catch plays an important role in the fisheries sectors of several coastal countries. Owing to its extensive migratory behavior and highly sensitive ecological traits, the population dynamics of this species are not only directly influenced by fishing intensity but also shaped by complex environmental factors such as variations in ocean temperature, salinity, and currents, all of which exhibit strong spatiotemporal variability. Therefore, establishing a scientific, systematic, and adaptive management framework is of great practical importance to ensure the sustainable development of Chub Mackerel resources.
To achieve effective management of this species, a systematic governance framework should be developed at multiple levels, encompassing fishing control, resource assessment and optimization, and regional cooperation mechanisms. Firstly, given the significant spatial and temporal differences in population distribution, differentiated quota strategies should be implemented based on seasons and geographical locations. Seasonal fishing bans or total allowable catch (TAC) limits should be enforced in high-density areas. At the same time, fishing intensity and vessel capacity should be strictly controlled to reduce the risk of overfishing. Secondly, to enhance the adaptability and foresight of management, key environmental factors such as sea temperature, salinity, and mixed-layer depth should be incorporated into the resource assessment framework, thereby promoting an ecosystem-coupled management approach. By systematically integrating high-resolution oceanographic variables, satellite environmental indices, and fleet behavior data, model prediction accuracy can be effectively improved, providing more reliable scientific support for fisheries resource management. Thirdly, full use should be made of the NPFC and its technical working groups to regularly update resource status assessments, facilitate joint data collection, model evaluation, and policy consultations among member states, and advance the process of transnational cooperative governance. Collectively, these comprehensive measures are expected to ensure the long-term sustainable utilization of Chub Mackerel resources while maintaining the economic viability of the fishery.

5. Conclusions

Based on the analysis of Chub Mackerel catch data in the Northwest Pacific from 2014 to 2022, this study reveals pronounced interannual fluctuations in catch volume, which reached the peak in 2017–2018 and significantly decreased in 2019. At the monthly scale, catches were concentrated from June to November, peaking in September–October. Among the four machine learning models evaluated, the RF model demonstrated the highest performance in predicting fishing ground distribution. Interpretability analysis based on the SHAP framework further indicated that temporal variables (year and month) contributed most significantly to the predictions, while environmental factors such as Chl-a, SSS, DO, and SST also played important roles.
While this study has yielded certain findings, it is not without limitations. The predictive performance of the model is, to some extent, contingent upon the type and quality of the input environmental variables. Certain potential critical factors (such as higher-resolution ocean dynamic processes) have yet to be incorporated into the analysis. Moreover, existing data still exhibit constraints in terms of temporal and spatial resolution, which may affect the model’s ability to accurately capture local-scale variations. Future research should integrate multi-source dynamic environmental data including ocean currents, mesoscale eddies, and oceanic fronts as well as socio-economic factors such as fisheries’ management policies and fishing vessel operation patterns, in order to construct a more robust and spatiotemporally adaptive prediction framework. Such an integrated modeling approach will provide more reliable scientific support for the sustainable management of fishery resources under climate change.

Author Contributions

L.Z.: Conceptualization, Data curation, Modeling analysis, Software, Writing—original draft. W.F.: Review and comment. F.T.: Conceptualization, Modeling analysis. Y.S.: Review and comment, Data processing. S.Z.: Review and comment. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the National Natural Science Foundation of China (32403030), National Key R&D Program of China (2023YFD2401305).

Institutional Review Board Statement

The data used in this study include commercial fishing operation records, and marine environmental data. As no animal experiments were involved, approval from the local ethics committee is not required.

Data Availability Statement

Data will be made available on request.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Zhao, G.; Zhang, H.; Tang, F. The spatio-temporal distribution and population dynamics of Chub mackerel (Scomber japonicus) in the high seas of the Northwest Pacific Ocean. Animals 2025, 15, 1135. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  2. Wang, L.; Ma, S.; Liu, Y.; Li, J.; Liu, S.; Lin, L.; Tian, Y. Fluctuations in the abundance of Chub mackerel in relation to climatic/oceanic regime shifts in the Northwest Pacific Ocean since the 1970s. J. Mar. Syst. 2021, 218, 103541. [Google Scholar] [CrossRef] [Scilit]
  3. Hong, J.; Kim, D.; Kim, D. Stock assessment of Chub mackerel (Scomber japonicus) in the Northwest Pacific Ocean based on catch and resilience data. Sustainability 2022, 15, 358. [Google Scholar] [CrossRef] [Scilit]
  4. Wang, Y.; Zheng, J.; Yu, C. Stock assessment of Chub mackerel (Scomber japonicus) in the central East China Sea based on length data. J. Mar. Biol. Assoc. 2013, 94, 211–217. [Google Scholar] [CrossRef] [Scilit]
  5. Chen, X.; Li, G.; Feng, B.; Tian, S. Habitat suitability index of Chub mackerel (Scomber japonicus) from July to September in the East China Sea. J. Oceanogr. 2008, 65, 93–102. [Google Scholar] [CrossRef] [Scilit]
  6. Xu, B.; Zhang, H.; Tang, F.H.; Sui, X.; Zhang, Y.Y.; Hou, G. Relationship between center of gravity and environmental factors of main catches of purse seine fisheries in the North Pacific high seas based on GAM. South China Fish. Sci. 2020, 16, 60–70. (In Chinese) [Google Scholar] [CrossRef]
  7. Okunishi, T.; Yokouchi, K.; Hasegawa, D.; Yokouchi, K.; Takasuka, A. Relationship between sea temperature variation and fishing ground formations of Chub mackerel in the Pacific Ocean off Tohoku. Bull. Jpn. Soc. Fish. Oceanogr. 2020, 84, 271–284. [Google Scholar] [CrossRef]
  8. Liang, M.; Liu, R.W.; Gao, R.; Xiao, Z.; Zhang, X.; Wang, H. A survey of distance-based vessel trajectory clustering. arXiv 2024, arXiv:2407.11084. [Google Scholar] [CrossRef] [Scilit]
  9. Guo, Y.; Liu, R.W.; Qu, J.; Lu, Y.; Zhu, F.; Lv, Y. Asynchronous trajectory matching-based multimodal maritime data fusion. IEEE Trans. Intell. Transp. Syst. 2023, 24, 12779–12792. [Google Scholar] [CrossRef] [Scilit]
  10. Behivoke, F.; Etienne, M.-P.; Guitton, J.; Randriatsara, R.M.; Ranaivoson, E.; Léopold, M. Estimating fishing effort in small-scale fisheries using GPS tracking data and random forests. Ecol. Indic. 2021, 123, 107321. [Google Scholar] [CrossRef] [Scilit]
  11. Gong, P.; Wang, D.; Yuan, H.; Chen, G.; Wu, R. Fishing ground forecast model of albacore tuna based on LightGBM. Fish. Sci. 2021, 40, 762–767. [Google Scholar]
  12. Han, H.; Shang, C.; Jiang, B.; Wang, Y.; Li, Y.; Xiang, D.; Zhang, H.; Shi, Y.; Jiang, K. Predicting fishing grounds of Chub mackerel using machine learning. Front. Mar. Sci. 2024, 11, 1451104. [Google Scholar] [CrossRef] [Scilit]
  13. Lundberg, S.M.; Lee, S.I. A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2017; pp. 4768–4777. [Google Scholar]
  14. Durap, A. Interpretable machine learning for coastal wind prediction. J. Coast. Conserv. 2025, 29, 24. [Google Scholar] [CrossRef] [Scilit]
  15. Shi, Y.; Yan, L.; Zhang, S.; Tang, F.; Yang, S.; Fan, W.; Han, H.; Dai, Y. Revealing the effects of environmental and spatio-temporal variables on Japanese sardine fishing grounds using interpretable machine learning. Front. Mar. Sci. 2025, 11, 1503292. [Google Scholar] [CrossRef] [Scilit]
  16. Lu, R.; Liu, S.; Duan, H.; Kang, W.; Zhi, Y. Combining SHAP and machine learning for desert type extraction. Remote Sens. 2024, 16, 4414. [Google Scholar] [CrossRef] [Scilit]
  17. Yu, W.; Guo, A.; Zhang, Y.; Chen, X.; Qian, W.; Li, Y. Climate-induced habitat suitability variations of Chub mackerel. Fish. Res. 2018, 207, 63–73. [Google Scholar] [CrossRef] [Scilit]
  18. Perrotta, R.G.; Viñas, M.D.; Hernández, M.; Tringali, L. Temperature conditions in Chub mackerel fishing grounds. Fish. Oceanogr. 2001, 10, 275–283. [Google Scholar] [CrossRef] [Scilit]
  19. Yasuda, T.; Kinoshita, J.; Niino, Y.; Okuyama, J. Vertical migration patterns in Chub mackerel. Prog. Oceanogr. 2023, 213, 103017. [Google Scholar] [CrossRef] [Scilit]
  20. Park, H.-S.; Song, S.H.; Jeong, J.M.; Yang, J.H.; Kim, C. Feeding characteristics of Chub mackerel. Water 2025, 17, 1804. [Google Scholar] [CrossRef] [Scilit]
  21. Yoon, S.J.; Kim, D.H.; Baeck, G.W.; Kim, J.W. Feeding habits of Chub mackerel. Korean J. Fish. Aquat. Sci. 2008, 41, 26–31. [Google Scholar] [CrossRef] [Scilit]
  22. Kume, G.; Shigemura, T.; Okanishi, M.; Hirai, J.; Shiozaki, K.; Ichinomiya, M.; Komorita, T.; Habano, A.; Makino, F.; Kobari, T. Distribution and feeding of Chub mackerel larvae. Front. Mar. Sci. 2021, 8, 725227. [Google Scholar] [CrossRef] [Scilit]
  23. Ayman, J.; Andreas, R.; Bouchta, E.M.; Abdellah, A.; Mustapha, S.; Khalid, M. Environmental factors and catches of Atlantic Chub mackerel. Egypt. J. Aquat. Biol. Fish. 2024, 28, 1727–1750. [Google Scholar]
  24. Politikos, D.V.; Petasis, G.; Katselis, G. Interpretable machine learning for hypoxia prediction. Ecol. Inform. 2021, 66, 101480. [Google Scholar] [CrossRef] [Scilit]
  25. Chambers, P.A.; Prepas, E.E.; Hamilton, H.R.; Bothwell, M.L. Current velocity and aquatic macrophytes. Ecol. Appl. 1991, 1, 249–257. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  26. Kane, J. Zooplankton abundance trends. ICES J. Mar. Sci. 2007, 64, 909–919. [Google Scholar] [CrossRef] [Scilit]
  27. Scornet, E.; Biau, G.; Vert, J.P. Consistency of random forests. Ann. Stat. 2015, 43, 1716–1741. [Google Scholar] [CrossRef] [Scilit]
  28. Chen, T.; Guestrin, C. XGBoost: A scalable tree boosting system. In Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining; Association for Computing Machinery: New York, NY, USA, 2016; pp. 785–794. [Google Scholar] [CrossRef] [Scilit]
  29. Prokhorenkova, L.; Gusev, G.; Vorobev, A.; Dorogush, A.V.; Gulin, A. CatBoost: Unbiased boosting. In NeurIPS; Curran Associates, Inc.: Red Hook, NY, USA, 2018. [Google Scholar]
  30. Lundberg, S.M.; Erion, G.; Chen, H.; DeGrave, A.; Prutkin, J.M.; Nair, B.; Katz, R.; Himmelfarb, J.; Bansal, N.; Lee, S.-I. From local explanations to global understanding with explainable AI for trees. Nat. Mach. Intell. 2020, 2, 56–67. [Google Scholar] [CrossRef] [Scilit]
  31. Benesty, J.; Chen, J.; Huang, Y.; Cohen, I. Pearson correlation coefficient. In Noise Reduction in Speech Processing; Springer: Berlin, Germany, 2009; pp. 1–4. [Google Scholar]
  32. O’Brien, R.M. A caution regarding variance inflation factors. Qual. Quant. 2007, 41, 673–690. [Google Scholar] [CrossRef] [Scilit]
  33. Obi, J.C. A comparative study of classification metrics. World J. Adv. Eng. Technol. Sci. 2023, 8, 308–314. [Google Scholar] [CrossRef] [Scilit]
  34. Kanamori, Y.; Takasuka, A.; Nishijima, S.; Okamura, H. Climate change shifts spawning ground of Chub mackerel. Mar. Ecol. Prog. Ser. 2019, 624, 155–166. [Google Scholar] [CrossRef] [Scilit]
  35. Kamimura, Y.; Taga, M.; Yukami, R.; Watanabe, C.; Furuichi, S. Density dependence of Chub mackerel. ICES J. Mar. Sci. 2021, 78, 3254–3264. [Google Scholar] [CrossRef] [Scilit]
  36. Zhao, G.Q.; Wu, Z.L.; Cui, X.S.; Tang, F.H.; Fan, W. Spatial-temporal patterns of Chub mackerel fishing ground. Haiyang Xuebao 2022, 44, 22–35. [Google Scholar]
  37. Wang, L.M.; Li, Y.; Zhang, R.; Tang, F.H.; Fan, W. Relationship between Chub mackerel distribution and seawater temperature. Ocean Univ. China 2019, 49, 29–38. (In Chinese) [Google Scholar] [CrossRef]
  38. Nikoloudakis, N.; Skaug, H.J.; Olafsdottir, A.H.; Jansen, T.; Jacobsen, J.A.; Enberg, K. Drivers of mackerel distribution. ICES J. Mar. Sci. 2019, 76, 530–548. [Google Scholar] [CrossRef] [Scilit]
  39. Hiyama, Y.; Yoda, M.; Ohshimo, S. Stock size fluctuations in Chub mackerel. Fish. Oceanogr. 2002, 11, 347–353. [Google Scholar] [CrossRef] [Scilit]
  40. Zhang, J.; Fan, D.; He, H.; Xiao, B.; Xiong, Y.; Shi, J. Forecasting albacore fishing grounds using machine learning. Appl. Sci. 2023, 13, 5485. [Google Scholar] [CrossRef] [Scilit]
  41. Fang, X.Y.; Chen, X.J.; Ding, Q. Fishing ground prediction of Dosidicus gigas. J. Guangdong Ocean Univ. 2014, 34, 67–73. [Google Scholar]
  42. Zhang, T.; Guo, H.; Song, L.; Yuan, H.; Sui, H.; Li, B. Evaluating the importance of vertical environmental variables for albacore fishing grounds in the tropical Atlantic Ocean using machine learning and Shapley additive explanations (SHAP). Fish. Oceanogr. 2025, 34, e12701. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Distribution of Chub Mackerel fisheries in the study area.
Figure 1. Distribution of Chub Mackerel fisheries in the study area.
Fishes 11 00274 g001
Figure 2. Trend in Chinese lighted purse seine vessels (2014–2022).
Figure 2. Trend in Chinese lighted purse seine vessels (2014–2022).
Fishes 11 00274 g002
Figure 3. Histogram of potential marine environmental factors. SST: sea surface temperature; SSS: sea surface salinity; Chl-a: chlorophyll-a concentration; CV: current velocity; DO: dissolved oxygen; SLA: sea level anomaly; MLD: mixed layer depth.
Figure 3. Histogram of potential marine environmental factors. SST: sea surface temperature; SSS: sea surface salinity; Chl-a: chlorophyll-a concentration; CV: current velocity; DO: dissolved oxygen; SLA: sea level anomaly; MLD: mixed layer depth.
Fishes 11 00274 g003
Figure 4. Pearson correlation matrix of the different variables.
Figure 4. Pearson correlation matrix of the different variables.
Fishes 11 00274 g004
Figure 5. VIF for each feature.
Figure 5. VIF for each feature.
Fishes 11 00274 g005
Figure 6. Interannual total catch (a) and monthly catch (b) of Chub Mackerel from 2014 to 2022.
Figure 6. Interannual total catch (a) and monthly catch (b) of Chub Mackerel from 2014 to 2022.
Fishes 11 00274 g006
Figure 7. Comparison chart of assessment metrics for the four models.
Figure 7. Comparison chart of assessment metrics for the four models.
Fishes 11 00274 g007
Figure 8. Four models’ ROC curves with AUC values.
Figure 8. Four models’ ROC curves with AUC values.
Fishes 11 00274 g008
Figure 9. SHAP swarm map based on RF model prediction.
Figure 9. SHAP swarm map based on RF model prediction.
Fishes 11 00274 g009
Figure 10. SHAP bar chart based on RF model prediction.
Figure 10. SHAP bar chart based on RF model prediction.
Fishes 11 00274 g010
Figure 11. SHAP global decision diagram based on RF model prediction. Colors indicate model output values from low (blue) to high (red).
Figure 11. SHAP global decision diagram based on RF model prediction. Colors indicate model output values from low (blue) to high (red).
Fishes 11 00274 g011
Figure 12. SHAP force diagram of RF-based model predictions. (a) High-abundance sample; (b) low-abundance sample. Colors indicate feature contributions from negative (blue) to positive (red).
Figure 12. SHAP force diagram of RF-based model predictions. (a) High-abundance sample; (b) low-abundance sample. Colors indicate feature contributions from negative (blue) to positive (red).
Fishes 11 00274 g012
Figure 13. SHAP waterfall plot of RF-based model predictions.
Figure 13. SHAP waterfall plot of RF-based model predictions.
Fishes 11 00274 g013
Table 1. Basic statistics of each potential marine environmental factor.
Table 1. Basic statistics of each potential marine environmental factor.
FactorMeanMinMaxSkewness
SLA0.12−0.3710.2
Chl-a0.510.082.11.51
SSS33.4332.1534.910.19
SST15.814.8126.19−0.28
DO245.92208.8313.510.34
MLD17.3210.53138.712.51
CV0.2401.931.76
Table 2. Specific hyperparameters and evaluation metrics of the four models.
Table 2. Specific hyperparameters and evaluation metrics of the four models.
ModelsAccuracyPrecisionRecallF1-ScoreROC-AUCHyperparameter
XGBoost76.3172.9582.9577.630.8384learning_rate: 0.05,
max_depth: 6,
n_estimators: 100
LightGBM76.2472.7683.1877.620.8382learning_rate: 0.05,
max_depth: 6,
n_estimators: 100,
num_leaves: 50
RF76.3372.8183.3677.730.8393max_depth: 10,
n_estimators: 200
CatBoost76.0972.8382.5277.370.8363depth: 6,
iterations: 200,
learning_rate: 0.05
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Zhang, L.; Fan, W.; Tang, F.; Shi, Y.; Zhang, S. Prediction of High-Abundance Fishing Grounds for Chub Mackerel (Scomber japonicus) in the Northwest Pacific Ocean and Its Environmental Drivers Based on Interpretable Machine Learning Model. Fishes 2026, 11, 274. https://doi.org/10.3390/fishes11050274

AMA Style

Zhang L, Fan W, Tang F, Shi Y, Zhang S. Prediction of High-Abundance Fishing Grounds for Chub Mackerel (Scomber japonicus) in the Northwest Pacific Ocean and Its Environmental Drivers Based on Interpretable Machine Learning Model. Fishes. 2026; 11(5):274. https://doi.org/10.3390/fishes11050274

Chicago/Turabian Style

Zhang, Leilei, Wei Fan, Fenghua Tang, Yongchuang Shi, and Shengmao Zhang. 2026. "Prediction of High-Abundance Fishing Grounds for Chub Mackerel (Scomber japonicus) in the Northwest Pacific Ocean and Its Environmental Drivers Based on Interpretable Machine Learning Model" Fishes 11, no. 5: 274. https://doi.org/10.3390/fishes11050274

APA Style

Zhang, L., Fan, W., Tang, F., Shi, Y., & Zhang, S. (2026). Prediction of High-Abundance Fishing Grounds for Chub Mackerel (Scomber japonicus) in the Northwest Pacific Ocean and Its Environmental Drivers Based on Interpretable Machine Learning Model. Fishes, 11(5), 274. https://doi.org/10.3390/fishes11050274

Article Metrics

Back to TopTop