Ship Equipment Order Target Price Prediction: An Interpretable Model Based on Boruta–Lasso and CatBoost-SHAP
Abstract
1. Introduction
2. Literature Review
2.1. Current Status of Research on Ship Equipment Cost and Price Prediction
2.2. Progress in Feature Selection Methods for High-Dimensional Data
2.3. Research Status of Integrated Prediction and Interpretability Methods
2.4. Aspects of Existing Research That Need Further Improvement
3. Method
3.1. Symbol Definition
3.2. Data Preprocessing
3.2.1. Min–Max Normalization
3.2.2. Handling Missing Values and Outliers
3.2.3. Dataset Partitioning
3.3. Feature Selection
3.3.1. Boruta Feature Selection
3.3.2. Lasso Regularization Sparse Selection
3.3.3. Final Feature Set Construction
3.4. CatBoost Model Training
3.5. Hyperparameter Tuning
3.6. Final Independent Test
3.6.1. Coefficient of Determination R2
3.6.2. Root Mean Square Error (RMSE)
3.6.3. Mean Absolute Error (MAE)
3.7. SHAP Interpretability Analysis
3.7.1. SHAP Value Calculation Formula
3.7.2. Statistical Response Transition Point and Uncertainty Estimation
3.8. Overall Modeling Process and Pseudocode
| Algorithm 1: Ship Equipment Order Target Price Prediction Model |
| Input: Original feature matrix X, equipment price vector y Output: Optimal prediction model, core feature subset, SHAP statistical response transition point, and 95% confidence intervals 1. Data Preprocessing: Perform price deflation processing, outlier detection, and correction. 2. Three-layer strict dataset splitting: Following the principle of no sample overlap and complete isolation, completely preventing data leakage: 2.1 Divide the 33 original real samples into a real training set and an independent test set at an 8:2 ratio; the test set remains completely sealed and is not involved in feature selection, hyperparameter optimization, or model training, only used for final unbiased generalization performance evaluation. 2.2 Further split the real training set at an 87.5:12.5 ratio into a real sub-training set and an independent validation set; the validation set is only used for hyperparameter optimization, not for feature selection or model training. 2.3 Merge interpolated augmented samples with the real sub-training set to construct the final modeling training set. 3. Feature Standardization: Fit the min-max normalization rule only based on the final training set, then consistently transform the training set, validation set, and test set. 4. Two-stage Feature Selection: Execute Boruta–Lasso two-stage feature selection only within the final training set to identify the core feature subset for the model. 5. Hyperparameter Optimization: Conduct grid search-based hyperparameter optimization using the independent validation set. 6. Model Training: Train the CatBoost model on the final training set using the optimal hyperparameters. 7. Independent Test Evaluation: On the completely isolated independent test set, use the coefficient of determination R2, root mean square error (RMSE), and mean absolute error (MAE) to conduct an unbiased evaluation of model generalization ability. 8. Model Interpretability Analysis: Calculate feature contributions based on Tree-SHAP, apply LOWESS curve smoothing, and extract and output the SHAP statistical response transition point and 95% confidence intervals. |
4. Results and Discussion
4.1. Data Collection and Preprocessing
4.1.1. Sample Composition and Data Sources
4.1.2. Data Integrity
4.1.3. Comparable Treatment of Destroyers and Frigates
4.1.4. Price Conversion and Temporal Deflation
4.2. Feature Selection Using a Two-Stage Boruta–Lasso Model
4.2.1. Stage One: Preliminary Feature Importance Screening Based on the Boruta Algorithm
- Base model configuration: A RandomForestRegressor model is constructed with the number of trees set to n_estimators = 100 and the maximum tree depth set to max_depth = 12 to prevent overfitting, a fixed random seed random_state = 42 to ensure reproducibility, and multi-core parallel training enabled to accelerate computation.
- Algorithm parameter optimization: To avoid excessively eliminating weakly relevant but engineering-significant features, this study specifically optimizes the core parameters of the Boruta algorithm. The percentile threshold perc is reduced from the default 100 to 60, so that a real feature only needs to exceed the median importance of shadow features to be retained. The significance level alpha is increased from 0.05 to 0.2, relaxing the statistical significance threshold. The combination of both enhances the ability to identify weakly correlated features, preventing the omission of key factors with engineering value, such as displacement and radar configuration, making it more suitable for ship equipment price prediction scenarios.
- Screening results: After the algorithm was executed, important features were identified and confirmed using boruta.support_, and ultimately 16 important features were selected from the original feature set, completing the preliminary dimensionality reduction in high-dimensional features. The selected features are X1, X2, X3, X4, X8, X9, X10, X11, X12, X13, X16, X20, X21, X22, X23, and X24. These features will all enter the subsequent Lasso regression stage for further feature refinement and modeling through regularization.
4.2.2. Second Stage: Feature Sparsity Selection Based on Lasso Regularization
- Feature standardization: Apply min–max normalization to the 16 features already selected by Boruta to eliminate scale differences due to different units, mapping the features to the [0, 1] range and ensuring that subsequent regularized training is not affected by feature scale. The standardization process strictly follows machine learning protocols, using only the statistics from the training set to transform the validation and test sets, avoiding data leakage.
- Optimal regularization parameter tuning: The regularization parameter was optimized using LassoCV combined with RepeatedKFold (five-fold cross-validation, with three repetitions), setting alphas = np.logspace(−4, 0, 50) to construct a logarithmically spaced grid for a global search of regularization intensity. Using the cross-validated mean squared error (MSE) as the evaluation metric, the parameter corresponding to the minimum error was found to be = 0.0001. Meanwhile, according to the 1-SE rule, a sparser and more generalizable regularization parameter of = 0.0024 was selected. Considering the model’s robustness and feature reduction requirements in the context of ship equipment price prediction, the final optimal regularization parameter was chosen to achieve greater feature sparsity and multicollinearity suppression while ensuring that predictive accuracy meets the requirements.
- Sparse selection results: Based on the training set, a Lasso model was constructed to remove redundant features whose coefficients were shrunk to 0, ultimately retaining 11 non-zero coefficient features, including X1, X3, X10, X11, X12, X19, X20, X21, X22, X23, and X24, achieving feature dimensionality reduction and multicollinearity suppression.
4.2.3. Intersection of Two-Stage Features and Construction of Final Feature Set
4.2.4. Visualization Analysis of Feature Selection Results
4.3. Nonlinear Association Between Core Features and Target Prices of Ship Equipment Orders and Statistical Response Analysis
4.4. Model Performance
4.5. Model Explanation
4.5.1. Individual Explanation
4.5.2. Comprehensive Explanation
4.5.3. Main Feature Dependency
- Statistical response of full-load displacement (X1)
- 2.
- Statistical Response of the Number of Vertical Launch Missile System Units (X11)
- 3.
- Statistical Response of Phased Array Radar Quantity (X20)
- 4.
- Statistical Response of Equipment Combat Power (X24)
4.6. Model Performance Comparison
4.6.1. Overall Model Performance Comparison
4.6.2. Ablation Study
4.7. Policy Research
- Construct a hierarchical classification pricing standard based on core statistical association characteristics. Based on key related characteristics such as full-load displacement, number of missile vertical launch units, phased array radar configuration, and comprehensive combat effectiveness, a benchmark price system for ship equipment ordering by tonnage, configuration, and combat power level should be established to replace the traditional subjective estimation model and improve the standardization, consistency and traceability of pricing.
- Implement precise control across the whole chain of cost-related elements. Focus on high-cost modules such as hull structures, power systems, shipboard electronic equipment and weapon systems, and reduce the unit price of core equipment through large-scale procurement, modular design, and localized substitution. At the same time, implement on-demand configuration for high-value equipment such as vertical launch systems and phased array radars to avoid functional redundancy and inflated costs.
- Establish a data-driven closed loop for price demonstration and review. Build a standardized database of technical parameters, configuration schemes and purchase prices of ship equipment, and embed the Boruta–Lasso-CatBoost-SHAP integrated model into the equipment pricing demonstration process to form a closed loop of “feature screening–high-precision prediction–explainable attribution–review decision-making”, improving the scientificity and transparency of pricing decisions.
- Optimize equipment efficiency–cost allocation based on nonlinear statistical transition points. Relying on the inflection point law of statistical response revealed by SHAP dependency (displacement, vertical unit, and transition points of combat effectiveness), the equipment configuration structure is optimized under the premise of meeting combat needs, the optimal cost-effective interval is identified, and the balanced matching between combat effectiveness and procurement cost is realized.
- Improve dynamic price correction and medium- and long-term forecasting mechanisms. Integrate external factors such as industrial producer price index, supply chain fluctuations, and technology iteration to construct a dynamic price adjustment model. The forecasting framework of this paper is used to predict the cost trend of the whole-lifecycle of ship equipment, which provides stable support for equipment development planning, budgeting and fund coordination.
4.8. Discussion of Research Limitations and Shortcomings
- Limited scale of original real samples is. In this study, there were only 33 original real samples. Although data augmentation through interpolation expanded the dataset to 198 samples for modeling, the interpolated samples are only a means of data augmentation and are not considered independent empirical samples. The limited amount of original data may still pose certain risks of fluctuation in the model’s generalization ability and the estimation of statistical response transition points.
- Price data come from public sources, and some data cannot be independently verified. Ship procurement prices involve defense secrets. This study used only public quotations and contract prices, without obtaining internal data such as actual shipyard costs, labor, and materials. As a result, some price information is difficult to cross-verify.
- Historical military procurement prices have potential biases. Military procurement is influenced by non-technical factors such as international politics, arms trade policies, batch sizes, and cooperation agreements. Historical prices may contain certain non-market biases that could affect the stability of statistical associations.
- Inflation and currency exchange introduce uncertainty. The research involves multiple currencies and cross-year prices. Although exchange rate conversion and inflation adjustment were applied, statistical errors in equipment price indices and military inflation coefficients across different countries may still bring uncertainty into the price benchmarks.
- Multicollinearity in ship design variables. Equipment characteristics such as displacement, main dimensions, propulsion, and combat capability are highly correlated. Although multicollinearity has been significantly alleviated through Boruta–Lasso selection, reducing the average VIF from 18.01 to 8.68, it could not be completely eliminated, which may affect the independence of feature contribution attribution.
- Insufficient external independent validation. This study only carried out internal training–test split validation on a self-constructed dataset and did not use an independent external dataset for cross-dataset validation. The model’s extrapolation and generalization capabilities still need to be further tested.
- Limited generalizability to other types of ships. The study focuses on destroyers and frigates, excluding aircraft carriers, supply ships, submarines, landing ships, and other vessel types. The model and statistical conclusions cannot be directly generalized to the pricing of all ship equipment.
5. Conclusions
- The Boruta–Lasso two-stage feature selection is effective and reliable, which can achieve feature sparseness while retaining key information and significantly alleviates multicollinearity and redundant interference. The VIF test shows that the average VIF of the whole feature set decreases from 18.01 to 8.68, indicating that multicollinearity is significantly improved. The overall performance is better than that of the single-feature selection methods and the full-feature modeling scheme.
- The CatBoost model optimized by GridSearchCV has excellent prediction accuracy. Based on 33 real ship data, the small sample data were enhanced to 198 modeling samples by interpolation, and the interpolated samples were not regarded as independent empirical samples. The optimal single-run test results of the model were R2 = 0.8949, RMSE = 0.0554, and MAE = 0.0476, while the average results across 10 replicates were R2 = 0.8828, RMSE = 0.0586, and MAE = 0.0529. The corresponding standard deviation was 0.0000 after retaining four decimal places because the fluctuation was very small. Compared with CatBoost, XGBoost, NGBoost, random forest and other models, the accuracy, stability and generalization ability of the model in this paper are more advantageous, and the performance is outstanding in small-sample scenarios. Paired t-test results showed that the performance improvement were statistically significant (p < 0.05).
- SHAP can be interpreted and analyzed to clearly reveal the statistical correlation law of prices. Full-load displacement, number of missile vertical launch system units, number of phased array radars, and combat effectiveness of equipment are the core characteristics showing strong statistical correlations with ship order prices. Based on the LOWESS and SHAP = 0 intersection methods, the transition points and 95% confidence intervals for each feature statistical response were obtained, revealing clear nonlinear statistical response laws. The above transition points are data-driven statistical thresholds and do not represent engineering or economic breakpoints.
- Ship equipment prices show increasing marginal statistical response with respect to key characteristics. By reasonably identifying and configuring the statistical response transition points, an optimal balance between efficiency and cost can be achieved, providing a quantitative basis for equipment scheme optimization and cost control.
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Li, Y.; Hu, C.; Gao, L. Analysis of the Reform of Equipment Pricing Work under the New Situation. China Mil.-Civ. Transit. 2025, 13, 13–16. (In Chinese) [Google Scholar]
- He, D.; Sun, S.; Xie, L. Multi-target regression based on multi-layer sparse structure and its application in warships scheduled maintenance cost prediction. Appl. Sci. 2022, 13, 435. [Google Scholar] [CrossRef]
- Mun, J. Empirical cost estimation for US Navy ships. Univers. J. Manag. 2019, 7, 152–176. [Google Scholar] [CrossRef]
- Moran, D. The Benefits and Drawback of Standardization of Cost Estimation in Naval Surface Warfare Centers. Ph.D. Thesis, Acquisition Research Program, Monterey, CA, USA, 2024. [Google Scholar]
- Kaluzny, B.L.; Wang, J.; Chen, H. An application of data mining algorithms for shipbuilding cost estimation. J. Cost Anal. Parametr. 2011, 4, 2–30. [Google Scholar] [CrossRef]
- Kursa, M.B.; Rudnicki, W.R. Boruta: Wrapper Algorithm for All Relevant Feature Selection. J. Stat. Softw. 2010, 36, 1–13. [Google Scholar] [CrossRef]
- Tibshirani, R. Regression Shrinkage and Selection Via the Lasso. J. R. Stat. Soc. Ser. B Stat. Methodol. 1996, 58, 267–288. [Google Scholar] [CrossRef]
- Huang, J.; Liu, W. Comparison of Machine Learning Models for Predicting Stroke Risk in Hypertensive Patients: Lasso Regression Model, Random Forest Model, Boruta Algorithm Model, and Boruta Algorithm Combined with Lasso Regression Model. Medicine 2025, 104, e45678. [Google Scholar] [CrossRef]
- Cabral, J.; Costa, C.; Silva, A. Comparison of Feature Selection Methods—Modelling COPD Outcomes. Mathematics 2024, 12, 1345. [Google Scholar] [CrossRef]
- Ali, N.M.; Salleh, N.S.; Omar, Z. Comparison of Microarray Breast Cancer Classification Using Support Vector Machine and Logistic Regression with LASSO and Boruta Feature Selection. Indones. J. Electr. Eng. Comput. Sci. 2020, 20, 712–720. [Google Scholar] [CrossRef]
- Mustapha, S.M.F.D.S. Predictive Analysis of Students’ Learning Performance Using Data Mining Techniques: A Comparative Study of Feature Selection Methods. Appl. Syst. Innov. 2023, 6, 89. [Google Scholar] [CrossRef]
- Demir, S.; Şahin, E.K. An Investigation of Feature Selection Methods for Soil Liquefaction Prediction Based on Tree-Based Ensemble Algorithms Using AdaBoost, Gradient Boosting, and XGBoost. Neural Comput. Appl. 2022, 35, 3173–3190. [Google Scholar] [CrossRef]
- Prokhorenkova, L.; Gusev, G.; Vorobev, A.; Dorogush, A.V.; Gulin, A. CatBoost: Unbiased Boosting with Categorical Features. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, Montreal, QC, Canada, 3–8 December 2018; pp. 6638–6648. [Google Scholar]
- Zhang, L.; Jánošík, D. Enhanced short-term load forecasting with hybrid machine learning models: CatBoost and XGBoost approaches. Expert Syst. Appl. 2024, 241, 122686. [Google Scholar] [CrossRef]
- Fan, Z.; Jin, G.; Weng, S. Complementary CatBoost based on residual error for student performance prediction. Pattern Recognit. 2025, 161, 111265. [Google Scholar] [CrossRef]
- Mosca, E.; Cagnina, L.; Errecalde, M. SHAP-based explanation methods: A review for NLP interpretability. In Proceedings of the 29th International Conference on Computational Linguistics, Gyeongju, Republic of Korea, 12–17 October 2022; pp. 5678–5690. [Google Scholar]
- Van den Broeck, G.; Lykov, A.; Schleich, M.; Suciu, D. On the tractability of SHAP explanations. J. Artif. Intell. Res. 2022, 74, 851–886. [Google Scholar] [CrossRef]
- Li, Z. Extracting spatial effects from machine learning model using local interpretation method: An example of SHAP and XGBoost. Comput. Environ. Urban Syst. 2022, 96, 101845. [Google Scholar] [CrossRef]
- Wang, H.; Zhang, Y.; Li, X. Feature selection strategies: A comparative analysis of SHAP-value and importance-based methods. J. Big Data 2024, 11, 44. [Google Scholar] [CrossRef]
- Wang, Y.; Wei, R.; Sun, S. GM Estimation Study of Ship Equipment Usage and Maintenance Costs. China Shiprep. 2006, 19, 46–48. (In Chinese) [Google Scholar]
- Huang, W.; Liao, H. Application of the Improved GM(1,1) Model in Predicting Special Repair Costs. J. Wuhan. Univ. Technol. 2007, 29, 100–102. (In Chinese) [Google Scholar]
- Liu, M. Application of the Improved GM(1,1) Model in Predicting Ship Maintenance Costs. Ship Electron. Eng. 2010, 30, 151–154. (In Chinese) [Google Scholar]
- Li, Z. Research on Prediction of Ship Equipment Maintenance Support Costs Based on Grey Theory. Knowl. Econ. 2012, 24, 94–106. (In Chinese) [Google Scholar]
- Zhang, J.; Geng, J.; Sun, L. Prediction and Analysis of Ship Hull Construction Costs Based on GM(1,n) Model. In Proceedings of the Life-Cycle Cost Technology and Coordinated Development, Wuhan, China, 24–26 August 2010; pp. 204–208. (In Chinese) [Google Scholar]
- Wu, B. Prediction of Ship Equipment Maintenance Costs Based on Combined Optimization Theory. Ship Sci. Technol. 2018, 40, 31–33. (In Chinese) [Google Scholar]
- He, P.; Sun, S. Ship Equipment Maintenance Cost Prediction Model Based on Improved GM(0,N). Ship Electron. Eng. 2022, 42, 151–154. (In Chinese) [Google Scholar]
- Liu, X.; Li, J. Prediction and Analysis of Ship Construction Costs Based on Neural Network. J. Nav. Univ. Eng. 2001, 13, 96–98. (In Chinese) [Google Scholar]
- Lin, M.; Zhang, S. Research on Shipbuilding Cost Prediction Based on Support Vector Machine. In Proceedings of the Life Cycle Cost Technology and Circular Economy, Wuhan, China, 01 November 2008; pp. 52–56. (In Chinese) [Google Scholar]
- Wang, W.; Peng, R. Research on the Application of Combined Prediction of Batch Shipbuilding Costs. Ship Sci. Technol. 2009, 31, 124–127. (In Chinese) [Google Scholar]
- Jiang, T.; Zhang, H. Research on Combined Prediction of Ship Equipment Construction Costs Based on Information Entropy. Ship Sci. Technol. 2011, 33, 127–130. (In Chinese) [Google Scholar]
- Dong, P.; Leng, J.; Luo, Z. Research on Combined Prediction Method of Shipbuilding Costs Based on Support Vector Machine. Ship Sci. Technol. 2011, 299, 13–15. (In Chinese) [Google Scholar]
- Yang, J.; Su, X.; Wu, F.; Zhang, Y. Shipbuilding Cost Estimation Based on XGBoost-SHAP Model. Ship Eng. 2026, 48, 129–139+177. [Google Scholar]
- Sun, S.; He, D.; Li, J. Prediction of Ship Class Repair Costs Based on Soft-impute Technology and MLSR Algorithm. J. Nav. Univ. Eng. 2023, 35, 40–45. (In Chinese) [Google Scholar]
- Lin, M.; Wang, C.; Xie, L. Case-Based Reasoning Prediction of Ship Equipment Maintenance Costs Based on Dual Similarity Retrieval. J. Nav. Univ. Eng. 2022, 34, 68–73. (In Chinese) [Google Scholar]
- Zhang, J.; Zhang, Q. Estimation of Ship Maintenance Costs Based on BP Neural Network. Ship Sci. Technol. 2019, 41, 220–222. (In Chinese) [Google Scholar]
- Wu, P.-C.; Lin, C.-Y. Feasibility and Cost-Benefit Analysis of Methanol as a Sustainable Alternative Fuel for Ships. J. Mar. Sci. Eng. 2025, 13, 973. [Google Scholar] [CrossRef]
- Palmén, M.; Lotrič, A.; Laakso, A.; Bolbot, V.; Elg, M.; Valdez Banda, O.A. Selecting Appropriate Energy Source Options for an Arctic Research Ship. J. Mar. Sci. Eng. 2023, 11, 2337. [Google Scholar] [CrossRef]
- Zheng, Q.; Sun, L.; Chang, S.; Xing, H. Techno-Economic Analysis of Multi-Purpose Heavy-Lift Vessels Using Methanol as Fuel. J. Mar. Sci. Eng. 2025, 13, 1234. [Google Scholar] [CrossRef]
- Alblas, G.; Pruijn, J. Are current shipbuilding cost estimation methods ready for a sustainable future? A literature review of cost estimation methods and challenges. Int. J. Ship Prod. Des. 2024, 71, 3–28. [Google Scholar] [CrossRef]
- Jin, Y.; Zhang, C. Development of Costing and Budget Control Strategy for Shipbuilding Based on Machine Learning. J. Comb. Math. Comb. Comput. 2025, 127, 7063–7082. [Google Scholar] [CrossRef]
- Jeong, B.; Wang, H.; Oguz, E.; Zhou, P. An effective framework for life cycle and cost assessment for marine vessels aiming to select optimal propulsion systems. J. Clean. Prod. 2018, 187, 111–130. [Google Scholar] [CrossRef]
- Petersen, B.; Efatmaneshnik, M. Optimising Warship Lifecycle Value: A Real Options Approach. In Proceedings of the 2024 IEEE International Symposium on Systems Engineering (ISSE), Rome, Italy, 16–19 October 2024; IEEE; 2024, pp. 1–8. [Google Scholar]
- Xiao, R.; Lin, M.; Tan, X. Single Ship Life Cycle Maintenance Cost Prediction Based on Variant Error Correction Model. In Proceedings of the 5th International Conference on Marine Systems Engineering and Computing, Dalian, China, 6–8 June 2025; pp. 239–245. [Google Scholar]
- Wang, Y.W.; Li, X.; Zhang, H. An XGBoost-SHAP Model for Energy Demand Prediction With Boruta–Lasso Feature Selection. IEEE Access 2025, 13, 135806–135821. [Google Scholar] [CrossRef]
- Ileri, K. Comparative analysis of CatBoost, LightGBM, XGBoost, RF, and DT methods optimised with PSO to estimate the number of k-barriers for intrusion detection in wireless sensor networks. Int. J. Mach. Learn. Cybern. 2025, 16, 6937–6956. [Google Scholar] [CrossRef]
- Qiu, J.; Li, Y.; Chen, L. Application of a Multi-Algorithm-Optimized CatBoost Model in Predicting the Strength of Multi-Source Solid Waste Backfilling Materials. Big Data Cogn. Comput. 2025, 9, 203. [Google Scholar] [CrossRef]
- Mouad, B.H.I.H.; El, M.A.; Ouahbi, F. Temporal Pattern-Aware Temperature Forecasting Using CatBoost: A Hybrid Machine Learning Approach. Results Eng. 2026, 27, 110212. [Google Scholar] [CrossRef]
- Hancock, J.T.; Khoshgoftaar, T.M.; Liang, Q. A problem-agnostic approach to feature selection and analysis using SHAP. J. Big Data 2025, 12, 12. [Google Scholar] [CrossRef]
- Kruk, M. SHAP-NET, a network based on Shapley values as a new tool to improve the explainability of the XGBoost-SHAP model for the problem of water quality. Environ. Model. Softw. 2025, 188, 106403. [Google Scholar] [CrossRef]
- Qi, X.; Wang, Y.; Zhang, L. Machine learning and SHAP value interpretation for predicting comorbidity of cardiovascular disease and cancer with dietary antioxidants. Redox Biol. 2025, 79, 103470. [Google Scholar] [CrossRef] [PubMed]
- Bernal, L.; Rastelli, G.; Pinzi, L. Improving machine learning classification predictions through SHAP and features analysis interpretation. J. Chem. Inf. Model. 2025, 65, 11716–11732. [Google Scholar] [CrossRef]
- Salih, A.M.; Rashid, T.A.; Ali, S.S. A perspective on explainable artificial intelligence methods: SHAP and LIME. Adv. Intell. Syst. 2025, 7, 240030. [Google Scholar] [CrossRef]
- Li, Y.; Lin, J.; Wang, X. Research on the Formation Mechanism and Demonstration Method of Equipment Ordering Target Price. Aviat. Weapons 2024, 31, 133–138. (In Chinese) [Google Scholar]
- Jane’s Information Group. Jane’s Fighting Ships 2020–2021; Jane’s Publishing: London, UK, 2020. [Google Scholar]
- Li, Y. Guide to Appreciating Main Combat Ships (Collector’s Edition), 2nd ed.; Tsinghua University Press: Beijing, China, 2018. (In Chinese) [Google Scholar]
- Leng, J.; Qi, H.; Dong, P. Calculation Method of Engineering Value Ratio for the Construction Cost of Large Ship Platforms. Syst. Eng. Electron. 2010, 32, 557–561+565. (In Chinese) [Google Scholar]
- Feng, W.; Kuang, H.; Wu, H.; Zhang, Y. Analysis of Factors Affecting New Ship Price Fluctuations Based on Coupling. Syst. Eng. Theory Pract. 2016, 36, 2879–2888. (In Chinese) [Google Scholar]













| Type | Factor | Unit |
|---|---|---|
| Military Requirements | X5 Maximum Speed | knots |
| X6 Endurance (Maximum Range) | nautical miles | |
| X10 Self-sufficiency | days | |
| X11 Number of Vertical Launch Missile System Units | units | |
| X12 Number of Vertical Launch Missile Systems | sets | |
| X13 Number of Anti-Ship Missile Launchers | sets | |
| X14 Number of Air Defense Missile Launchers | sets | |
| X15 Number of Anti-Submarine Missile Launchers | sets | |
| X16 Number of Torpedo Launchers | sets | |
| X17 Number of Anti-Submarine Depth Charge Launchers | sets | |
| X18 Anti-Submarine Rocket Launchers | sets | |
| X19 Number of Jammer Rocket Launchers | sets | |
| X20 Number of Phased Array Radars | sets | |
| X21 Number of Other Radars | sets | |
| X22 Number of Sonar Systems | sets | |
| X23 Number of Naval Guns | sets | |
| X24 Equipped Combat Strength | combat capability | |
| Technical Solution | X1 Full-Load Displacement | tons |
| X2 Ship Length | m | |
| X3 Ship Width | m | |
| X4 Draft | m | |
| X7 Crew Size | persons | |
| Industrial Manufacturing Level | X8 Total Power of Ship Propulsion System | kW |
| X9 Installed Generator Power | kW | |
| Market Economic Environment | Reflected by the Price Index and Exchange Rate Index, Internal Characteristic Variables are not Included. |
| Algorithm | Parameter | Parameter Grid | Optimal Value |
|---|---|---|---|
| CatBoost | learning_rate | [0.005, 0.01, 0.02, 0.05, 0.1] | 0.1 |
| iterations | [100, 200, 500, 1000] | 500 | |
| depth | [1, 2, 9, 10] | 2 |
| Evaluation | Optimal Single Result | Results of 10 Repeated Validations (Mean ± Standard Deviation) |
|---|---|---|
| R2 | 0.8949 | 0.8828 ± 0.0000 |
| RMSE | 0.0554 | 0.0586 ± 0.0000 |
| MAE | 0.0476 | 0.0529 ± 0.0000 |
| Model | Optimal Single Result | Results of 10 Repeated Validations (Mean ± Standard Deviation) | p-Value | Significance | ||||
|---|---|---|---|---|---|---|---|---|
| R2 | RMSE | MAE | R2 | RMSE | MAE | |||
| Grid-CatBoost | 0.8949 | 0.0554 | 0.0476 | 0.8828 ± 0.0000 | 0.0586 ± 0.0000 | 0.0529 ± 0.0000 | —— | —— |
| CatBoost | 0.8627 | 0.0634 | 0.0559 | 0.8459 ± 0.0000 | 0.0671 ± 0.0000 | 0.0617 ± 0.0000 | 0.021 | p < 0.05 |
| NGBoost | 0.8044 | 0.0756 | 0.0553 | 0.8045 ± 0.0002 | 0.0756 ± 0.0000 | 0.0553 ± 0.0003 | 0.0004 | p < 0.001 |
| XGBoost | 0.8353 | 0.0694 | 0.0530 | 0.8215 ± 0.0105 | 0.0722 ± 0.0022 | 0.0549 ± 0.0022 | 0.003 | p < 0.01 |
| RF | 0.8511 | 0.0660 | 0.0468 | 0.8457 ± 0.0091 | 0.0671 ± 0.0020 | 0.0486 ± 0.0014 | 0.008 | p < 0.01 |
| Feature Selection | R2 | RMSE | MAE | Number of Features | Average VIF | Highest VIF | p-Value | Significance |
|---|---|---|---|---|---|---|---|---|
| Boruta–Lasso | 0.8949 | 0.0554 | 0.0476 | 10 | 8.68 | 32.03 | —— | —— |
| Boruta | 0.8333 | 0.0698 | 0.0520 | 16 | 11.05 | 69.21 | 0.0012 | p < 0.01 |
| Lasso | 0.8744 | 0.0606 | 0.0481 | 11 | 8.27 | 33.45 | 0.0178 | p < 0.05 |
| Full features | 0.8612 | 0.0637 | 0.0486 | 24 | 18.01 | 119.95 | 0.0068 | p < 0.01 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Li, K.; Sun, S.; Zhu, C.; Zhang, Y. Ship Equipment Order Target Price Prediction: An Interpretable Model Based on Boruta–Lasso and CatBoost-SHAP. J. Mar. Sci. Eng. 2026, 14, 949. https://doi.org/10.3390/jmse14100949
Li K, Sun S, Zhu C, Zhang Y. Ship Equipment Order Target Price Prediction: An Interpretable Model Based on Boruta–Lasso and CatBoost-SHAP. Journal of Marine Science and Engineering. 2026; 14(10):949. https://doi.org/10.3390/jmse14100949
Chicago/Turabian StyleLi, Kai, Shengxiang Sun, Chen Zhu, and Ying Zhang. 2026. "Ship Equipment Order Target Price Prediction: An Interpretable Model Based on Boruta–Lasso and CatBoost-SHAP" Journal of Marine Science and Engineering 14, no. 10: 949. https://doi.org/10.3390/jmse14100949
APA StyleLi, K., Sun, S., Zhu, C., & Zhang, Y. (2026). Ship Equipment Order Target Price Prediction: An Interpretable Model Based on Boruta–Lasso and CatBoost-SHAP. Journal of Marine Science and Engineering, 14(10), 949. https://doi.org/10.3390/jmse14100949

