A Machine Learning-Based Predictive Model for Maintenance Management of Combustion Engines in the Agricultural Sector
Abstract
1. Introduction
2. Related Work
2.1. Evolution of Maintenance Strategies and the Role of Data-Driven Methods
2.2. Machine Learning Applied to Internal Combustion Engine Diagnostics
2.3. Gradient Boosting Methods in Industrial Maintenance
2.4. Predictive Maintenance in Agricultural Contexts and Identified Gaps
3. Materials and Methods
3.1. Problem Formulation
- Dataset and observation space. Let denote a fleet of stationary combustion engines. For each engine , a sequence of operational records is available, ordered chronologically. The full dataset is:
- Feature transformation. The raw transactional record is transformed via a sliding-window function :
- Prediction horizon and label definition. The target variable is defined over a forward-looking window days. A record is labelled positive if an unscheduled corrective maintenance intervention occurs within the next days:
- Optimization objective. Given the asymmetric cost structure of predictive maintenance, the primary criterion is:
- Temporal data partition. Let denote the split point (October 2022):
3.2. Research Design
3.3. Dataset and Preprocessing
3.4. Feature Engineering
3.5. Exclusion of Label-Associated Event Descriptors
3.6. Definition of the Prediction Horizon
- Logistical-Operational Criterion: In agricultural contexts with large fleets, the supply chain for critical spare parts requires extended planning times. A shorter window would be insufficient for supply management, whereas the 60-day horizon enables tactical planning of downtime without affecting production peaks.
- Feature-Horizon Alignment: The 60-day horizon matches the time scale of the engineered predictors, which encode cumulative load and failure-rate trajectories over multi-week windows. The study does not define or evaluate a taxonomy of failure types; the contribution of the trajectory-based features to predictive performance is quantified directly in the ablation study (Section 4.4).
3.7. Usage Profile Segmentation (Clustering)
3.8. Model Configuration and Algorithms
3.9. Training and Validation Strategy
- Scenario A: Classic 80/20 split (training/testing) with 3-fold cross-validation ().
- Scenario B: Strict 60/40 chronological (temporal) split with 5-fold cross-validation (). To rigorously evaluate the predictive horizon, the split was performed chronologically: historical records from December 2018 up to October 2022 formed the training set (60%), while the subsequent records from November 2022 to June 2025 constituted the test set (40%). This strict temporal ordering ensures that no future information is available during training, effectively preventing data leakage and simulating a real-world deployment scenario.
3.10. Evaluation Metrics and Class Balancing Techniques
Statistical Significance Testing
4. Results and Discussion
4.1. Characterization of Operational Patterns
4.2. Comparison of Evaluation Protocols
4.3. Predictive Performance Evaluation
4.4. Component Validation Under Leave-Engine-Group-Out Cross-Validation
4.5. Model Interpretability
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AUC | Area Under the Curve |
| CRISP-DM | Cross-Industry Standard Process for Data Mining |
| FN | False Negatives |
| FP | False Positives |
| LightGBM | Light Gradient Boosting Machine |
| LOGO | Leave-Engine-Group-Out (engine-grouped cross-validation) |
| MPC | Model-Based Predictive Control |
| OOF | Out-Of-Fold (pooled cross-validated predictions) |
| PdM | Predictive Maintenance |
| RF | Random Forest |
| ROC | Receiver Operating Characteristic |
| TN | True Negatives |
| TP | True Positives |
Appendix A. Feature Engineering Dictionary
| Variable | Description | Unit |
|---|---|---|
| Group 1: Raw operational inputs | ||
| potencia_hp | Rated engine power (manufacturer specification) | HP |
| horas_motor_actuales | Current hour-meter reading at the time of intervention | Hours |
| tipo_mant_actual | Type of current intervention (0 = Preventive, 1 = Corrective) (excluded from the final feature space, Section 3.5) | Cat. |
| actividad_actual_enc | Specific maintenance activity, target-encoded by failure rate (excluded from the final feature space, Section 3.5) | Ratio |
| Group 2: Time-based features | ||
| dias_desde_ultimo_mant | Days elapsed since the last service intervention | Days |
| horas_desde_ultimo_mant | Engine hours accumulated since the last service | Hours |
| antiguedad_motor_en_dias | Total operational age of the asset since commissioning | Days |
| dias_desde_primer_mant | Days elapsed since the first recorded maintenance event | Days |
| año_mantenimiento | Calendar year in which the intervention was recorded | Year |
| trimestre_mantenimiento | Quarter of the year in which the intervention occurred | Cat. |
| mes_mantenimiento | Month in which the intervention was recorded | Cat. |
| Group 3: Cumulative maintenance history | ||
| num_mantenimientos_previos | Total number of prior interventions recorded for the engine | Count |
| num_correctivos_previos | Total number of prior corrective interventions | Count |
| conteo_mantenimientos_acum | Accumulated maintenance count up to the current record | Count |
| correctivos_acumulados | Accumulated corrective interventions over engine lifetime | Count |
| horas_totales_acumuladas | Total accumulated operating hours over engine lifetime | Hours |
| ratio_fallas | Historical failure frequency (failures/total interventions) | Ratio |
| ratio_fallas_acumuladas | Cumulative failure rate up to the current record | Ratio |
| horas_por_falla_correctiva | Mean operating hours elapsed between corrective failures | Hrs/Fail |
| Group 4: Sliding-window features | ||
| media_movil_horas_3 | 3-event rolling mean of engine hours at intervention | Hours |
| std_dev_dias_3 | 3-event rolling standard deviation of inter-service intervals | Days |
| fallas_ultimos_3_mant | Number of corrective events in the last 3 interventions | Count |
| Group 5: Station-level contextual features | ||
| fallas_historicas_estacion | Historical corrective interventions at the pumping station prior to record | Count |
| total_fallas_estacion_hist | Total cumulative corrective interventions recorded at the station | Count |
| promedio_horas_estacion | Mean operating hours across all engines at the station | Hours |
| Group 6: Outlier indicator flags | ||
| outlier_horas_motor_act. | Binary flag indicating anomalous hour-meter reading (IQR method) | Binary |
| outlier_dias_ultimo_mant | Binary flag indicating anomalous inter-service day interval | Binary |
| outlier_horas_ultimo_mant | Binary flag indicating anomalous hours since last service | Binary |
| Group 7: Brand dummy variables (one-hot encoding) | ||
| marca_CATERPILLAR | Binary indicator: engine manufactured by Caterpillar | Binary |
| marca_CUMMINS | Binary indicator: engine manufactured by Cummins | Binary |
| marca_DETROIT | Binary indicator: engine manufactured by Detroit | Binary |
| marca_MARATHON | Binary indicator: engine manufactured by Marathon | Binary |
| marca_MAXX-FORCE | Binary indicator: engine manufactured by Maxx-Force | Binary |
| marca_MWM | Binary indicator: engine manufactured by MWM | Binary |
| marca_SIEMENS | Binary indicator: engine manufactured by Siemens | Binary |
| marca_US_ELEC_MOTOR | Binary indicator: engine manufactured by US Electrical Motor | Binary |
| marca_WEICHAI | Binary indicator: engine manufactured by Weichai | Binary |
| Group 8: Categorical encodings | ||
| modelo_target_enc | Engine model encoded by observed failure rate | Ratio |
| estacion_hacienda_enc | Hacienda identifier encoded by observed failure rate | Ratio |
| estacion_tipo_tipo | Station type label-encoded (irrigation/drainage) | Cat. |
| tipo_mant_actual_mant | Maintenance type label-encoded (0 = Preventive, 1 = Corrective) (excluded from the final feature space, Section 3.5) | Cat. |
References
- Abbasi, R.; Martinez, P.; Ahmad, R. The digitization of agricultural industry—A systematic literature review on agriculture 4.0. Smart Agric. Technol. 2022, 2, 100042. [Google Scholar] [CrossRef] [Scilit]
- Zonta, T.; da Costa, C.A.; da Rosa Righi, R.; de Lima, M.J.; da Trindade, E.S.; Li, G.P. Predictive maintenance in the Industry 4.0: A systematic literature review. Comput. Ind. Eng. 2020, 150, 106889. [Google Scholar] [CrossRef] [Scilit]
- Selcuk, S. Predictive maintenance, its implementation and latest trends. Proc. Inst. Mech. Eng. Part B J. Eng. Manuf. 2017, 231, 1670–1679. [Google Scholar] [CrossRef] [Scilit]
- Carvalho, T.P.; Soares, F.A.A.M.N.; Vita, R.; da P. Francisco, R.; Basto, J.P.; Alcalá, S.G.S. A systematic literature review of machine learning methods applied to predictive maintenance. Comput. Ind. Eng. 2019, 137, 106024. [Google Scholar] [CrossRef] [Scilit]
- Amruthnath, N.; Gupta, T. A research study on unsupervised machine learning algorithms for early fault detection in predictive maintenance. In Proceedings of the 2018 5th International Conference on Industrial Engineering and Applications, ICIEA 2018, Singapore, 26–28 April 2018; pp. 355–361. [Google Scholar] [CrossRef] [Scilit]
- Roman, A.; Rahman, M.M.; Haider, S.A.; Akram, T.; Naqvi, S.R. Integrating Feature Selection and Deep Learning: A Hybrid Approach for Smart Agriculture Applications. Algorithms 2025, 18, 222. [Google Scholar] [CrossRef] [Scilit]
- Bhatt, A.N.; Shrivastava, N. Application of Artificial Neural Network for Internal Combustion Engines: A State of the Art Review. Arch. Comput. Methods Eng. 2022, 29, 897–919. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Mohasab, H.; Abouelsoud, M.; Shmroukh, A.N.; Ghazaly, N. Application of Artificial Neural Networks in Predicting Internal Combustion Engine Performance and Emission Characteristics: A Review of Key Methodologies and Findings. Int. J. Robot. Control Syst. 2024, 4, 1967–2025. [Google Scholar] [CrossRef] [Scilit]
- Giordano, D.; Giobergia, F.; Pastor, E.; Macchia, A.L.; Cerquitelli, T.; Baralis, E.; Mellia, M.; Tricarico, D. Data-driven strategies for predictive maintenance: Lesson learned from an automotive use case. Comput. Ind. 2022, 134, 103554. [Google Scholar] [CrossRef] [Scilit]
- Niroomand, N.; Bach, C. Integrating Machine Learning for Predicting Internal Combustion Engine Performance and Segment-Based CO2 Emissions Across Urban and Rural Settings. IEEE Access 2024, 12, 66223–66236. [Google Scholar] [CrossRef] [Scilit]
- Jaramillo, I.F.; Villarroel-Molina, R.; Pico, B.R.; Redchuk, A. A Comparative Study of Classifier Algorithms for Recommendation of Banking Products. In Proceedings of the Trends and Applications in Information Systems and Technologies; Springer: Cham, Switzerland, 2021; Volume 1366, pp. 253–263. [Google Scholar] [CrossRef] [Scilit]
- Aliramezani, M.; Koch, C.R.; Shahbakhti, M. Modeling, diagnostics, optimization, and control of internal combustion engines via modern machine learning techniques: A review and future directions. Prog. Energy Combust. Sci. 2022, 88, 100967. [Google Scholar] [CrossRef] [Scilit]
- Breiman, L. Random Forests. Mach. Learn. 2001, 45, 5–32. [Google Scholar] [CrossRef] [Scilit]
- Snyder, T.; Ali, A. Predictive Analytics and Diagnostics Drive Effectiveness in Condition Based Monitoring. In Proceedings of the ASME 2010 Internal Combustion Engine Division Fall Technical Conference, San Antonio, TX, USA, 12–15 September 2010; pp. 975–981. [Google Scholar] [CrossRef] [Scilit]
- Torres, N.N.S.; Lima, J.G.; Maciel, J.N.; Gazziro, M.; Filho, A.C.L.; Souto, C.R.; Salvadori, F.; Junior, O.H.A. Non-Invasive Techniques for Monitoring and Fault Detection in Internal Combustion Engines: A Systematic Review. Energies 2024, 17, 6164. [Google Scholar] [CrossRef] [Scilit]
- Ke, G.; Meng, Q.; Finley, T.; Wang, T.; Chen, W.; Ma, W.; Ye, Q.; Liu, T.Y. LightGBM: A Highly Efficient Gradient Boosting Decision Tree. In Proceedings of the Advances in Neural Information Processing Systems; Guyon, I., Luxburg, U.V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., Garnett, R., Eds.; Curran Associates, Inc.: Red Hook, NY, USA, 2017; Volume 30. [Google Scholar]
- Schröer, C.; Kruse, F.; Gómez, J.M. A Systematic Literature Review on Applying CRISP-DM Process Model. Procedia Comput. Sci. 2021, 181, 526–534. [Google Scholar] [CrossRef] [Scilit]
- DeLong, E.R.; DeLong, D.M.; Clarke-Pearson, D.L. Comparing the Areas under Two or More Correlated Receiver Operating Characteristic Curves: A Nonparametric Approach. Biometrics 1988, 44, 837–845. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Demšar, J. Statistical Comparisons of Classifiers over Multiple Data Sets. J. Mach. Learn. Res. 2006, 7, 1–30. [Google Scholar]
- Maione, F.; Lino, P.; Maione, G.; Giannino, G. A Machine Learning Framework for Condition-Based Maintenance of Marine Diesel Engines: A Case Study. Algorithms 2024, 17, 411. [Google Scholar] [CrossRef] [Scilit]
- Jain, A.K. Data clustering: 50 years beyond K-means. Pattern Recognit. Lett. 2010, 31, 651–666. [Google Scholar] [CrossRef] [Scilit]
- Pedregosa, F.; Varoquaux, G.; Gramfort, A.; Michel, V.; Thirion, B.; Grisel, O.; Blondel, M.; Müller, A.; Nothman, J.; Louppe, G.; et al. Scikit-learn: Machine Learning in Python. J. Mach. Learn. Res. 2011, 12, 2825–2830. [Google Scholar]
- Guanin-Fajardo, J.H.; Guaña-Moya, J.; Casillas, J. Predicting Academic Success of College Students Using Machine Learning Techniques. Data 2024, 9, 60. [Google Scholar] [CrossRef] [Scilit]
- Efron, B.; Tibshirani, R.J. An Introduction to the Bootstrap; Chapman & Hall/CRC: New York, NY, USA, 1993. [Google Scholar] [CrossRef] [Scilit]
- Jaramillo, I.F.; Torres-Lindao, V.; Cevallos, J. Bibliometric Analysis of Scientific Production on Multicriteria Decision-Making (MCDM) Methods. In Proceedings of the International Conference on Applied Technologies; Botto-Tobar, M., Lema Moreta, L., Zambrano Vizuete, M., León, M., Torres-Carrion, S., Durakovic, P., Eds.; Springer: Cham, Switzerland, 2026; pp. 222–235. [Google Scholar] [CrossRef] [Scilit]







| Hyperparameter | Model | Search Space | Selected Value |
|---|---|---|---|
| n_estimators | Random Forest | {50, 100, 300} | 300 |
| max_depth | Random Forest | {3, 5, None} | None |
| min_samples_split | Random Forest | {2, 5} | 2 |
| min_samples_leaf | Random Forest | {1, 2} | 1 |
| n_estimators | LightGBM | {50, 100, 500} | 500 |
| learning_rate | LightGBM | {0.05, 0.1} | 0.05 |
| num_leaves | LightGBM | {15, 31, 63} | 63 |
| min_child_samples | LightGBM | {20, 30} | 20 |
| scale_pos_weight | LightGBM | {0.8762, 1.051} | 1.051 |
| n_estimators | XGBoost | {50, 100, 300} | 300 |
| learning_rate | XGBoost | {0.05, 0.1} | 0.05 |
| max_depth | XGBoost | {3, 5, 6} | 6 |
| min_child_weight | XGBoost | {1, 5} | 1 |
| scale_pos_weight | XGBoost | {0.8762, 1.051} | 0.8762 |
| iterations | CatBoost | {50, 100, 300} | 300 |
| learning_rate | CatBoost | {0.05, 0.1} | 0.05 |
| depth | CatBoost | {3, 5, 6} | 6 |
| l2_leaf_reg | CatBoost | {1, 3} | 1 |
| scale_pos_weight | LightGBM/XGBoost | (computed) | 1.051/0.8762 |
| class_weight | Random Forest | ‘balanced’ | ‘balanced’ |
| auto_class_weights | CatBoost | ‘Balanced’ | ‘Balanced’ |
| Model | AUC | Recall | F1-Score | Precision |
|---|---|---|---|---|
| Random Forest (proposed) | 0.90 | 84.2% | 84.5% | 84.8% |
| LightGBM | 0.89 | 77.6% | 82.5% | 88.2% |
| XGBoost | 0.89 | 75.9% | 81.7% | 88.5% |
| CatBoost | 0.89 | 68.9% | 78.0% | 89.7% |
| Model | AUC | AUC/Fold () | Recall | F1 | p (AUC) | p (F1) |
|---|---|---|---|---|---|---|
| Random Forest (proposed) | 0.956 | 90.7% | 89.6% | — | — | |
| CatBoost | 0.953 | 88.8% | 89.5% | 0.042 | 0.826 | |
| XGBoost | 0.952 | 89.2% | 89.2% | 0.038 | 0.644 | |
| LightGBM | 0.951 | 88.0% | 88.5% | 0.011 | 0.120 |
| Configuration | AUC | Recall | F1-Score | Precision | p vs. M4 |
|---|---|---|---|---|---|
| M1: Raw features, no clustering | 0.900 | 86.1% | 84.0% | 82.0% | <10−26 |
| M2: FE features, no clustering | 0.956 | 90.2% | 89.2% | 88.3% | 0.511 |
| M3: Raw features + clustering | 0.900 | 85.8% | 83.7% | 81.7% | <10−26 |
| M4: FE features + clustering (full) | 0.956 | 90.7% | 89.6% | 88.5% | — |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Jaramillo, I.-F.; Orozco-Iguasnia, W.; Quinteros, R.P.A.; Villarroel-Molina, R.R.; Vilcacundo-Chiluisa, A. A Machine Learning-Based Predictive Model for Maintenance Management of Combustion Engines in the Agricultural Sector. Algorithms 2026, 19, 618. https://doi.org/10.3390/a19080618
Jaramillo I-F, Orozco-Iguasnia W, Quinteros RPA, Villarroel-Molina RR, Vilcacundo-Chiluisa A. A Machine Learning-Based Predictive Model for Maintenance Management of Combustion Engines in the Agricultural Sector. Algorithms. 2026; 19(8):618. https://doi.org/10.3390/a19080618
Chicago/Turabian StyleJaramillo, Ivan-Fredy, Walter Orozco-Iguasnia, Rubén Patricio Alcocer Quinteros, Ricardo Rafael Villarroel-Molina, and Alejandro Vilcacundo-Chiluisa. 2026. "A Machine Learning-Based Predictive Model for Maintenance Management of Combustion Engines in the Agricultural Sector" Algorithms 19, no. 8: 618. https://doi.org/10.3390/a19080618
APA StyleJaramillo, I.-F., Orozco-Iguasnia, W., Quinteros, R. P. A., Villarroel-Molina, R. R., & Vilcacundo-Chiluisa, A. (2026). A Machine Learning-Based Predictive Model for Maintenance Management of Combustion Engines in the Agricultural Sector. Algorithms, 19(8), 618. https://doi.org/10.3390/a19080618

