Symptom-Based Lung Cancer Prediction Using Ensemble Learning with Threshold Optimization and Interpretability
Abstract
1. Introduction
- •
- Creation of an effective and scalable machine learning pipeline to model lung cancer based on survey data.
- •
- State-of-the-art comparison between ensemble learning models and principled validation-based model selection.
- •
- The decision threshold is optimally determined to maximize classification performance on unseen data.
- •
- Complete analysis, such as calibration, subgroup analysis, feature ablation, and explainability.
2. Related Work
2.1. Traditional Statistical Approaches
2.2. Machine Learning Methods for Lung Cancer Prediction
2.3. Class Imbalance, Threshold Selection, and Interpretability
3. Materials and Methods
3.1. Dataset Description
3.2. Data Preprocessing
3.3. Model Development
3.4. Decision Threshold Optimization
3.5. Evaluation Metrics
4. Results
4.1. Model Performance on Validation and Test Sets
4.2. Performance Impact of Decision Threshold Optimization
4.3. Final Test Set Performance
4.4. Error Analysis
4.5. Model Explainability and Feature Importance
4.6. Calibration and Reliability
4.7. Threshold Sensitivity
4.8. Subgroup Performance
4.9. Risk Stratification Analysis
5. Discussion
6. Conclusions and Future Work
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Siegel, R.L.; Miller, K.D.; Jemal, A. Cancer statistics, 2018. CA A Cancer J. Clin. 2018, 68, 7–30. [Google Scholar] [CrossRef] [Scilit]
- Wilson, B.E.; Wright, K.; Sengar, M.; Sullivan, R.; Pearson, S.-A.; Barton, M.B.; Gyawali, B.; De Vries, E.; Moja, L.; Pramesh, C.S.; et al. Analysis of 2023 World Health Organization cancer Essential Medicines List and concordance with resource-stratified guidelines. JNCI J. Natl. Cancer Inst. 2025, 117, djaf100. [Google Scholar] [CrossRef] [Scilit]
- N.L.S.T.R. Team. Reduced lung-cancer mortality with low-dose computed tomographic screening. N. Engl. J. Med. 2011, 365, 395–409. [Google Scholar] [CrossRef] [Scilit]
- Aberle, D.R.; Adams, A.M.; Berg, C.D.; Black, W.C.; Clapp, J.D.; Fagerstrom, R.M.; Gareen, I.F.; Gatsonis, C.; Marcus, P.M. Results of the two incidence screenings in the National Lung Screening Trial. N. Engl. J. Med. 2013, 369, 920–931. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Topol, E. Deep Medicine: How Artificial Intelligence Can Make Healthcare Human Again; Basic Books: Hachette UK, 2019. [Google Scholar]
- Abdullah, D.M.; Abdulazeez, A.M.; Sallow, A.B. Lung cancer prediction and classification based on correlation selection method using machine learning techniques. Qubahan Acad. J. 2021, 1, 141–149. [Google Scholar] [CrossRef] [Scilit]
- Dutta, B. Comparative Analysis of Machine Learning and Deep Learning Models for Lung Cancer Prediction Based on Symptomatic and Lifestyle Features. Appl. Sci. 2025, 15, 4507. [Google Scholar] [CrossRef] [Scilit]
- Chicco, D.; Jurman, G. The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. BMC Genom. 2020, 21, 6. [Google Scholar] [CrossRef] [Scilit]
- Lever, J.; Krzywinski, M.; Altman, N. Points of significance: Classification evaluation. Nat. Methods 2016, 13, 603–604. [Google Scholar] [CrossRef] [Scilit]
- He, H.; Garcia, E.A. Learning from imbalanced data. IEEE Trans. Knowl. Data Eng. 2009, 21, 1263–1284. [Google Scholar] [CrossRef] [Scilit]
- Japkowicz, N.; Stephen, S. The class imbalance problem: A systematic study. Intell. Data Anal. 2002, 6, 429–449. [Google Scholar] [CrossRef] [Scilit]
- Fawcett, T. An introduction to ROC analysis. Pattern Recognit. Lett. 2006, 27, 861–874. [Google Scholar] [CrossRef] [Scilit]
- Prokhorenkova, L.; Gusev, G.; Vorobev, A.; Dorogush, A.V.; Gulin, A. CatBoost: Unbiased boosting with categorical features. In Advances in Neural Information Processing Systems, Proceedings of the Annual Conference on Neural Information Processing Systems Montreal, QC, Canada, 3–8 December2018; NeurIPS: Red Hook, NY, USA, 2018; Volume 31. [Google Scholar]
- Peto, R.; Darby, S.; Deo, H.; Silcocks, P.; Whitley, E.; Doll, R. Smoking, smoking cessation, and lung cancer in the UK since 1950: Combination of national statistics with two case-control studies. Br. Med. J. 2000, 321, 323–329. [Google Scholar] [CrossRef] [Scilit]
- Cox, D.R. Regression models and life-tables. J. R. Stat. Soc. Ser. B (Methodol.) 1972, 34, 187–202. [Google Scholar] [CrossRef] [Scilit]
- Tammemägi, M.C.; Katki, H.A.; Hocking, W.G.; Church, T.R.; Caporaso, N.; Kvale, P.A.; Chaturvedi, A.K.; Silvestri, G.A.; Riley, T.L.; Commins, J.; et al. Selection criteria for lung-cancer screening. N. Engl. J. Med. 2013, 368, 728–736. [Google Scholar] [CrossRef] [Scilit]
- Mohanambal, K.; Nirosha, Y.; Roshini, E.O.; Punitha, S.; Shamini, M. Lung cancer detection using machine learning techniques. Int. J. Adv. Res. Electr. Electron. Instrum. Eng. 2019, 8, 266–271. [Google Scholar]
- Al-Ameer, A.A.A.; Hussien, G.A.; Al Ameri, H.A. Lung cancer detection using image processing and deep learning. Indones. J. Electr. Eng. Comput. Sci 2022, 28, 987–993. [Google Scholar] [CrossRef] [Scilit]
- Dietterich, T.G. Ensemble methods in machine learning. In Proceedings of the International Workshop on Multiple Classifier Systems, Cagliari, Italy, 21–23 June 2000; pp. 1–15. [Google Scholar]
- Hussain, L.; Almaraashi, M.S.; Aziz, W.; Habib, N.; Abbasi, S.-U.-R.S. Machine learning-based lung cancer detection using reconstruction independent component analysis and sparse filter features. Waves Random Complex Media 2024, 34, 226–251. [Google Scholar] [CrossRef] [Scilit]
- Litjens, G.; Kooi, T.; Bejnordi, B.E.; Setio, A.A.A.; Ciompi, F.; Ghafoorian, M.; Van Der Laak, J.A.; Van Ginneken, B.; Sánchez, C.I. A survey on deep learning in medical image analysis. Med. Image Anal. 2017, 42, 60–88. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chawla, N.V.; Bowyer, K.W.; Hall, L.O.; Kegelmeyer, W.P. SMOTE: Synthetic minority over-sampling technique. J. Artif. Intell. Res. 2002, 16, 321–357. [Google Scholar] [CrossRef] [Scilit]
- Yang, Z.; Liu, N.; Hu, X.B.; Jin, F. Tutorial on deep learning interpretation: A data perspective. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management, Atlanta, GA, USA, 17–21 October 2022; pp. 5156–5159. [Google Scholar]
- Nemlander, E.; Rosenblad, A.; Abedi, E.; Ekman, S.; Hasselström, J.; Eriksson, L.E.; Carlsson, A.C. Lung cancer prediction using machine learning on data from a symptom e-questionnaire for never smokers, formers smokers and current smokers. PLoS ONE 2022, 17, e0276703. [Google Scholar] [CrossRef] [Scilit]
- Nabeel, S.M.; Bazai, S.U.; Alasbali, N.; Liu, Y.; Ghafoor, M.I.; Khan, R.; Ku, C.S.; Yang, J.; Shahab, S.; Por, L.Y. Optimizing lung cancer classification through hyperparameter tuning. Digit. Health 2024, 10, 20552076241249661. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Dritsas, E.; Trigka, M. Lung cancer risk prediction with machine learning models. Big Data Cogn. Comput. 2022, 6, 139. [Google Scholar] [CrossRef] [Scilit]
- Lundberg, S.M.; Lee, S.-I. A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems, Proceedings of the Annual Conference on Neural Information Processing Systems, Long Beach, CA, USA, 4–9 December 2017; NeurIPS: Red Hook, NY, USA, 2017; Volume 30. [Google Scholar]
- Husaini, Y.N.A. Lung Cancer Survey Dataset. Kaggle. 2025. Available online: https://www.kaggle.com/code/oman0086/stbased-lung-cancer (accessed on 2 December 2025).
- Mamun, M.; Farjana, A.; Al Mamun, M.; Ahammed, M.S. Lung cancer prediction model using ensemble learning techniques and a systematic review analysis. In Proceedings of the 2022 IEEE World AI IoT Congress (AIIoT), Seattle, WAS, USA, 6–9 June 2022; pp. 187–193. [Google Scholar]
- Vieira, E.; Ferreira, D.; Neto, C.; Abelha, A.; Machado, J. Data mining approach to classify cases of lung cancer. In Proceedings of the World Conference on Information Systems and Technologies, Azores, Portugal, 30 March–1 April 2021; pp. 511–521. [Google Scholar]
- Maurya, S.P.; Sisodia, P.S.; Mishra, R.; Singh, D.P. Performance of machine learning algorithms for lung cancer prediction: A comparative approach. Sci. Rep. 2024, 14, 18562. [Google Scholar] [CrossRef] [Scilit] [PubMed]













| Variable | Actual | Encoded | Missing-Entry Treatment |
|---|---|---|---|
| LUNG_CANCER (target) | “YES”, “NO” | NO → 0, YES → 1 | The released dataset contained no missing entries; nevertheless, the pipeline includes checks to drop missing targets and impute missing predictors using the training-set median |
| AGE | numeric (years) | kept numeric (no recoding) | |
| GENDER | categorical (Male/Female) | Male → 1, Female → 0 | |
| SMOKING | {1, 2} (survey code) | 1 → 0 (No), 2 → 1 (Yes) | |
| All other variables | {1, 2} | 1 → 0, 2 → 1 |
| Dataset | Accuracy | ROC-AUC | Precision | Recall | F1-Score | Specificity |
|---|---|---|---|---|---|---|
| Validation | 0.9516 ± 0.84 | 0.9375 ± 0.96 | 0.943 ± 1.02 | 0.981 ± 0.73 | 0.962 ± 0.81 | 0.625 ± 4.21 |
| Test | 0.9516 ± 0.91 | 0.9375 ± 1.03 | 0.9474 ± 1.18 | 1.0000 ± 0.00 | 0.9730 ± 0.74 | 0.6250 ± 4.12 |
| Threshold | Accuracy | Precision | Recall | F1-Score | FN Count |
|---|---|---|---|---|---|
| 0.50 (default) | 0.8871 | 0.9636 | 0.8889 | 0.9254 | 6 |
| 0.19 (optimized) | 0.9516 | 0.9474 | 1.0000 | 0.9730 | 0 |
| Error Type | Count | Rate (%) |
|---|---|---|
| True Positive (TP) | 54 | 87.1 |
| True Negative (TN) | 5 | 8.1 |
| False Positive (FP) | 3 | 4.8 |
| False Negative (FN) | 0 | 0.0 |
| Rank | Feature | Importance Value | Effect Direction |
|---|---|---|---|
| 1 | Coughing | 0.184 | ↑ increases risk |
| 2 | Wheezing | 0.162 | ↑ increases risk |
| 3 | Alcohol Consumption | 0.109 | ↑ increases risk |
| 4 | Swallowing Difficulty | 0.093 | ↑ increases risk |
| 5 | Allergy | 0.071 | ↑ increases risk |
| Risk Group | N | Cancer Prevalence (%) | FN Rate (%) |
|---|---|---|---|
| Low | 9 | 22.2 | 0.0 |
| Medium | 17 | 58.8 | 0.0 |
| High | 36 | 94.4 | 0.0 |
| Ref. | Dataset/Protocol (as Reported) | Best Model Reported | Accuracy (%) | ROC-AUC (%) | Precision (%) | Recall/Sensitivity (%) | F1-Score (%) | Specificity (%) |
|---|---|---|---|---|---|---|---|---|
| [29] | Lung Cancer Survey (309); SMOTE + 10-fold CV | XGBoost | 94.42 | 98.14 | 95.66 | 94.46 | 94.74 | — |
| [30] | Lung Cancer Survey (reported); protocol not fully specified in excerpt | ANN | 93.00 | — | 91.00 | 96.00 | — | 90.00 |
| [31] | Small “binary characteristics” dataset; best accuracy reported as 92.86% | KNN | 92.86 | — | — | — | — | — |
| Proposed (SETI-LC): Lung Cancer Survey (309); stratified 60/20/20; class weights; threshold optimized | CatBoost | 95.16 | 93.75 | 94.74 | 98.9 | 97.30 | 62.50 | |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the author. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Husaini, Y.A. Symptom-Based Lung Cancer Prediction Using Ensemble Learning with Threshold Optimization and Interpretability. Information 2026, 17, 172. https://doi.org/10.3390/info17020172
Husaini YA. Symptom-Based Lung Cancer Prediction Using Ensemble Learning with Threshold Optimization and Interpretability. Information. 2026; 17(2):172. https://doi.org/10.3390/info17020172
Chicago/Turabian StyleHusaini, Yousuf Al. 2026. "Symptom-Based Lung Cancer Prediction Using Ensemble Learning with Threshold Optimization and Interpretability" Information 17, no. 2: 172. https://doi.org/10.3390/info17020172
APA StyleHusaini, Y. A. (2026). Symptom-Based Lung Cancer Prediction Using Ensemble Learning with Threshold Optimization and Interpretability. Information, 17(2), 172. https://doi.org/10.3390/info17020172

