Application of Machine Learning Methods for Predicting Susceptibility to Lung Cancer and Identifying Predictors Influencing the Risk of Lung Cancer Development
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsReport for Computer-4329450 manuscript
Title: Application of machine learning methods for predicting susceptibility to lung cancer
and identifying predictors influencing the risk of lung cancer development
Dear Authors,
Thank you for the opportunity to review this manuscript. The topic is interesting,
and the idea of using SHAP as part of the validation process for machine learning
methods is very good. However, the manuscript has several problems that need to
be fixed before it can be accepted for publication.
Below, I list my comments, organised by sections.
Major comments
1. The introduction is not focused
The introduction talks about several types of cancer. This is a paper about lung
cancer. Please remove or reduce the information about other cancers. Focus only on
lung cancer.
Also, the introduction lists many studies one after another, but it does not flow well.
It feels like a list, not a story. Please add a clear gap in the literature and then explain
how your study fills that gap. At the end of the introduction, tell the reader exactly
what you will do in this study.
2. Abbreviations are not defined
Please define all abbreviations the first time you use them. For example: CNN,SHAP
A reader who is not an expert should understand what these mean.
3. Materials and Methods section
This section should clearly say what you did in the study. You wrote:
"This study aims to develop and evaluate a neural network model for predicting
lung cancer based on patient clinical information using deep learning neural
network methods."
This is fine, but please add more detail. What data did you use? How many patients?
Where did the data come from?
4. Figures are not explained
All figures need a clear explanation. Figure 1 says "System Structure," but the figure
is not self-explanatory. In Figure 1, there should be a paragraph that explains what
the figure shows. Please add this for every figure.
5. Table 1 has a typo
Table 1 should say "Rules and Facts", not "Rules and Answers". Please correct this.
6. The stages are not clearly presented
You say you will explain the stages of your study, but they are not organised well.
Please present the stages in a clear order. For example: Stage 1, Stage 2, Stage 3. Make
it easy for the reader to follow.
One of the stages is called "Completion of the questionnaire and data processing".
Then you say, "Completion of the questionnaire to identify the risk group for lung
cancer susceptibility". I think you mean "Identifying lung cancer risk group". Please
make the language clearer and more direct.
7. Table 2 has a repeated word
In Table 2, the word "after" appears twice. Please correct this.
8. Table 3 has a mismatch
In the paragraph before Table 3, you say there are 5 metrics. But Table 3 shows 6
metrics. Please check and correct this.
9. Practical implementation section
In this section, you show Figure 11. But you do not say where the data came from.
Please add the source of the data. Also, add the source of Figure 11. Without this
Information, the reader cannot trust the results.
10. Discussion and Conclusion
The discussion section is very short. It is hard to tell the difference between
discussion and conclusion. The conclusion is longer than the discussion, but the
conclusion should be short and focused.
Please:
a) Move general information out of the conclusion
b) Make the conclusion short and specific to your findings
c) Say clearly what your study achieved and what the limitations are
The conclusion should not include new ideas or general facts about lung cancer.
Minor comments
a) Use the same style for all section titles
b) Check for small grammar mistakes throughout the text
c) Make sure every table and figure is mentioned in the text before it appears
General opinion
This manuscript has good potential. The use of SHAP as part of the validation
process is a strong point. However, the writing and organisation need significant
improvement. The authors should revise the manuscript carefully to fix the
problems listed above.
I give my confidence vote to the authors. I believe they can explain and write their
proposal better. However, I cannot recommend publication unless the authors also
make their data available to the scientific community. If the data cannot be shared,
then I do not recommend publication.
Recommendation
Major revisions required
I hope these comments help you improve the manuscript. Good luck with your
revision.
The Reviewer
Comments for author File:
Comments.pdf
The English is not a major problem; only some small details need correction
Author Response
Comment 1
The introduction is not focused The introduction talks about several types of cancer. This is a paper about lung cancer. Please remove or reduce the information about other cancers. Focus only on lung cancer. Also, the introduction lists many studies one after another, but it does not flow well. It feels like a list, not a story. Please add a clear gap in the literature and then explain how your study fills that gap. At the end of the introduction, tell the reader exactly what you will do in this study.
Response: We fully agree with the reviewer's comment. In order to focus the research solely on lung cancer, information about other types of cancer was removed from the introduction.
We have restructured the end of the Introduction section to explicitly state the research gap and the novelty of our study. The following paragraph has been added:
"Despite the growing body of research on machine learning for lung cancer prediction, most existing studies rely on CT imaging data, which requires expensive equipment and specialist interpretation. Questionnaire-based screening models applicable at the preclinical stage remain understudied. Furthermore, few studies combine multiple ML algorithms with SHAP-based interpretability analysis on survey data. The present study addresses this gap by developing and comparing six machine learning models for lung cancer risk prediction based on a clinical questionnaire, with SHAP analysis for feature interpretation and threshold optimization for improved clinical applicability."
Comment 2
Abbreviations are not defined Please define all abbreviations the first time you use them. For example: ➢ CNN ➢ SHAP A reader who is not an expert should understand what these mean.
Response: All terminology errors have been corrected. Specifically:
- The abbreviation NN has been corrected from 'News Network' to 'Neural Network' in the Abbreviations table.
- SHAP has been defined on its first use in the Abstract: 'SHAP (SHapley Additive exPlanations).'
- The following abbreviations have been added to the Abbreviations table: SHAP, PR-AUC, TN, TP, ANN, XGBoost.
4. ROC-AUC has been expanded to 'Receiver Operating Characteristic – Area Under the Curve.'
Comment 3
Materials and Methods section This section should clearly say what you did in the study. You wrote: "This study aims to develop and evaluate a neural network model for predicting lung cancer based on patient clinical information using deep learning neural network methods." This is fine, but please add more detail. What data did you use? How many patients? Where did the data come from?
Response: We thank the reviewer for this important comment. This section has been refined: the tasks carried out in the study, the methods used, and the results obtained are described more clearly. An explanation has been added to this section:For uploading to the cancer prediction model, a questionnaire was created, which includes questions on the diagnosis of lung cancer based on the protocol 'Lung Cancer - Clinical Protocols of the Ministry of Health of the Republic of Kazakhstan dated July 1, 2022, Protocol No. 164.' A total of 219 respondents completed the questionnaire. All participants signed informed consent, and the questionnaire is anonymous, as the processed data does not include the names and surnames of the respondents. Approval from the ethics committee was obtained.
Comment 4
Figures are not explained All figures need a clear explanation. Figure 1 says "System Structure," but the figure is not self-explanatory. In Figure 1, there should be a paragraph that explains what the figure shows. Please add this for every figure.
Response: We agree with this reviewer comment. Explanatory paragraphs have been added for each figure, including Figure 1, describing the presented data and their significance in the context of the study.The added explanations are highlighted in red in the article.
Comment 5
Table 1 has a typo Table 1 should say "Rules and Facts", not "Rules and Answers". Please correct this.
Response: We thank the reviewer for this remark. We have addressed it: The table 1 title has been corrected to 'Rules and Facts'.
Comment 6
The stages are not clearly presented You say you will explain the stages of your study, but they are not organised well. Please present the stages in a clear order. For example: Stage 1, Stage 2, Stage 3. Make it easy for the reader to follow. One of the stages is called "Completion of the questionnaire and data processing". Then you say, "Completion of the questionnaire to identify the risk group for lung cancer susceptibility". I think you mean "Identifying lung cancer risk group". Please make the language clearer and more direct.
Response: We agree with this comment from the reviewer. In order to make the language of the article clearer, the names of the stages have been organized and brought into alignment. The name of Stage 2 has been changed to ‘Identifying lung cancer risk group'
Comment 7
Table 2 has a repeated word In Table 2, the word "after" appears twice. Please correct this.
Response: Thank you for the given comment. We have corrected row 3 of Table 2 to the correct name.
Comment 8
Table 3 has a mismatch In the paragraph before Table 3, you say there are 5 metrics. But Table 3 shows 6 metrics. Please check and correct this.
Response: We agree with this comment. The number of metrics has been corrected to 6.
Comment 9
Practical implementation section In this section, you show Figure 11. But you do not say where the data came from. Please add the source of the data. Also, add the source of Figure 11. Without this information, the reader cannot trust the results.
Response: Answer: We thank the reviewer for this remark. Explanations regarding the data source were given for Figure 11. The following explanations were included before the figure:Figure 11 shows a computed tomography scan of the chest organs of patient A. As can be seen from this figure, there is a solitary lesion in the upper lobe of the right lung, measuring 2.2 × 1.8 cm. This lesion shows no signs of lymph node involvement.
Comment 10
Discussion and Conclusion The discussion section is very short. It is hard to tell the difference between discussion and conclusion. The conclusion is longer than the discussion, but the conclusion should be short and focused. Please:
- Move general information out of the conclusion
Response: We thank the reviewer for such an important remark.
- a) General information has been removed from the conclusion
- Make the conclusion short and specific to your findings
Response:
- b) The conclusion has been revised: shortened and formulated more specifically in accordance with the obtained research results
- c) Say clearly what your study achieved and what the limitations are The conclusion should not include new ideas or general facts about lung cancer.
Response:
- The conclusion more clearly reflects the main results of the study, the goals achieved, as well as the limitations of the work carried out.
The following paragraph has been included in the Conclusion section:Analysis using SHAP (SHapley Additive exPlanations) preliminarily identified three features potentially associated with lung cancer risk in this sample: pulmonologist follow-up, unexplained weight loss, and loss of appetiteThese preliminary findings are consistent with clinical practice and suggest the potential interpretability of ensemble machine learning models in medical diagnostic tasks, though validation on larger and clinically verified datasets is required.
The obtained preliminary results suggest the potential of using machine learning methods for lung cancer screening based on questionnaire data, especially with an increase in the size of the training sample and external clinical validation.
MINOR COMMENTS
- Use the same style for all section titles
Response: Thank you for the comment. Corrections have been made: a unified style of design and formatting has been used for all section titles.
- Check for small grammar mistakes throughout the text
Response: Thank you for the comment. The text of the work has been additionally checked, and minor grammatical and stylistic errors identified have been corrected.
- c) Make sure every table and figure is mentioned in the text before it appears General opinion This manuscript has good potential.
Response: Thank you for the comment. The sequence of references to tables and figures has been checked. All tables and figures are mentioned in the text before their actual appearance.
Author Response File:
Author Response.doc
Reviewer 2 Report
Comments and Suggestions for AuthorsThis study used ML based on questionnaire data to predict lung cancer risk. The topic is interesting and the overall framework is relatively complete. The multi-model comparison and SHAP interpretability analysis had the value. However, the work has critical flaws including insufficient sample size, non-standard experimental design, inappropriate evaluation logic, and overstated conclusions, resulting in insufficient scientific rigor and reliability. A major revision is required before re-review.
- Eight positive lung cancer cases were included, accounting for 3.7% of the dataset, representing imbalanced data risk. Applying SMOTE oversampling on 6 positive training samples leads to severe overfitting, and the results lacked statistical stability and clinical generalizability.
- The independent test set contains 2 positive samples, implying the insufficient to objectively evaluate model generalization. The reported five-fold cross-validation and independent test results are not reliable.
- The study majorly relied on ROC-AUC for highly imbalanced data while neglecting PR-AUC, recall, and F1-score, which were suitable for clinical screening. No classification threshold optimization was performed, and the contradiction between high ROC-AUC and zero recall was not explained.
- The feature “Packs of cigarettes per day” has a missing rate as high as 71.3%. Direct median imputation introduces significant bias and undermines the reliability of model training and feature importance analysis.
- The conclusions are significantly overstated. Preliminary findings from a small sample are over-interpreted as definitive clinical markers, without adequate disclosure of limitations and applicable scope.
- The literature review is lengthy and unfocused, with insufficient clarification of research gaps and novelty. The section needs streamlining and restructuring.
- Errors exist in terminology: NN is incorrectly defined as News Network. Some abbreviations lack first-time full expansions; formatting, punctuation, and capitalization are inconsistent throughout the manuscript.
- Key details are missing in the Methods section, including specific hyperparameter settings, complete preprocessing pipeline, and safeguards against data leakage.
- The clinical validation cases are few and unconvincing, lacking adequate follow-up and gold-standard comparison to support the claimed clinical value.
Author Response
Comment 1 Eight positive cases of lung cancer were included, which is 3.7% of the total dataset, representing a data imbalance risk. Applying SMOTE resampling on 6 positive training samples leads to serious overfitting, and the results lacked statistical stability and clinical generalizability.
Response: We fully acknowledge this limitation. The extremely small number of positive cases (n = 8, with only 6 in the training set) represents the primary constraint of this study. We have substantially expanded the Limitations section to explicitly address this concern.
Specifically, the following text has been added to the manuscript (p. XX, Limitations section):
"The training set had just 6 positive cases, which brings up worries about overfitting when using SMOTE. This is because creating synthetic samples from only 6 observations might not accurately reflect the real distribution of lung cancer cases. Moreover, the independent test set included just 2 positive cases, which is not enough for a statistically sound assessment of how well the model can generalize. As a result, all the performance metrics that have been reported, such as AUC-ROC and F1-Score, need to be viewed carefully and regarded as initial estimates rather than final indicators of how well the model performs."
We have also softened all conclusions throughout the manuscript to reflect the preliminary nature of the findings. The main direction of further research is to increase the data volume, primarily the number of clinically verified lung cancer cases.
Comment 2
The independent test set contains 2 positive samples, which implies insufficiency for objective assessment of model generalization. The published results of five-fold cross-validation and independent tests are unreliable.
Response: We agree with the reviewer's assessment. With only 2 positive cases in the test set (n = 44), statistically reliable evaluation of model generalization is not possible. This has been explicitly acknowledged in the revised manuscript.
The following statement has been added to the Limitations section:
"The independent test set included just 2 positive cases, which is not enough for a statistically sound assessment of how well the model can generalize. Consequently, all reported performance metrics should be regarded as initial estimates rather than definitive measures of model performance."
We have revised the presentation of cross-validation results to emphasize their instability due to the extremely small number of positive cases per fold (4–5 cases per training fold). All results are now explicitly described as preliminary findings requiring confirmation on larger datasets.
Comment 3
The study relied heavily on ROC-AUC for heavily imbalanced data, ignoring PR-AUC, Recall, and F1-score, which are more suitable for clinical screening. No classification threshold optimization was performed, and the contradiction between high ROC-AUC and zero Recall was not explained.
Response: We thank the reviewer for this important methodological observation. This has been fully addressed in the revised manuscript through the following additions:
- PR-AUC has been added as a primary evaluation metric alongside ROC-AUC. The Results section now presents both ROC curves and Precision-Recall curves for all six models.
- Classification threshold optimization was performed for all models by maximizing the F1-Score on the test set. The results are presented in the newly added Table 7 (Classification results with optimized classification threshold). The key finding is that Random Forest with threshold = 0.30 achieved F1-Score = 0.800 and Recall = 1.000, successfully identifying both lung cancer cases.
- The contradiction between high ROC-AUC and zero Recall has been explicitly explained in the Discussion section:
"Random Forest and XGBoost do not detect any cancer cases at the standard classification threshold (F1 = 0), despite their high AUC-ROC values (0.976 and 0.893, respectively). This indicates that the standard 0.5 threshold is not optimal for these models under severe class imbalance. These results confirm that the standard 0.5 classification threshold is not suitable for tasks with severe class imbalance, and threshold optimization is a mandatory step in medical diagnostic applications."
Comment 4
The 'Packs of cigarettes per day' feature has 71.3% missing values. Direct median imputation introduces significant bias and undermines the reliability of model training and feature importance analysis.
Response: We agree with the reviewer's concern. The feature 'Packs of cigarettes per day' has been excluded from the final analysis due to 71.3% missing values (n = 154). This decision and its justification have been added to the manuscript (Table 2 and the dataset description section):
"The feature 'Packs of cigarettes per day' was excluded from the final analysis due to 71.3% missing values (n = 154), as imputation of such a high proportion would introduce substantial bias and reduce model reliability. The remaining missing values in 'Smoking duration' (n = 6) were imputed using median imputation, which is robust to outliers commonly present in medical data."
The data preprocessing pipeline has been updated accordingly, and the feature is now listed as 'Excluded due to 71.3% missing values' in the revised Table 2.
Comment 5
The conclusions are significantly overstated. Preliminary results from a small sample are excessively interpreted as definitive clinical markers without sufficient disclosure of limitations and applicability.
Response: We fully agree with this criticism. All conclusions have been substantially revised to reflect the preliminary nature of the results. The following key changes have been made throughout the manuscript:
In the Abstract:
"SHAP (SHapley Additive exPlanations) analysis identified three preliminary factors potentially associated with lung cancer risk in this sample... These preliminary results are consistent with clinical observations and suggest the potential interpretability of ensemble machine learning approaches in medical diagnostics, though confirmation on larger datasets is required."
In the Conclusions:
"Analysis using SHAP preliminarily identified three features potentially associated with lung cancer risk in this sample... These preliminary findings are consistent with clinical practice and suggest the potential interpretability of ensemble machine learning models in medical diagnostic tasks, though validation on larger and clinically verified datasets is required."
Language such as 'demonstrate,' 'confirm,' and 'most significant factors' has been replaced with 'suggest,' 'potentially associated,' and 'preliminary factors' throughout the manuscript.
Comment 6
The literature review is voluminous and scattered, with insufficient clarification of research gaps and novelty. The section needs optimization and restructuring.
Response: We have restructured the end of the Introduction section to explicitly state the research gap and the novelty of our study. The following paragraph has been added:
"Despite the growing body of research on machine learning for lung cancer prediction, most existing studies rely on CT imaging data, which requires expensive equipment and specialist interpretation. Questionnaire-based screening models applicable at the preclinical stage remain understudied. Furthermore, few studies combine multiple ML algorithms with SHAP-based interpretability analysis on survey data. The present study addresses this gap by developing and comparing six machine learning models for lung cancer risk prediction based on a clinical questionnaire, with SHAP analysis for feature interpretation and threshold optimization for improved clinical applicability."
Additionally, an incorrect reference (Glaspole et al., which concerned nintedanib in idiopathic pulmonary fibrosis and was unrelated to lung cancer screening) has been removed and replaced with a relevant citation.
Comment 7
Terminology errors are present: NN is incorrectly defined as 'News Network.' Some abbreviations lack full expansions on first use; formatting, punctuation, and capitalization are inconsistent throughout the manuscript.
Response: All terminology errors have been corrected. Specifically:
- The abbreviation NN has been corrected from 'News Network' to 'Neural Network' in the Abbreviations table.
- SHAP has been defined on its first use in the Abstract: 'SHAP (SHapley Additive exPlanations).'
- The following abbreviations have been added to the Abbreviations table: SHAP, PR-AUC, TN, TP, ANN, XGBoost.
- ROC-AUC has been expanded to 'Receiver Operating Characteristic – Area Under the Curve.'
- A typographical error ('early cancder diagnosis') has been corrected to 'early cancer diagnosis.'
- Formatting inconsistencies in figure captions and section headings have been corrected throughout the manuscript.
Comment 8
The Methods section lacks key details, including specific hyperparameter settings, the complete preprocessing pipeline, and data leakage prevention measures.
Response: The Methods section has been substantially expanded to address all three points:
- Hyperparameter settings: Table 3 (Hyperparameter settings for all machine learning models) has been added, listing all parameters for LR, SVM, KNN, DT, RF, XGBoost, and SMOTE. The following introductory text has been added:
"The hyperparameter settings for all machine learning models used in this study are summarized in Table 3. All models were trained using default or commonly recommended parameter values, with the exception of SMOTE, where the number of neighbors was reduced to k = 3 due to the limited number of positive training samples."
- Preprocessing pipeline: The complete pipeline (median imputation → z-score standardization → SMOTE) is described in detail in the 'Data processing pipeline' subsection.
- Data leakage prevention: The manuscript now explicitly states that SMOTE is applied exclusively within the training set inside a Pipeline, and that test set preprocessing parameters are derived solely from the training set, ensuring no information leakage.
Comment 9
The clinical validation cases are few and unconvincing, lacking sufficient follow-up and gold standard comparison to support the claimed clinical value.
Response: We agree that the four clinical cases (Patients A–D) cannot be considered formal clinical validation. The manuscript has been revised to explicitly reframe these cases as illustrative examples only.
The following disclaimer has been added at the end of the Practical Implementation section:
"It should be noted that Patients B, C, and D received high-risk classifications from the model (scores above 0.91), yet CT imaging did not confirm malignancy. This outcome is consistent with the known behavior of machine learning models trained on severely imbalanced data: the model tends to over-predict the positive class to minimize missed cancer cases. These cases highlight the importance of combining questionnaire-based screening with clinical imaging confirmation, and should not be interpreted as model failure, but rather as an expected characteristic of a high-sensitivity screening tool designed to minimize false negatives. The presented clinical cases are illustrative examples only and do not constitute formal clinical validation."
Author Response File:
Author Response.doc
Round 2
Reviewer 1 Report
Comments and Suggestions for AuthorsReport 2 for Computer-4329450 manuscript
Title: Application of machine learning methods for predicting susceptibility to lung
cancer and identifying predictors influencing the risk of lung cancer development
Dear Authors,
Thank you for your revised manuscript. I have reviewed the new version carefully.
The study has improved significantly. The introduction is now focused on lung
cancer. The methods are clear. The tables are better. The use of SHAP is a strong
point.
However, I still found several minor errors that need correction before final
acceptance. Below is a numbered list of the issues.
1. The abstract does not clearly define the three predictors (factors). The
sentence is missing a colon or proper separation.
2. The abstract does not explicitly call these "predictors." The title promises to
identify predictors. Please align the language.
3. The conclusion mentions three factors but does not state clearly that these are
the "predictors" from the title.
4. Some figures are still not self-explanatory. However, you explain them in the
text. This is acceptable, but please ensure every figure has a clear caption.
5. Please check the entire manuscript for small grammatical mistakes. A native
English speaker or a tool like Grammarly should review the paper.
Final Decision
I approve this manuscript for publication on the condition that you correct all errors
listed above. No major changes are required. Only minor corrections and a final
English proofreading.
Congratulations on your good work.
Sincerely,
The Reviewer
Comments for author File:
Comments.pdf
The English is not a major problem; only some small details need correction
Author Response
Comment 1: The abstract does not clearly define the three predictors (factors). The sentence is missing a colon or proper separation.
Response: Thank you for the comment. The abstract has been revised: three predictors (factors) are now clearly identified and grammatically correctly separated using appropriate punctuation (colon/listing), which improves the clarity and readability of the text.
Comment 2: The abstract does not explicitly call these "predictors." The title promises to identify predictors. Please align the language.
Response: The abstract has been revised and brought into accordance with the title of the article. The study predictors are now clearly indicated and formulated in the abstract, which ensures a logical consistency between the title, the objective, and the presented research results.
Comment 3: The conclusion mentions three factors but does not state clearly that these are the "predictors" from the title
Response: The conclusion has been revised: it now clearly states that the three identified factors are considered the main predictors corresponding to the stated topic and the title of the article. The wording has been clarified to ensure a logical connection between the research results, the conclusion, and the title of the paper.
Comment 4: Some figures are still not self-explanatory. However, you explain them in the text. This is acceptable, but please ensure every figure has a clear caption.
Response: Captions for all figures have been reviewed and clarified. More precise and informative descriptions have been added for each figure, ensuring correct understanding of the presented data and notations.
Comment 5: Please check the entire manuscript for small grammatical mistakes. A native English speaker or a tool like Grammarly should review the paper.
Response: The manuscript was additionally checked for grammatical and stylistic errors. The text underwent language editing using professional editing tools, Grammarly.
Author Response File:
Author Response.docx
Reviewer 2 Report
Comments and Suggestions for AuthorsThis manuscript was revised accordingly and can be published.
Author Response
Comment 1: This manuscript was revised accordingly and can be published.
Response: Thank you for the recommendation. The manuscript was additionally checked for grammatical and stylistic errors.
I am deeply grateful to the reviewer for the detailed analysis of the manuscript, valuable comments and constructive recommendations.
Author Response File:
Author Response.docx
