Previous Article in Journal
Targeting Fungal Adaptive Networks and Emerging Molecular Targets for Next-Generation Antifungal Therapeutics
 
 
Article
Peer-Review Record

Deep Learning-Based Molecular Generation for Lung Cancer Therapeutics

Drugs Drug Candidates 2026, 5(3), 48; https://doi.org/10.3390/ddc5030048
by Mohavia Ben Amid Sinon * and Uche A. K. Chude-Okonkwo
Reviewer 1: Anonymous
Reviewer 2: Anonymous
Drugs Drug Candidates 2026, 5(3), 48; https://doi.org/10.3390/ddc5030048
Submission received: 7 May 2026 / Revised: 17 August 2026 / Accepted: 19 August 2026 / Published: 26 August 2026
(This article belongs to the Section In Silico Approaches in Drug Discovery)

Round 1

Reviewer 1 Report

Comments and Suggestions for Authors

The manuscript presents an AI-driven molecular generation framework integrating fragment-based design, RGCN, and WGAN approaches for the discovery of novel P2X7 receptor-targeted compounds for lung cancer therapy. However, following queries are expected to be resolved by the authors.

  1. Abstract – The authors should explicitly state that the biological activity of the generated compounds was not experimentally evaluated and that the conclusions are based solely on computational analyses, including docking studies.
  2. Introduction – The authors should provide a stronger justification for selecting the P2X7 receptor over other established lung cancer targets and clarify how the proposed framework improves upon previously reported RGCN- or GAN-based molecular generation approaches.
  3. Methodology – The study was trained on only 5,000 SMILES derived from nine seed molecules. The authors should discuss the potential impact of this limited chemical diversity on model performance and justify why this dataset size is sufficient for training the proposed framework.
  4. Methodology – No comparison with existing generative models (e.g., VAE, MolGAN, MedGAN) is presented. Including a benchmark comparison would allow readers to assess the relative performance of the proposed method.
  5. Results – The highest Tanimoto similarity achieved was 0.59, below the stated target threshold of 0.65. The authors should discuss the implications of this result more critically rather than attributing it solely to computational limitations.
  6. Docking Analysis – The docking study was performed on only eighteen generated molecules selected based on Tanimoto similarity. The rationale for selecting these compounds should be clarified, and additional top-ranked molecules should be evaluated to better support the conclusions.
  7. Docking Analysis – The reported improvement in docking scores relative to seed molecules appears modest. Statistical analysis should be provided to determine whether the observed differences are meaningful.
  8. Results and Discussion – The study focuses primarily on QED, LogP, and Lipinski parameters. Inclusion of additional metrics such as synthetic accessibility, PAINS filtering, or ADMET predictions would strengthen the assessment of drug-likeness.
  9. Conclusion – The statement that the generated molecules represent promising therapeutic candidates is not fully supported by the presented data. The conclusions should be limited to computational prioritization until further validation is performed.
  10. References – Several formatting inconsistencies were noted in the reference list and should be corrected according to the journal guidelines.

Author Response

Comments 1: [Abstract – The authors should explicitly state that the biological activity of the generated compounds was not experimentally evaluated and that the conclusions are based solely on computational analyses, including docking studies.]

 

Response 1: [Thank you for the comment. This has been added: “Nevertheless, the biological activity of the generated compounds remains unvalidated experimentally and the findings are based solely on computational analyses.” Lines 20-21]

 

Comments 2: [Introduction – The authors should provide a stronger justification for selecting the P2X7 receptor over other established lung cancer targets and clarify how the proposed framework improves upon previously reported RGCN- or GAN-based molecular generation approaches.]

 

Response 2: [Thank you for the valuable comment. Additional literature on P2X7R has been added to strengthen its selection as a lung cancer target. The proposed framework improves on previous RGCN- and GAN-based methods by using a fragment-based generation strategy to address limited data and increase the diversity of training molecules. Lines 31-46]

 

 

Comments 3: [Methodology – The study was trained on only 5,000 SMILES derived from nine seed molecules. The authors should discuss the potential impact of this limited chemical diversity on model performance and justify why this dataset size is sufficient for training the proposed framework.]

 

Response 3: [Thank you for the suggestion. This has been rectified with the paragraph “A dataset containing 100,000 unique SMILES was generated using a fragment-based molecular generation method [33] from nine P2X7R-targeting seed molecules. Due to computational resource limitations, a random subset of 5,000 SMILES was selected for this study. While this reduction improves training efficiency, it may limit the coverage of the chemical space represented in the dataset and constrain the model’s ability to fully represent all chemically relevant structures associated with the P2X7R. However, random sampling helps mitigate selection bias and increases the likelihood that the subset retains the statistical characteristics of the original dataset. We acknowledge that this reduction may still restrict the diversity of chemical space explored. The selected molecules’ sizes and characteristics were analysed using the open-source cheminformatics toolkit RDKit.” Lines 330-341]

 

 

Comments 4: [Methodology – No comparison with existing generative models (e.g., VAE, MolGAN, MedGAN) is presented. Including a benchmark comparison would allow readers to assess the relative performance of the proposed method.]

Response 4: [Thank you for this valuable suggestion. We attempted to benchmark our method against MolGAN and MedGAN; however, these models were unable to generate valid molecules from the dataset used in this study. Therefore, a fair comparison was not possible. We have acknowledged this limitation in the revised manuscript and identified benchmarking with compatible generative models as future work.in the conclusion. Lines 545-547]

 

Comments 5: [Results – The highest Tanimoto similarity achieved was 0.59, below the stated target threshold of 0.65. The authors should discuss the implications of this result more critically rather than attributing it solely to computational limitations.]

 

Response 5:[I appreciate the comment. This has been replaced with the paragraph “The ideal Tanimoto similarity score should be ≥ 0.65; however, our highest score is only 0.59 when compared to the seed data molecules. This result indicates that the generated compounds exhibit moderate structural similarity to the seed data molecules. This outcome may indicate that the model is exploring novel regions of chemical space while remaining partially anchored to the seed data molecules. In de novo drug design, such moderate similarity can be interpreted as a trade-off between novelty and structural conservation, where lower similarity values may reflect increased molecular diversity. This might also be due to the model being trained on just 5,000 SMILES from the 100,000 generated fragmented molecules. This limited training set may cause the model to lack adequate structural diversity, resulting in lower Tanimoto scores even when generated molecules remain structurally similar to the seeds. Despite the score not reaching the ideal Tanimoto similarity threshold, the generated molecules show sufficient diversity, a key indicator of chemical novelty. Two molecules with the highest Tanimoto are represented in Figure 7.” Lines 208-220]

 

Comments 6: [Docking Analysis – The docking study was performed on only eighteen generated molecules selected based on Tanimoto similarity. The rationale for selecting these compounds should be clarified, and additional top-ranked molecules should be evaluated to better support the conclusions.]

 

Response 6: [Thank you for this comment. It has been addressed with the paragraph “After performing the Tanimoto similarity test between the novel molecules and the seed molecules, eighteen novel compounds with the highest Tanimoto scores were selected, along with the six seed molecules to which they showed the highest similarity, for docking against P2X7R.This criterion ensures that the evaluated molecules retain partial structural correspondence to known P2X7R active molecules while still exhibiting novelty. This approach was adopted to prioritise compounds most likely to occupy relevant binding regions of the receptor.” Lines 223-229

And additional analysis have been performed on top-ranked molecules . Table 6 and 7]

 

Comments 7: [Docking Analysis – The reported improvement in docking scores relative to seed molecules appears modest. Statistical analysis should be provided to determine whether the observed differences are meaningful.]

 

Response 7 : [Thank you for the suggestion. The Statistics are provided in Table 5 and last paragraph of section 3.4.2 and  we also added the paragraph “The p-value obtained from the docking scores is 0.714, which is greater than 0.05.This indicates that the difference in docking scores between novel and seed molecule are not statistically significant. Which indicated that sore differences are likely due to chance rather than a statistically meaningful improvement in binding.” Lines 249-252]

 

 

Comment 8: [Results and Discussion – The study focuses primarily on QED, LogP, and Lipinski parameters. Inclusion of additional metrics such as synthetic accessibility, PAINS filtering, or ADMET predictions would strengthen the assessment of drug-likeness.]

Response:[I appreciate the comment. We expanded the results by incorporating Synthetic Accessibility (SA), PAINS filtering, and ADMET-related analyses using ADMETlab 3.0 for the eighteen novel compounds with the highest Tanimoto scores generated compounds. Section 3.3.7

Comments 9: [Conclusion – The statement that the generated molecules represent promising therapeutic candidates is not fully supported by the presented data. The conclusions should be limited to computational prioritization until further validation is performed.]

 

Response 9: [Thank you for the comment. I replaced last two paragraph of the conclusion with “Despite its success, the study was constrained by limited computational resources and the relatively small training dataset, which may have affected the range of molecular diversity. Future work should focus on scaling the model with larger, experimentally annotated datasets, incorporating reinforcement learning for bioactivity optimization, and validating top candidates through molecular docking and wet-lab screening.

In conclusion, this research establishes a proof-of-concept for integrating deep generative models with graph-based learning in fragment-based drug discovery. The proposed framework successfully generated chemically valid, novel, and drug-like molecules and enabled the computational prioritization of compounds with favourable predicted interactions with the P2X7 receptor. However, these findings are based solely on in silico analyses and do not establish therapeutic efficacy. Further experimental validation is required to determine the biological activity, safety, and clinical relevance of the generated molecules.”]

 

Comment 10: [References – Several formatting inconsistencies were noted in the reference list and should be corrected according to the journal guidelines.]

 

Response 10: [Thank you for the comment. The references has been review and doi were also added.]

Reviewer 2 Report

Comments and Suggestions for Authors

The study employs deep learning to generate novel molecules for Lung Cancer

Therapeutics.

The study is interesting and well conducted. Some comments are the following.

1.In Introduction and related work there is adequate (or maybe too much ) information on the deep learning techniques but very little is devoted to the target P2X7R receptor.

2.Which criteria were used to select the 9 compounds of the seed datasets.

3.Figure 1 is not clear to me. It presents examples of novel compounds? Structure 4-6 are not unique or even valid.

4.Both Figures 2 – 5 and Figures 13-16 show the distribution of counts, logP QED, It is not clear to me which molecules were included in Figure 2-5 and which in Figures 13-16

5.Which crystal structure of P2X7R receptor. Was used from PDB. Was it in complex with a substrate. More information on molecular docking should be given. A Figure with the best docked compound should be given.

  1. Section 3.3.3. The authors state ….. further confirming the strong predicted oral bioavailability . This is not exactly correct. Compliance with the rule of 5 is desirable but it does not guarantee oral availability.

 

Minor comments: Some grammar errors or typos should be checked. e.g. novelette instead of novelty

Comments on the Quality of English Language

Some grammar errors and typos should be checked

Author Response

Comments  1.In Introduction and related work there is adequate (or maybe too much ) information on the deep learning techniques but very little is devoted to the target P2X7R receptor.

Response 1 :[`Thank you for this suggestion. More literature has been added for the P2X7R in the second paragraph of the introduction. Lines 31-46]

 

Comments 2: [Which criteria were used to select the 9 compounds of the seed datasets.]

Response 2: [Thank you for the comment. This has been addressed with the paragraph “A dataset D of nine chemically diverse molecules, which has been reported as P2X7R modulators and studied in literature [4,5 ,8– 10 ,39 – 41 ], was obtained from ChEMBL and DrugBank. The drug molecules with their respective representation from SMILES to molecular graph are presented in Figure 13. Lines 301-305]

 

Comments 3: [is not clear to me. It presents examples of novel compounds? Structure 4-6 are not unique or even valid.]

Response 3: [I appreciate the comment. Figure 1 shows a randomly selected set of molecules generated by the model. The molecules were considered novel if they were no duplicate and they were validated for chemical correctness using RDKit.]

Comments 4: [Both Figures 2 – 5 and Figures 13-16 show the distribution of counts, logP QED, It is not clear to me which molecules were included in Figure 2-5 and which in Figures 13-16]

Response 4:[Thank you for the comment. This has been addressed “Figures 2–5 present the novel generated molecules and Figures 15–18 present results from 5,000 randomly selected SMILES, which have been renamed for clarity.”]

Comments 5:[ Which crystal structure of P2X7R receptor. Was used from PDB. Was it in complex with a substrate. More information on molecular docking should be given. A Figure with the best docked compound should be given.]

Response 5: [I appreciate the suggestion. More information on the P2x7R and the docking has been added in 4.10. Molecular Docking. Figures of the two best docked compounds has been added. Lines 500-512]

 

Comments 6: [Section 3.3.3. The authors state ….. further confirming the strong predicted oral bioavailability . This is not exactly correct. Compliance with the rule of 5 is desirable but it does not guarantee oral availability.]

Response 6 :[Thank you for this valuable comment. I added the paragraph “Overall, the physicochemical profiles suggest that the dataset represents a collection of drug-like small molecules with favourable characteristics for oral favourable properties for oral drug development; however, Lipinski’s Rule of Five compliance is only indicative and does not guarantee oral bioavailability.” Line 189-192]

Minor comments: Some grammar errors or typos should be checked. e.g. novelette instead of novelty

Response:[Thank you for the suggestion. This was checked and fixed]

Comments on the Quality of English Language Some grammar errors and typos should be checked

 

Response: [Thank you for this suggestion. Grammar errors and typos were checked and fixed]

Round 2

Reviewer 2 Report

Comments and Suggestions for Authors

Although the authors have addressed most comments  and the manuscript has been improved they did not provide satisfactory answer to comment 3. Structures 4-6 are not single molecules. They are two or three molecules/fragments . Structure 4 are two molecfules, Structure 5 are two molecules and a third wrong molecule (CH4- does not exist ) Structure 6 are three molecules. 

Also in Figure 8 compound A are two molecules. 

Often models generate wrong structures. The authors should check with some person proficient in chemistry.

In my opinion the manuscript cannot be published.

Comments on the Quality of English Language

Some grammar errors and typos should be checked

Author Response

Comments 1: Although the authors have addressed most comments  and the manuscript has been improved they did not provide satisfactory answer to comment 3. Structures 4-6 are not single molecules. They are two or three molecules/fragments . Structure 4 are two molecfules, Structure 5 are two molecules and a third wrong molecule (CH4- does not exist ) Structure 6 are three molecules. 

Also in Figure 8 compound A are two molecules. 

Often models generate wrong structures. The authors should check with some person proficient in chemistry.

In my opinion the manuscript cannot be published.

 

Response 2: Thank you for your valuable comment. To resolve structures, we did rerun the model at 500 epoch instead of 300 previously which generated better results. The generated molecules were then filter and we kept molecules with valid structure. Therefore, we had to rerun most of the analyses due to the new generated molecules. The molecular docking was done using PyRx as we had issues with the previous docking software schrödinger used. The new docking software produced better docking results for both novel and seed molecules as we did the docking on the entire protein P2X7R structure. All figures were reproduced with valid molecular structures.

Round 3

Reviewer 2 Report

Comments and Suggestions for Authors

The authors have addressed the issue with the chemical structures. So, now the manuscript could be published. Some typos and grammatical errors should still be checked.

Comments on the Quality of English Language

Some grammar errors and typos should be checked

Author Response

Comment 1: Some grammar errors and typos should be checked

Response 1: I would like to thank you for your valuable comment. Grammar and typos have been checked and resolved.

Back to TopTop