OPLE: Drug Discovery Platform Combining 2D Similarity with AI to Predict Off-Target Liabilities
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsThe off-target liabilities present significant challenges for successful drug development. It is essential to develop computer models that can quickly differentiate between successful and unsuccessful small molecule drug candidates. The authors of the manuscript have established a relationship between similarity and activity by using a probability assignment curve. Predictions based on the similarity to known active compounds were combined with the calibrated probability estimates of ML models using a probability theory formula. The approach proposed in the manuscript provides researchers an opportunity to predict safety liability likelihood based on molecular similarities and machine learning model predictions.
Two small questions:
In this manuscript compounds were classified as active if the IC50 values were less than or equal to 100 nM. But, the percentage of active and inactive molecules ranges from 3 to 55%%. How good is the choice of such a single threshold value for all targets? For example, all six models for the targets ALPHA1A (8% active, 92% inactive), ALPHA2A (6% active, 94% inactive), GABAA (2% active, 98% inactive) show significantly worse accuracy for (Figure 3).
What is the purpose of using DUD-E ("For each target, DUD-E decoys were either retrieved from or generated based on active com pounds using the DUD-E website. [21]"), especially for calibrating Tanimoto similarity, when the reference [21] states: "In the final decoy procedure, ECFP4 fingerprints were generated by Scitegic’s Pipeline Pilot for ligands and potential decoys. The decoys were sorted by their maximum Tc to any ligand, and the most dissimilar 25% were retained through this dissimilarity filter."
Author Response
Please see the attachment
Author Response File:
Author Response.docx
Reviewer 2 Report
Comments and Suggestions for AuthorsThe manuscript presents interesting and valuable work with clear potential for publication. The topic is relevant, the methodology is generally sound, and the results contribute meaningful insights to the field. However, several issues require clarification or correction to improve the scientific rigor and overall quality of the paper.
Please refer to the detailed comments provided in the attached commnts . Addressing these points will significantly enhance the manuscript's clarity, coherence, and methodological robustness.
I believe that with the appropriate revisions, the paper can be considered for acceptance.
-
Line 2–3: The title is somewhat long. Consider rephrasing for clarity: e.g., “Combining 2D similarity with AI to predict off-target liabilities for the drug discovery platform OPLE.”
-
Lines 5–7: Affiliations formatting is inconsistent. Standardize spacing and numbering; include ORCIDs if required.
-
Lines 8–31 (Abstract):
-
Sentences are long and reduce readability; consider breaking them up.
-
There is redundancy in terms like “off-target liabilities.”
-
Clarify the research gap and what limitations of current approaches OPLE addresses.
-
Include brief mention of ML methods and evaluation metrics for clarity.
-
“Belief theory formula” should be briefly explained.
-
Statements like “relationship between similarity and activity was established” are vague—quantify if possible.
-
Conclusion statements such as “potential to be a valuable early alert tool” should be supported with evidence.
-
Keywords could be expanded to include terms like “Tanimoto similarity,” “ECFP,” “safety pharmacology,” etc.
-
-
Lines 34–90 (Introduction):
-
Paragraphs are too long; break into shorter sections.
-
Some sentences are repetitive regarding drug discovery challenges and AI advantages.
-
Explicitly define the research gap and limitations of current methods.
-
Clarify terms like “off-target compound”—it is the activity, not the compound, that is the concern.
-
Informal expressions (“ML models are not omniscient”) should be revised.
-
Fenfluramine example is good but can be shortened; ensure consistent citations.
-
SafetyScreen panel types are listed without explanation; briefly clarify differences.
-
Explain why combining ECFP similarity with ML is novel.
-
Provide total dataset size including inactive compounds.
-
Early reporting of recall >0.8 should be moved to Results or Abstract.
-
-
Lines 99–124 (Section 2.1 – OPLE active library):
-
This reads more like Methods; consider moving procedural details.
-
Only active compounds are described; include inactive compounds.
-
Clarify why EMERALD “took precedence” and how duplicates/conflicts were handled.
-
Justify IC50 ≤ 100 nM as cutoff for actives.
-
Provide details on data curation: unit standardization, duplicates, and BindingDB integration.
-
Figures 1a and 1b need descriptive captions and summary statistics.
-
Address reproducibility for proprietary datasets.
-
-
Lines 133–154 (Section 2.2 – Global probability assignment curve):
-
Briefly explain Supplemental Figure 1 for clarity.
-
Clarify definition of similarity bins and fraction active.
-
Justify active/inactive/decoy pair thresholds and DUD-E decoy selection.
-
Explain why a global curve was used instead of per-target curves; provide quantitative support.
-
Clarify what “better agreement with Muchmore curve parameters” means.
-
Figure 2 should be summarized with key findings.
-
Consider moving methodological details to Methods; Results should focus on interpretation.
-
-
Lines 181–213 (Section 2.3 – ML probability calibration):
-
Clarify BECFP definition and how ML probabilities (BML) are combined.
-
Briefly explain Hooper’s rule and belief theory for readers unfamiliar with them.
-
Justify selecting top ML algorithm using MCC; mention other metrics if relevant.
-
Explain why XGBoost, Random Forest, and Gradient Boost performed best.
-
Provide justification for accuracy ≥75% and MCC ≥0.4 thresholds.
-
Support claim that poor performance was due to small datasets or class imbalance.
-
Clarify why ML outputs are not true probabilities and implications for belief combination.
-
Quantify dependence between BECFP and BML (e.g., Spearman ρ, mutual information).
-
Explain rationale for using RDKit descriptors instead of fingerprints.
-
Summarize key findings from calibration analyses (Supplemental Figure 3).
-
Correct abrupt sentence endings and ensure smooth paragraph transitions.
-
-
General (all sections):
-
Simplify long sentences to improve readability.
-
Ensure figures, tables, and supplementary materials are complete, labeled, and discussed.
-
Address minor typographical issues (e.g., hyphenation, line breaks).
-
Author Response
Please see the attachment.
Author Response File:
Author Response.docx
Reviewer 3 Report
Comments and Suggestions for AuthorsThe manuscript describes OPLE as a platform that predicts off-target liabilities by merging ECFP6 Tanimoto similarity, traditional machine learning classifiers, and belief theory. The authors build a global similarity–activity relationship, apply various probability calibration methods, and report high recall across external ChEMBL datasets. Although the topic is of utmost importance in safety pharmacology and cheminformatics, the manuscript is criticized for significant conceptual flaws, limited novelty, poor methodological choices, and overinterpretation of the data. These problems lead to the conclusion that the manuscript fails to satisfy the publication criteria of Pharmaceuticals.
- Fundamental conceptual flaw: Global ECFP similarity–activity curve usage. The authors assume one common universal probability curve for all >40 protein targets (pages 4–5). This is an invalid assumption from both biological and chemical perspectives: the steepness of SAR and the similarity–activity relationships are very different across GPCRs, ion channels, kinases, and transporters. Enforcing a single SC50 value (0.281) across all target families imposes artificial, unjustified homogeneity. The outcome is mathematically convenient but scientifically misleading. This by itself disapproves of the reliability of all OPLE predictions. The problem is severe and needs to be resolved before the paper is re-evaluated.
- Arbitrary and unsupported definition of “active” (IC50 ≤ 100 nM). The decision to label compounds with IC50 ≤ 100 nM as “active” (page 3) is not backed by any reasoning. This limit is not suitable for safety pharmacology: Numerous adverse effects due to off-target action that are of clinical importance take place at affinities in the micromolar range. The 100 nM threshold wrongfully cuts off the significant weak binders, thus affecting both model training and applicability domain negatively.
- The authors of the paper recognize the existence of an unbalanced dataset (p. 5) however they are silent on the following practices and techniques: class weights SMOTE or similar techniques random removal of a certain percentage of the majority class use of algorithms that are robust to imbalanced datasets This is a serious shortcoming in terms of contemporary ML practice and it also leads to higher than actual recall values being reported.
- On page 9, the authors propose a Tanimoto similarity threshold of 0.7 for applicability. Nevertheless, on page 5, it is stated that the dataset has an average similarity of only 0.08 ± 0.02. This inconsistency leads to the following conclusions: the model can be used only for an extremely small part of the chemical space, OPLE is not capable of supporting scaffold hopping or novel chemotypes. The authors must conduct a thorough applicability-domain analysis; at present, OPLE's usable range is immensely overstated.
- The evaluation concentrates mainly on recall (see Figure 3, p. 7). This is not the right way: High recall with non-controlled false-positive rates is not beneficial for safety prediction. The discussion regarding Precision, AUC, MCC per target, and calibration curves is far too brief. There are no confusion matrices, no error analysis, and no performance stratification by target family. This selective reporting skews the narrative towards OPLE.
-
Even with the description of the “novel methodology” (page 2), the manuscript is really nothing more than a reimplementation of Muchmore et al. (2008): Same similarity pairing logic, Same ECFP6 fingerprinting, Same belief theory equation, Similar fusion of independent evidence streams. It is very hard to say that the use of isotonic calibration instead of Platt scaling is a new approach. The paper seems to be more of a description of a proprietary internal tool than the presentation of a new scientific research.
To merit publication, the authors must:
1. Rebuild the similarity–activity curves on a per-target basis or scientifically justify global fusion with extensive benchmarking.
2. Redefine activity thresholds to reflect safety pharmacology reality possibly using multi-class potency bins or continuous regression.
3. Properly address data imbalance with modern ML techniques and report robustness.
4. Provide complete performance metrics including:
-
per-target ROC–AUC
-
per-target precision
-
calibration plots
-
confusion matrices
-
applicability-domain sensitivity analysis
5. Clarify novelty, explicitly distinguishing OPLE from Muchmore et al.
6. Include a prospective or retrospective real-world validation to demonstrate practical benefit.
7.Plagiarism detection:
-
Author Response
Please see the attachment.
Author Response File:
Author Response.docx
Reviewer 4 Report
Comments and Suggestions for AuthorsThe study addresses a relevant topic, but the concerns identified affect the scientific rigor and reliability of the conclusions. For this reason, I recommend a Major Revision, as substantial methodological clarification and restructuring are required before the manuscript can be considered for publication.
Abstract
Line 10 “blemish” is a non-scientific term; I suggest replacing it.
Line 12 “skyrocketed.” Again, it is informal. I suggest replacing it.
Lines 27–28: I consider it necessary to add information briefly: recall over which dataset, how many models, and which targets?
Introduction
Line 35–41 I consider that this information is basic and very general, present in most articles; it could be summarized to give space to other topics of importance in this article.
Line 42–43 off-target compound is incorrect. The off-target refers to the interaction, not the compound.
Line 44–52 I believe this would be the moment to talk a bit more about these systems and briefly describe the difference between ML and AI and how OPLE complements it.
Line 66–73 These are terms that the general reader may not know; it would be enriching for the manuscript to mention what they measure.
Line 90–96 Results should not be included in the introduction.
Results and Discussion
Line 101–107 Please explain how redundant or conflicting data are handled and indicate averages, best curve, minimum value, and exclusion criteria.
Line 107–111 The number of “successful” vs “unsuccessful” is not quantified, nor how those terms were defined. I recommend a table to explain it.
Line 112–116 The data come from laboratories, assay formats, and very heterogeneous conditions (BindingDB). There is no discussion of “batch effect” nor how this heterogeneity may impact the model.
Figure 1 Inactives are not shown nor the active/inactive ratio per target, therefore it is not possible to evaluate class imbalance. It also does not show any metric of chemical space coverage.
Line 150–153. This is an important methodological point, but there is no analysis showing that the global curve does not deteriorate performance in targets with a good amount of data.
Figure 2: Please report fitting metrics.
Line 183–191 The assumption of independence of beliefs is mentioned, but it is not rigorously demonstrated; I would expect to see average values of ρ, MI per target, and predefined thresholds.
Line 197–198 In my opinion, an in-depth discussion of problematic cases is missing, which are precisely the most important for designing improvements.
Figure 3: No statistical significance analysis is shown between methods. In this way it remains at a descriptive level and would be subjective.
Line 242–24:5 Was there a strong trade-off with precision? A high recall but with very low precision may cause the appearance of false positives.
Figure 4 and description, this block should not be included. It is 100% methodological, not a result nor a discussion.
Lines 282–288. This discussion is very important, but I consider it necessary to delve deeper into the limitations since there are several.
Materials and Methods
Lines 289–297 Please mention access criteria and type of assay since it is not possible to replicate the model without access to the source or without equivalent experimental standards.
Lines 299–300 DUD-E decoys are optimized for docking benchmarking, not for ML predictive off-target activity. This could introduce biases; please justify it.
Lines 336–349 It is important to describe the dataset size used for calibration vs that used for training; likewise, there is a risk that if the data used for calibration were not separated from the original training, the calibration may be overfit.
Author Response
Please see the attachment.
Author Response File:
Author Response.docx
Round 2
Reviewer 4 Report
Comments and Suggestions for AuthorsAfter the revision, the authors have addressed most of the observations and suggestions; they expanded the methods section, incorporated calibration analyses, and clarified assumptions. These changes result in a substantial improvement in the scientific quality of the article. Some areas for improvement remain (additional metrics, scaffold analysis), but overall, the manuscript is ready for publication.
