Next Article in Journal
Neurofunctional and Clinical Effects of Intranasal Human Recombinant Nerve Growth Factor in Children with Acquired Brain Injury
Previous Article in Journal
Combined Lung Immune Prognostic Index (LIPI)-Glasgow Prognostic Score (GPS) as a Prognostic Tool in Extensive-Stage Small-Cell Lung Cancer Treated with First-Line Chemo-Immunotherapy
Previous Article in Special Issue
Thiazole as a Promising Scaffold for the Treatment of Schistosomiasis: In Vitro and In Vivo Activity Against Different Developmental Stages of Schistosoma mansoni
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Machine Learning-Based QSAR Models for Discovery of Inhibitors Targeting Leishmania infantum Amastigotes

by
Naivi Flores-Balmaseda
1,†,
Julio A. Rojas-Vargas
2,†,
Susana Rojas-Socarrás
1,
Facundo Pérez-Giménez
3,
Francisco Torrens
4 and
Juan A. Castillo-Garit
5,*
1
Unit of Computer-Aided Molecular ‘‘Biosilico” Discovery and Bioinformatic Research (CAMD-BIR Unit), Departamento de Farmacia, Facultad de Química-Farmacia, Universidad Central ‘‘Marta Abreu” de Las Villas, Santa Clara 54830, Cuba
2
Doctorado en Informática Aplicada a Salud y Medio Ambiente, Universidad Tecnológica Metropolitana, Ignacio Valdivieso 2409, San Joaquín, Santiago 8940577, Chile
3
Unidad de Investigación de Diseño de Fármacos y Conectividad Molecular, Departamento de Química Física, Facultad de Farmacia, Universitat de València, 46100 Valencia, Spain
4
Institut Universitari de Ciència Molecular, Universitat de València, Edifici d’Instituts de Paterna, P.O. Box 22085, E-46071 Valencia, Spain
5
Instituto Universitario de Investigación y Desarrollo Tecnológico (IDT), Universidad Tecnológica Metropolitana, Ignacio Valdivieso 2409, San Joaquín, Santiago 8940577, Chile
*
Author to whom correspondence should be addressed.
These authors contributed equally to this work.
Pharmaceuticals 2026, 19(4), 588; https://doi.org/10.3390/ph19040588
Submission received: 24 February 2026 / Revised: 1 April 2026 / Accepted: 4 April 2026 / Published: 7 April 2026
(This article belongs to the Special Issue Advances in Antiparasitic Drug Research)

Abstract

Background/Objectives: Leishmaniasis is a group of diseases caused by obligate intracellular parasites of the Leishmania genus and is classified by the World Health Organization as a category I neglected tropical disease. Leishmania infantum predominantly affects children under five years of age and shows an increasing incidence of cutaneous and visceral forms. The development of new therapeutic alternatives remains challenging, making in silico approaches valuable for accelerating antileishmanial drug discovery. This study aimed to identify new compounds with potential activity against Leishmania infantum amastigotes using artificial intelligence-based classification models. Methods: A curated database of compounds with reported biological activity was constructed. Molecular representation employed zero- to two-dimensional descriptors calculated with Dragon software (v 7.0.10). Unsupervised k-means cluster analysis was applied to define training and external prediction sets. Supervised models were developed on the WEKA platform using IBk, J48, multilayer perceptron, and sequential minimal optimization algorithms. Model performance was assessed through internal cross-validation and external validation procedures. Results: All models achieved classification accuracies above eighty percent for both training and prediction sets, indicating consistent predictive performance and good generalization ability. The validated models were applied to virtual screening of the DrugBank database and a collection of synthetic compounds. This screening campaign enabled the identification of one hundred twenty compounds with potential activity against the amastigote form of Leishmania infantum. Conclusions: Artificial intelligence-based QSAR models proved to be useful tools for prioritizing antileishmanial candidates. The integration of molecular descriptors, machine learning, and virtual screening offers an efficient strategy for drug discovery.

Graphical Abstract

1. Introduction

Parasites of the genus Leishmania are intracellular protozoa belonging to the family Trypanosomatidae and are responsible for a group of complex clinical manifestations collectively known as leishmaniasis. This disease is characterized by severe alterations of the host immune system and presents a wide spectrum of clinical outcomes, ranging from self-healing cutaneous lesions to fatal visceral infections [1,2]. According to the World Health Organization, leishmaniasis is classified as a neglected tropical disease (NTD), predominantly affecting tropical and subtropical regions and impacting more than one billion people annually [3]. Its designation as “neglected” arises from the limited attention it has historically received, as it mainly affects low-income populations in developing countries with restricted access to adequate healthcare systems [4]. In recent years, leishmaniasis has gained increased relevance due to a marked rise in its global incidence and geographical expansion, currently being endemic in 98 countries across five continents, with approximately 350 million people living in areas at risk of infection [3,5].
The genus Leishmania comprises 53 recognized species, of which 29 are distributed in the Old World (southern Europe, Africa, the Middle East, Central Asia, and the Indian subcontinent), 20 in the New World (Central and South America), three are present in both regions, and only one species has been reported in Australia [6]. Among these, 31 species are pathogenic to mammals, and 20 are directly responsible for the diverse clinical forms of leishmaniasis in humans [6,7]. The parasite exhibits two main morphological forms during its life cycle: the extracellular, flagellated promastigote, predominant in the insect vector, and the intracellular, non-flagellated amastigote, which proliferates within vertebrate host cells [8,9]. Transmission occurs through the bite of infected female hematophagous dipterans of the genus Phlebotomus in the Old World and Lutzomyia in the New World [10,11]. Due to its zoonotic nature, leishmaniasis affects both humans and domestic animals, particularly dogs, which constitute the principal urban reservoir, alongside several wild mammalian species [10]. The complex biological cycle of Leishmania, involving multiple parasite species, vectors, and reservoirs, gives rise to a broad range of clinical syndromes affecting the skin, mucosae, and internal organs [12]. In the Americas, the disease presents with high prevalence and wide distribution, further exacerbated by social, economic, and environmental risk factors that significantly increase population vulnerability [13].
Among the pathogenic species, Leishmania infantum is of particular importance due to its presence in both the Old and New Worlds and its role as the etiological agent of zoonotic visceral leishmaniasis [12,14]. This species primarily infects macrophages in hematopoietic organs such as the bone marrow, spleen, and liver, leading to severe organ dysfunction [11]. In humans, infection by L. infantum may result in clinical presentations ranging from cutaneous leishmaniasis to the highly lethal visceral form. Although leishmaniasis was traditionally considered a rural disease, it is now well established in urban environments, where domestic dogs act as the main reservoir [13]. The discovery of new antileishmanial compounds remains a major scientific challenge. The diversity of Leishmania species, the parasite’s complex life cycle, and the heterogeneity of clinical manifestations hinder the identification of effective lead compounds [15]. Moreover, current therapeutic options are limited by high toxicity, severe adverse reactions, prolonged treatment regimens, and the emergence of drug resistance [16,17]. These limitations are further aggravated by an increasing parasite burden, the risk of HIV co-infection, and reduced drug responsiveness due to suboptimal dosing strategies [18].
Although no autochthonous cases of leishmaniasis have been reported in Cuba, the disease is of considerable national interest due to its global epidemiological relevance, the country’s geographical position, and the internationalist principles of its public health system [19]. Additionally, recent scientific and technological advances have introduced novel approaches for addressing unresolved aspects of leishmaniasis management [20]. In this context, there is an urgent need to discover new molecular entities with antileishmanial activity that exhibit improved therapeutic indices, reduced toxicity, and lower susceptibility to resistance [16]. To date, no compounds fully meeting these criteria have been identified, and traditional drug discovery approaches based on trial-and-error remain time-consuming and resource-intensive [15].
Computer-aided drug design (CADD) offers a faster and more cost-effective alternative for the discovery and optimization of bioactive compounds [21]. In silico methodologies encompass a wide range of computational techniques used to support drug discovery, with quantitative structure–activity/property relationship (QSAR/QSPR) modeling playing a central role [22,23]. These approaches have significantly contributed to the development of several drugs currently available on the market [24]. At the Faculty of Chemistry and Pharmacy of the Universidad Central “Marta Abreu” de Las Villas, the Computer-Aided Molecular Design and Bioinformatics Research Group (CAMD-BIR Unit) has reported relevant advances in the application of computational and graph-theoretical QSAR methodologies for the rational design of potentially bioactive organic compounds [25,26,27,28,29].
In a previous study, we developed predictive models for antileishmanial activity against Leishmania amazonensis using curated datasets and computational approaches [30]. However, the present study focuses on Leishmania infantum, a parasite species with important biological and clinical differences. While L. amazonensis is mainly associated with cutaneous and diffuse cutaneous leishmaniasis, L. infantum is one of the principal causative agents of visceral leishmaniasis, a systemic and potentially fatal disease that affects internal organs such as the spleen, liver, and bone marrow [5,31,32]. These differences imply distinct pharmacological contexts, including differences in therapeutic strategies, drug distribution requirements, and efficacy endpoints. Importantly, several studies have shown that susceptibility to antileishmanial drugs varies significantly among Leishmania species, with substantial interspecies differences in response to pentavalent antimonials, amphotericin B, and other compounds [15,16]. Such variability is associated with species-specific biological traits, and molecular mechanisms of drug response and resistance, including differential gene expression patterns linked to antimonial resistance in L. amazonensis [33]. Consequently, structure–activity relationships derived from datasets generated against one Leishmania species cannot be assumed to be directly transferable to another. In addition, the chemical space explored in the present dataset differs from that used in the previous study, as it includes compounds experimentally evaluated against L. infantum amastigotes, which may lead to different distributions of biological activity and relevant molecular descriptors. Consequently, the resulting structure–activity relationships may differ between datasets derived from different parasites [15,16]. Therefore, it is necessary to develop species-specific predictive models in order to capture the pharmacological and molecular determinants associated with activity against L. infantum, and to identify compounds with potential efficacy against visceral leishmaniasis, which remains one of the most severe and potentially fatal forms of the disease.
Based on the aforementioned considerations, the scientific problem addressed in this study is the limited efficacy and safety of current drugs used to treat leishmaniasis caused by Leishmania infantum, which necessitates the discovery of new active molecular entities. We hypothesize that it is possible to identify novel compounds with potential antileishmanial activity against L. infantum through the application of QSAR methodologies combined with artificial intelligence techniques [22,34]. Accordingly, the main objective of this work is to identify new antileishmanial compounds using in silico approaches.

2. Results and Discussion

2.1. Data Management, Curation, and Chemical Space Coverage

The construction of a robust and predictive QSAR model critically depends on the quality, consistency, and representativeness of the underlying dataset. In this study, an initial collection of 934 compounds was retrieved from PubChem bioassays reporting activity against Leishmania species. Following a rigorous curation process aimed at minimizing experimental noise and annotation inconsistencies, a final dataset of 437 compounds active against Leishmania infantum amastigotes was obtained, comprising 202 active and 235 inactive molecules. Activity classification was primarily based on IC50 values while also considering reported mechanisms of action and structural patterns, thereby reducing the risk of mislabeling compounds tested under heterogeneous experimental conditions.
The curated dataset exhibits a high degree of chemical and structural diversity, including multiple heterocyclic frameworks, aromatic and polyaromatic systems, aliphatic and cycloaliphatic motifs, steroid-like scaffolds, salicylic-acid derivatives, aromatic diamidines, substituted pyrimidines, and mixed saturated/unsaturated ring systems. The distribution of compounds in these major chemical classes is illustrated in Figure 1. As shown in the figure, the dataset is not dominated by a single type of scaffold but instead includes multiple structural classes, which supports the chemical heterogeneity of the compound collection. Such diversity is essential for developing QSAR models with broad applicability domains, particularly in neglected disease research, where chemical space exploration remains limited. The structural heterogeneity observed here provides a solid foundation for machine learning-based classification and reduces the likelihood of model bias toward narrow scaffold families. This structural diversity is further supported by the PCA-based chemical space analysis (Figure 2), which shows that the compounds occupy a large region of the descriptor space used for the development of the model.
Molecular representation was achieved through the calculation of 2489 zero- to two-dimensional molecular descriptors using DRAGON software [35]. These descriptors encompass constitutional properties, substructure fragment counts, functional group frequencies, topological indices, and connectivity-related parameters, collectively capturing both global and local molecular features. To mitigate multicollinearity and reduce the risk of overfitting, a correlation filter (|r| > 0.9) was applied, resulting in a reduced set of 1187 non-redundant descriptors. This dimensionality reduction step is particularly important for classification problems involving moderately sized datasets, as excessive descriptor redundancy can artificially inflate model performance while degrading generalizability.

2.2. Diversity-Preserving Dataset Splitting Strategy

To ensure that the predictive performance of the models reflects genuine generalization rather than favorable data partitioning, a diversity-driven splitting strategy was implemented. Instead of using random or purely stratified splits, cluster analysis was performed separately for active and inactive compounds using k-means cluster analysis (k-MCA) implemented in STATISTICA 8.0 [36], following previously established recommendations [37,38]. This approach ensures that each subset contains representative chemical diversity from the full dataset, thereby minimizing the risk of “easy splits” where structurally similar compounds are disproportionately grouped.
As illustrated in Figure 3, the clustering-based strategy resulted in 286 compounds assigned to the training set, 107 compounds to the prediction set, and 44 compounds to an external validation set. This design strengthens the methodological rigor of the study, as it forces the models to learn structure–activity relationships that are transferable across different regions of chemical space. From a QSAR validation standpoint, this step represents a key strength of the work and directly addresses common reviewer concerns regarding dataset bias and overoptimistic performance estimates.

2.3. Descriptor Selection and Algorithm-Specific Model Optimization

Given the heterogeneity of the descriptor space and the fundamentally different learning principles underlying each classification algorithm, descriptor selection was carried out independently for each model using the WEKA platform. Multiple evaluators, filters, and search strategies were applied to identify compact yet informative subsets of descriptors tailored to each algorithm.
Starting from the initial pool of 1187 filtered descriptors, preliminary reduced subsets were obtained for each classifier. These initial subsets consisted of 18 descriptors for IBk, 13 for J48, 20 for MLP, and 33 for SVM. These subsets were subsequently refined to obtain the final descriptor sets used for model construction. This algorithm-specific feature selection strategy is particularly important in comparative QSAR studies, as enforcing a single descriptor subset across heterogeneous machine learning algorithms often leads to suboptimal model performance and potentially misleading comparisons.
The novelty of the present modeling work is further emphasized by the limited availability of comparable classification studies based on artificial intelligence targeting Leishmania species. In this context, the work of Flores-Balmaseda et al. (2015) [39] represents the primary benchmark for comparison. The present study extends upon previous efforts by incorporating a carefully curated dataset, multiple layers of validation, and consensus modeling strategies.
After selecting the descriptors, the hyperparameters of each machine learning algorithm were optimized using the WEKA platform. For each classifier, a batch exploration of multiple combinations of parameters within predefined ranges was performed. The model optimization was performed using only the training set. The external test set was not used in the model optimization process and was reserved exclusively for the final evaluation of predictive performance. The use of an independent external validation set follows the best practices recommended for the development of QSAR models and provides a robust assessment of the predictive performance of the models, in accordance with the validation principles proposed by the Organization for Economic Co-operation and Development (OECD) [40].
The final models were built using optimized hyperparameters specific to each algorithm. The IBk classifier was configured with k = 3 nearest neighbors using the Manhattan distance. The J48 decision tree was generated using a pruning confidence factor of 0.09 and a minimum of five instances per leaf. The MLP model consisted of a single hidden layer containing 13 neurons and was trained for 500 epochs with a learning rate of 1.0 and momentum of 0.8. The SVM model was implemented using the SMO (sequential minimal optimization) algorithm with a radial basis function (RBF) kernel, with cost (C) and gamma parameters optimized during model development. The structural characteristics and hyperparameters of the developed models are summarized in Table S2 (Supporting Information), while the complete WEKA command-line configurations used for model training and testing are detailed in Table S1.

2.4. Classification Model Performance on Training and Prediction Sets

2.4.1. IBk (k-Nearest Neighbors) Model

The IBk classifier demonstrated the strongest overall performance among the evaluated models. As reported in Table 1, the model achieved 88.81% accuracy on the training set and 85.05% accuracy on the prediction set, with ROC-AUC values of 0.955 and 0.863, respectively. Importantly, sensitivity and specificity remained well balanced across both sets, indicating that the model does not favor one class disproportionately.
The low false-positive rates observed further support the suitability of the IBk model for antileishmanial activity classification. The descriptors used in this model reveal that topological descriptors, atom-centered fragments, and 2D autocorrelation indices play a central role in defining the local similarity relationships exploited by the kNN algorithm. Compared to the reference model by Flores-Balmaseda et al. [39], the present IBk model exhibits improved training performance while maintaining competitive predictive accuracy, highlighting the benefits of dataset curation and descriptor optimization.

2.4.2. J48 Decision Tree Model

The J48 decision tree model achieved 87.76% training accuracy and 84.11% prediction accuracy, with ROC-AUC values exceeding 0.88 (Table 2). A particularly notable result is the high training specificity (91.30%), indicating a strong capacity to correctly identify inactive compounds. This property is advantageous in early-stage virtual screening, where reducing false positives can significantly lower experimental costs.
Descriptor analysis indicates that constitutional descriptors, connectivity indices, and functional group counts dominate the decision process, enabling the extraction of interpretable structure–activity rules. From a scientific perspective, this interpretability adds value beyond raw predictive performance, as it facilitates hypothesis generation and guides rational compound optimization.

2.4.3. MLP Neural Network Model

The MLP model displayed stable and consistent performance across training and prediction sets, with accuracies of 83.22% and 83.18%, respectively (Table 3). Although slightly less accurate than IBk and J48, the MLP achieved competitive ROC-AUC values and demonstrated a more conservative classification profile, characterized by higher specificity than sensitivity in the prediction set.
This conservative classification behavior is advantageous when the objective is to prioritize compounds with a higher confidence of true biological activity, particularly in early-stage screening, where minimizing false-positive selections is critical. The descriptor analysis suggests that the neural network captures complex nonlinear relationships among topological and information-based descriptors. Notably, the MLP model outperforms the corresponding benchmark reported by Flores-Balmaseda et al. [39], supporting the consistency of the present modeling framework.

2.4.4. SVM (Support Vector Machine) Model

The SVM model yielded the lowest overall accuracy among the evaluated classifiers (81.47% training; 78.50% prediction, Table 4) but clearly excelled in terms of specificity and false-positive minimization. Training and prediction specificities exceeded 91%, with false-positive rates below 6%.
These results highlight the precision-oriented nature of the SVM classifier, which is particularly valuable in contexts where experimental validation resources are limited. Descriptors used in this model indicate a strong reliance on connectivity indices and functional group frequency descriptors, consistent with the margin-based decision boundaries characteristic of SVMs. As summarized in Figure 4, the SVM model complements the higher-accuracy classifiers by providing a highly conservative screening filter.

2.5. External and Internal Validation of Model Robustness

The primary criterion for the acceptance or rejection of a classification model relies on its performance on the external prediction set, which reflects the model’s predictive performance on unseen compounds. As previously described, the performance of each developed machine learning (ML) model was systematically evaluated using both validation and external datasets. The comparative results obtained for all models are summarized in Table 5. External validation using an independent set of 44 compounds provided an additional assessment of the predictive behavior of the developed models. As reported in Table 5 and Figure 5, the MLP model achieved the highest external accuracy (81.82%), indicating a relatively balanced classification performance. In contrast, the SVM model achieved 100% specificity, correctly classifying all inactive compounds in the external set. However, this result should be interpreted with caution due to the relatively small size of the external dataset. Moreover, the high specificity was accompanied by a lower sensitivity (57.14%) and a moderate MCC (0.64), indicating that the model tends to favor inactive predictions and may generate a higher number of false negatives. Therefore, rather than indicating apparent high predictive performance, the observed performance suggests that the SVM model prioritizes minimizing false positives. This behavior may be advantageous in screening scenarios where avoiding false positives is critical, but it also highlights the importance of considering complementary models with more balanced sensitivity–specificity profiles. According to the OECD principles for the validation of QSAR models, model performance should be interpreted by considering multiple statistical metrics rather than a single indicator [40].
Internal validation was conducted using 10-fold cross-validation on the training set, following established best practices [41,42]. The results summarized in Table 6 and Table 7 demonstrate that all models maintain stable performance across folds, with no evidence of severe overfitting. Notably, the SVM model exhibits the smallest discrepancies between training and cross-validation metrics, leading the authors to identify it as the most robust and reproducible model overall, despite its lower raw accuracy.
The comparative performance of the four machine learning classifiers highlights distinct trade-offs between predictive accuracy, sensitivity, and specificity (Table 7). The IBk model exhibited the highest overall quality (Q) under cross-validation (88.81%), accompanied by a strong balance between sensitivity (86.92%) and specificity (88.28%), indicating a robust ability to correctly classify both active and inactive compounds. However, its external sensitivity decreased notably (76.92%), together with an increased false-positive rate (FPR = 23.08%), suggesting a reduced generalization capability when applied to unseen data. In contrast, the J48 decision tree demonstrated a more conservative behavior, reflected in a lower cross-validation Q value (78.67%) but a markedly improved specificity in the external set (91.30%) and a reduced FPR (6.41%), which is advantageous for minimizing false-positive predictions. The MLP model showed moderate and more homogeneous performance across validation schemes, although it presented the highest FPR during cross-validation (23.08%), indicating a tendency to overpredict active compounds. Finally, the SVM classifier achieved a favorable compromise between specificity and error control, yielding the lowest FPR in both cross-validation (7.05%) and external validation (3.85%), albeit at the expense of lower sensitivity, particularly in the external set (63.85%). Overall, these results emphasize that while IBk achieves the highest classification performance, SVM and J48 provide more stringent classification with reduced false positives, which may be preferable in virtual screening campaigns where reliability and experimental cost reduction are critical.

2.6. Virtual Screening and Consensus-Based Hit Prioritization

The validated models were subsequently applied to a virtual screening campaign involving 5128 compounds, including 4660 DrugBank molecules [43,44,45] and 468 synthetic compounds from collaborating laboratories. Individual model predictions varied substantially, as shown in Figure 6, reflecting the distinct decision strategies of each classifier.
To increase confidence in the predicted hits, a consensus modeling approach was adopted. As summarized in Figure 7, 1335 compounds were predicted active by at least one model, while 120 compounds were predicted active by all four models, including 115 DrugBank compounds and five synthetic candidates. This progressive reduction highlights the effectiveness of consensus modeling in prioritizing a manageable and high-confidence subset of candidates for experimental validation, substantially reducing the number of compounds that would otherwise require experimental screening.
From a translational standpoint, the identification of DrugBank compounds within the four-model consensus is particularly significant, as these molecules may benefit from existing pharmacokinetic and safety data, thereby accelerating downstream experimental and clinical evaluation.
An examination of the descriptors selected across the different machine learning models provides insight into the structural features associated with antileishmanial activity. Several models include Kier–Hall valence connectivity indices (e.g., X0Av, X1Av, X2v, and X3Av), which describe molecular size, branching patterns, and overall topological complexity. These descriptors are usually related to physicochemical properties that influence membrane permeability and the ability of compounds to interact with biological targets in protozoan parasites.
Information theory-based descriptors, such as SIC5 and CIC5, capture aspects of molecular symmetry and structural diversity, suggesting that the spatial distribution of atoms and substituents may influence biological activity. In addition, Burden matrix eigenvalues (e.g., SpMin4_Bh(m), SpMax1_Bh(m), SpMin1_Bh(v), and SpMin4_Bh(e)) encode electronic and steric properties derived from atomic masses, electronegativity, polarizability, and van der Waals volumes, which are often associated with electronic distribution and intermolecular interactions involved in ligand–target binding.
Fragment-based descriptors and functional group counts further highlight chemically meaningful features within the dataset. In particular, descriptors associated with tertiary amines, amidine-type functionalities, and halogenated fragments indicate that specific chemical groups may contribute to activity by modulating lipophilicity, polarity, hydrogen-bonding capacity, and electrostatic interactions with biological targets.
Interestingly, several of the functional groups captured by the selected descriptors are also present in compounds identified among the predicted hits during the virtual screening, including molecules containing amine-based functionalities, heteroaromatic systems, and highly polar phosphate-containing groups such as those present in bisphosphonates.
It is important to note that the predicted hits include compounds belonging to several pharmacologically relevant categories, such as approved drugs, bioactive molecules in the investigational phase, synthetic compounds, and endogenous metabolites. The presence of approved drugs among the predicted compounds highlights the pharmacological diversity of the screened dataset and supports the biological plausibility of the computational predictions.
Among the approved drugs identified within the consensus set are disulfiram (DB00822), an aldehyde dehydrogenase inhibitor used for the treatment of alcohol use disorder, and several nitrogen-containing bisphosphonates such as alendronate (DB00630) and pamidronic acid (DB00282), which are widely used in the treatment of osteoporosis and disorders associated with increased bone resorption. The identification of bisphosphonate derivatives is particularly noteworthy because this class of compounds has previously demonstrated antiparasitic activity. In particular, aromatic bisphosphonate derivatives have been reported to inhibit parasite replication in Trypanosoma, Leishmania, Toxoplasma, and Plasmodium, in some cases exhibiting IC50 values in the nanomolar to low micromolar range [46].
Among the predicted compounds, other approved or clinically used drugs have been identified, such as the antiviral agent foscarnet (DB00529), the antineoplastic drugs pipobroman (DB00236) and busulfan (DB01008), the farnesyltransferase inhibitor lonafarnib (DB06448), and the marine-derived anticancer compound trabectedin (DB05109). The presence of multiple drugs with well-characterized pharmacological profiles highlights the potential opportunities for drug repurposing in the discovery of antileishmanial drugs.
The dataset also includes several bioactive molecules in the research or experimental phase that are currently under clinical or preclinical evaluation. Examples include 1-oleoyl-2-palmitoylphosphatidylcholine (DB05456), which has been investigated as a therapeutic candidate for acute coronary syndromes, QS-21 (DB05400), an immunological adjuvant evaluated in several clinical trials, and bevirimat (DB06581), which has been explored as a maturation inhibitor for HIV treatment. The presence of these molecules further illustrates the chemical and pharmacological diversity detected through computational screening.
In addition to the DrugBank-derived molecules, five synthetic compounds from collaborating laboratories were also predicted as active by all four models. One of these molecules contains a halogenated heteroaromatic core bearing two tert-butyl-substituted phenyl groups, a scaffold that combines high lipophilicity with electron-withdrawing substituents that may improve membrane permeability and protein binding. The remaining compounds share a triarylmethane-based scaffold functionalized with dimethylamino-substituted aromatic rings and a glycosyl moiety.
Triarylmethane derivatives are well known for their antimicrobial and antiparasitic properties, while the presence of dimethylamino aromatic groups may facilitate interactions with biological targets [47,48]. In addition, glycosylation and acetylation patterns may modulate physicochemical properties such as solubility, lipophilicity, and cellular uptake [49]. The identification of these compounds by all four models highlights their potential as promising candidates for further experimental evaluation.
Furthermore, the set of predicted hits contains small molecules with diverse structures and heterocyclic scaffolds commonly found in antimicrobial and antiparasitic drug discovery libraries. This structural diversity suggests that the models can identify compounds spanning multiple regions of biologically relevant chemical space, an important requirement for the discovery of structurally diverse antileishmanial candidates. Taken together, these descriptors suggest that antileishmanial activity in the analyzed dataset is influenced by a combination of molecular topology, electronic properties, and pharmacologically relevant functional groups.
From a translational perspective, it is important to recognize that drug discovery is inherently a multi-stage process in which computational predictions, in vitro assays, and in vivo studies provide complementary information. While experimental validation is essential, early-stage computational prioritization plays a critical role in reducing the number of candidate compounds and guiding subsequent biological evaluation. In this context, the present framework enables the identification of a reduced set of high-confidence candidates from large chemical libraries, thereby supporting more efficient downstream experimental efforts.
Therefore, the compounds prioritized in this study represent testable hypotheses that can guide future experimental validation efforts aimed at identifying novel antileishmanial agents.
Taken together, these results demonstrate that the integration of diversity-aware data splitting, algorithm-specific descriptor selection, multi-level validation, and consensus modeling yields a consistent and scientifically sound QSAR framework capable of supporting the identification and prioritization of novel antileishmanial candidates. The complementary strengths of the evaluated classifiers—accuracy-driven methods (IBk and J48), a balanced neural network model (MLP), and a margin-based support vector machine (SMO)—justify their combined use in virtual screening pipelines and provide a rational basis for prioritizing compounds for experimental testing.

3. Materials and Methods

3.1. Data Collection and Curation

All compounds included in this study, together with their biological activity profiles, were compiled from publicly available PubChem BioAssay records reporting experimental evaluations against Leishmania infantum amastigotes. These assays correspond to studies published in peer-reviewed, high-impact journals indexed in the Web of Science database, primarily between 1983 and 2016. When a compound appeared in more than one source, IC50 values derived from the most rigorous and clearly described experimental methodologies were selected. Additional metadata were extracted, including parasite life stage, assay conditions, reported IC50 values, molecular weight, and chemical structures derived from canonical SMILES representations. Compounds were classified as active when IC50 ≤ 1.5 μM and inactive otherwise, following thresholds commonly applied in antileishmanial screening studies [43,45,50].

3.2. Chemical Structure Representation and Curation

A total of 437 chemical structures were included in the dataset. Most structures were retrieved directly from PubChem [45,50], while the remaining compounds were manually drawn using ChemDraw Ultra 8.0 [51] and exported as *.sdf as MDL files. All structures were subsequently curated and standardized using Open Babel 2.4.1 [52], including the addition of explicit hydrogen atoms, charge neutralization, and removal of salts or counterions that could affect descriptor calculation. Duplicate structures were identified and removed using ISIDA [53] and EdiSDF [54], ensuring a non-redundant chemical dataset suitable for quantitative modeling.

3.3. Molecular Descriptor Calculation

Molecular descriptors encoding structural, physicochemical, and topological information were calculated using DRAGON Professional software v.7.0.10 [35]. Descriptor families spanning 0D to 3D representations were generated, resulting in an initial pool of 1 187 descriptors. Constant, near-constant descriptors and highly correlated variables (Pearson correlation coefficient > 0.90) were eliminated to reduce redundancy and minimize the risk of overfitting. Descriptor matrices were organized in spreadsheet format, with compounds represented as rows and descriptors as columns, supported by JChem for Excel for structure handling and visualization [23].

3.4. Statistical Analysis and Data Partitioning

To ensure adequate structural diversity and representative sampling, cluster analysis was performed separately for active and inactive compounds using k-means cluster analysis (k-MCA) implemented in STATISTICA 8.0 software [36]. Prior to clustering, all descriptor matrices were standardized. Compounds were randomly selected from each cluster to construct three non-overlapping subsets: a training set (65%), a prediction/test set (25%), and an external validation set (10%). This strategy ensured that chemical classes defined by clustering were proportionally represented across all subsets. Compounds included in the prediction and external validation sets were not used during model training [55,56].

3.5. Feature Selection and Machine Learning Modeling

QSAR classification models were developed using the WEKA 3.6 data mining platform [57]. As previously stated, the initial pool of calculated molecular descriptors was used as the set of input attributes after standard preprocessing steps, which included removing constant or near-constant descriptors and highly correlated variables. Feature selection was performed using the feature selection tools available in WEKA, combining filter-based and wrapper-based approaches. Multiple attribute evaluators and search strategies were explored for each learning algorithm to identify compact and informative subsets of descriptors that would maximize model performance. Four supervised learning techniques were evaluated: k-nearest neighbors (IBk) [58], decision trees (J48) [59], multilayer perceptron neural networks (MLPs) [60] and support vector machines (SVMs) [61]. It is important to note that the descriptor selection and model optimization procedures were performed using only the training set. The external prediction set was not used during descriptor selection or model development and was reserved exclusively for the final evaluation of the model’s predictive performance. This workflow follows the best practices recommended for the development and validation of QSAR models [30,62].

3.6. Model Validation and Performance Evaluation

Final models were selected according to the principle of parsimony, favoring those with the highest statistical significance and predictive performance using the smallest possible number of descriptors. A complete list of the molecular descriptors selected for each machine learning model is provided in the Supplementary Material (Table S3). Model performance was assessed using standard classification metrics, including overall accuracy (Q), sensitivity, specificity, false-positive rate (FPR), and the Matthews correlation coefficient (MCC). Internal validation was performed using 10-fold cross-validation, while predictive power was further assessed using both the independent prediction set and an external validation set composed of compounds not involved in model construction. These validation strategies are consistent with widely accepted recommendations for robust QSAR modeling [22,40,63].

3.7. Virtual Screening and Hit Identification

The DrugBank database was employed as the primary source for virtual screening. DrugBank integrates bioinformatic and chemoinformatic data for FDA-approved drugs, experimental compounds, and associated protein targets [64]. The database was screened using the validated QSAR models to identify compounds with potential activity against L. infantum amastigotes.
Additionally, a set of 468 synthetic compounds obtained from organic chemistry laboratories, 407 from the University of Rostock (Germany), and 61 from the Conservatoire National des Arts et Métiers (CNAM, France) was screened. The best identified hits were structurally analyzed and reported as promising candidates for future antileishmanial drug discovery efforts.

4. Conclusions

In this study, we developed and validated a set of machine learning-based QSAR models for the identification of compounds with potential activity against Leishmania infantum amastigotes. By integrating rigorous data curation, diversity-aware dataset splitting, algorithm-specific descriptor selection, multi-level validation, and consensus-based virtual screening, we established a robust and interpretable computational framework for the prioritization of antileishmanial candidates.
The application of this framework to a dataset of more than 5000 compounds enabled the identification of a reduced set of high-confidence candidates, including DrugBank molecules with known pharmacological profiles and structurally diverse synthetic compounds. The pharmacological interpretation of the predicted hits and the analysis of relevant molecular descriptors provide additional support for the biological plausibility of the results.
Importantly, the proposed approach is not intended to replace experimental validation, but to support early-stage drug discovery by focusing experimental efforts on the most promising candidates. In this context, the compounds identified in this study represent experimentally testable hypotheses that may contribute to the discovery of novel antileishmanial agents. Future work will focus on the experimental evaluation of these prioritized compounds, including in vitro assays against Leishmania infantum amastigotes, to further assess their biological relevance.
Overall, this work highlights the value of descriptor-based machine learning and consensus modeling as practical tools for antiparasitic drug discovery, particularly in scenarios where efficient prioritization strategies are essential to reduce experimental costs and accelerate the identification of candidate compounds.

Supplementary Materials

The following supporting information can be downloaded at https://www.mdpi.com/article/10.3390/ph19040588/s1, Figure S1: Representative consensus hits predicted as active by all four machine-learning models during the virtual screening; Table S1: Final machine learning model configurations implemented in WEKA 3.6; Table S2: Structural characteristics and main hyper parameters of the machine learning models used for QSAR classification; Table S3: Molecular descriptors used in each machine-learning model developed in this study; Excel File S1: curated database and predictions for each serie.

Author Contributions

Conceptualization, N.F.-B. and J.A.C.-G.; methodology, F.T.; software, F.P.-G.; validation, N.F.-B. and J.A.R.-V.; formal analysis, N.F.-B. and S.R.-S.; investigation, S.R.-S. and J.A.R.-V.; resources, F.P.-G.; data curation, S.R.-S.; writing—original draft preparation, N.F.-B. and J.A.R.-V.; writing—review and editing, F.T. and J.A.C.-G.; visualization, F.T.; supervision, F.P.-G.; project administration, J.A.C.-G.; funding acquisition, F.T. and J.A.C.-G. All authors have read and agreed to the published version of the manuscript.

Funding

This research was supported by the Competition for Research Regular Projects year 2023, code LPR23-11, Universidad Tecnologica Metropolitana, and also by the Competition for Research Assistant Funding UTEM, year 2024, code AI25-09, Universidad Tecnológica Metropolitana.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The original contributions presented in this study are included in the article/Supplementary Material.

Acknowledgments

Torrens, F. thanks the Universitat de Valencia for the Special Research Actions Funding 2024. J.A.C.-G. thanks the program ‘Estades Temporals per a Investigators Convidats’ for a fellowship to work at Valencia University in 2018.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Desjeux, P. Leishmaniasis: Current situation and new perspectives. Comp. Immunol. Microbiol. Infect. Dis. 2004, 27, 305–318. [Google Scholar] [CrossRef]
  2. Mann, S.; Frasca, K.; Scherrer, S.; Henao-Martínez, A.F.; Newman, S.; Ramanan, P.; Suarez, J.A. A Review of Leishmaniasis: Current Knowledge and Future Directions. Curr. Trop. Med. Rep. 2021, 8, 121–132. [Google Scholar] [CrossRef]
  3. WHO. Leishmaniasis—Fact Sheet. Available online: https://www.who.int/news-room/fact-sheets/detail/leishmaniasis (accessed on 19 February 2024).
  4. Ca, J.; Kumar, P.V.; Kandi, V.; N, G.; K, S.; Dharshini, D.; Batchu, S.V.C.; Bhanu, P. Neglected Tropical Diseases: A Comprehensive Review. Cureus 2024, 16, e53933. [Google Scholar] [CrossRef]
  5. Alvar, J.; Vélez, I.D.; Bern, C.; Herrero, M.; Desjeux, P.; Cano, J.; Jannin, J.; den Boer, M. Leishmaniasis worldwide and global estimates of its incidence. PLoS ONE 2012, 7, e35671. [Google Scholar] [CrossRef]
  6. Akhoundi, M.; Kuhls, K.; Cannet, A.; Votýpka, J.; Marty, P.; Delaunay, P.; Sereno, D. A Historical Overview of the Classification, Evolution, and Dispersion of Leishmania Parasites and Sandflies. PLoS Negl. Trop. Dis. 2016, 10, e0004349. [Google Scholar] [CrossRef] [PubMed]
  7. Ready, P.D. Epidemiology of visceral leishmaniasis. Clin. Epidemiol. 2014, 6, 147–154. [Google Scholar] [CrossRef] [PubMed]
  8. Bates, P.A. Transmission of Leishmania metacyclic promastigotes by phlebotomine sand flies. Int. J. Parasitol. 2007, 37, 1097–1106. [Google Scholar] [CrossRef]
  9. Kaye, P.; Scott, P. Leishmaniasis: Complexity at the host-pathogen interface. Nat. Rev. Microbiol. 2011, 9, 604–615. [Google Scholar] [CrossRef]
  10. Dantas-Torres, F.; Solano-Gallego, L.; Baneth, G.; Ribeiro, V.M.; de Paiva-Cavalcanti, M.; Otranto, D. Canine leishmaniosis in the Old and New Worlds: Unveiled similarities and differences. Trends Parasitol. 2012, 28, 531–538. [Google Scholar] [CrossRef] [PubMed]
  11. Killick-Kendrick, R. The biology and control of phlebotomine sand flies. Clin. Dermatol. 1999, 17, 279–289. [Google Scholar] [CrossRef]
  12. Chappuis, F.; Sundar, S.; Hailu, A.; Ghalib, H.; Rijal, S.; Peeling, R.W.; Alvar, J.; Boelaert, M. Visceral leishmaniasis: What are the needs for diagnosis, treatment and control? Nat. Rev. Microbiol. 2007, 5, 873–882. [Google Scholar] [CrossRef] [PubMed]
  13. Borges, M.S.; Niero, L.B.; da Rosa, L.D.S.; Citadini-Zanette, V.; Elias, G.A.; Amaral, P.A. Factors associated with the expansion of leishmaniasis in urban areas: A systematic and bibliometric review (1959–2021). J. Public Health Res. 2022, 11, 22799036221115775. [Google Scholar] [CrossRef] [PubMed]
  14. Rodrigues, A.V.; Valério-Bolas, A.; Alexandre-Pires, G.; Aires Pereira, M.; Nunes, T.; Ligeiro, D.; Pereira da Fonseca, I.; Santos-Gomes, G. Zoonotic Visceral Leishmaniasis: New Insights on Innate Immune Response by Blood Macrophages and Liver Kupffer Cells to Leishmania infantum Parasites. Biology 2022, 11, 100. [Google Scholar] [CrossRef] [PubMed]
  15. Croft, S.L.; Olliaro, P. Leishmaniasis chemotherapy—Challenges and opportunities. Clin. Microbiol. Infect. 2011, 17, 1478–1483. [Google Scholar] [CrossRef]
  16. Ponte-Sucre, A.; Gamarro, F.; Dujardin, J.C.; Barrett, M.P.; López-Vélez, R.; García-Hernández, R.; Pountain, A.W.; Mwenechanya, R.; Papadopoulou, B. Drug resistance and treatment failure in leishmaniasis: A 21st century challenge. PLoS Negl. Trop. Dis. 2017, 11, e0006052. [Google Scholar] [CrossRef]
  17. Sundar, S.; Chakravarty, J. Antimony Toxicity. Int. J. Environ. Res. Public Health 2010, 7, 4267–4277. [Google Scholar] [CrossRef]
  18. Alvar, J.; Cañavate, C.; Gutiérrez-Solar, B.; Jiménez, M.; Laguna, F.; López-Vélez, R.; Molina, R.; Moreno, J. Leishmania and human immunodeficiency virus coinfection: The first 10 years. Clin. Microbiol. Rev. 1997, 10, 298–319. [Google Scholar] [CrossRef]
  19. Scarpini, S.; Dondi, A.; Totaro, C.; Biagi, C.; Melchionda, F.; Zama, D.; Pierantoni, L.; Gennari, M.; Campagna, C.; Prete, A.; et al. Visceral Leishmaniasis: Epidemiology, Diagnosis, and Treatment Regimens in Different Geographical Areas with a Focus on Pediatrics. Microorganisms 2022, 10, 1887. [Google Scholar] [CrossRef]
  20. Ejazi, S.A.; Ali, N. Developments in diagnosis and treatment of visceral leishmaniasis during the last decade and future prospects. Expert Rev. Anti-Infect. Ther. 2013, 11, 79–98. [Google Scholar] [CrossRef]
  21. Lionta, E.; Spyrou, G.; Vassilatis, D.K.; Cournia, Z. Structure-based virtual screening for drug discovery: Principles, applications and recent advances. Curr. Top. Med. Chem. 2014, 14, 1923–1938. [Google Scholar] [CrossRef]
  22. Cherkasov, A.; Muratov, E.N.; Fourches, D.; Varnek, A.; Baskin, I.I.; Cronin, M.; Dearden, J.; Gramatica, P.; Martin, Y.C.; Todeschini, R.; et al. QSAR modeling: Where have you been? Where are you going to? J. Med. Chem. 2014, 57, 4977–5010. [Google Scholar] [CrossRef]
  23. Todeschini, R.; Consonni, V. Molecular Descriptors for Chemoinformatics; Wiley-VCH: Weinheim, Germany, 2009; I and II. [Google Scholar] [CrossRef]
  24. Lounkine, E.; Keiser, M.J.; Whitebread, S.; Mikhailov, D.; Hamon, J.; Jenkins, J.L.; Lavan, P.; Weber, E.; Doak, A.K.; Côté, S.; et al. Large-scale prediction and testing of drug activity on side-effect targets. Nature 2012, 486, 361–367. [Google Scholar] [CrossRef]
  25. Castillo-Garit, J.A.; Del Toro-Cortés, O.; Kouznetsov, V.V.; Puentes, C.O.; Romero Bohórquez, A.R.; Vega, M.C.; Rolón, M.; Escario, J.A.; Gómez-Barrio, A.; Marrero-Ponce, Y.; et al. Identification In Silico and In Vitro of Novel Trypanosomicidal Drug-Like Compounds. Chem. Biol. Drug Des. 2012, 80, 38–45. [Google Scholar] [CrossRef]
  26. Cañizares-Carmenate, Y.; Mena-Ulecia, K.; Perera-Sardiña, Y.; Torrens, F.; Castillo-Garit, J.A. An approach to identify new antihypertensive agents using Thermolysin as model: In silico study based on QSARINS and docking. Arab. J. Chem. 2019, 12, 4861–4877. [Google Scholar] [CrossRef]
  27. Castillo-Garit, J.A.; del Toro-Cortés, O.; Vega, M.C.; Rolón, M.; Rojas de Arias, A.; Casañola-Martin, G.M.; Escario, J.A.; Gómez-Barrio, A.; Marrero-Ponce, Y.; Torrens, F.; et al. Bond-based bilinear indices for computational discovery of novel trypanosomicidal drug-like compounds through virtual screening. Eur. J. Med. Chem. 2015, 96, 238–244. [Google Scholar] [CrossRef]
  28. Marrero-Ponce, Y.; Martínez-Albelo, E.; Casañola-Martín, G.; Castillo-Garit, J.; Echevería-Díaz, Y.; Zaldivar, V.; Tygat, J.; Borges, J.; García-Domenech, R.; Torrens, F.; et al. Bond-based linear indices of the non-stochastic and stochastic edge-adjacency matrix. 1. Theory and modeling of ChemPhys properties of organic molecules. Mol. Div. 2010, 14, 731–753. [Google Scholar] [CrossRef]
  29. Castillo-Garit, J.A.; Marrero-Ponce, Y.; Torrens, F.; García-Domenech, R.; Rodríguez-Borges, J.E. Applications of Bond-Based 3D-Chiral Quadratic Indices in QSAR Studies Related to Central Chirality Codification. QSAR Comb. Sci. 2009, 28, 1465–1477. [Google Scholar] [CrossRef]
  30. Castillo-Garit, J.A.; Flores-Balmaseda, N.; Álvarez, O.; Pham-The, H.; Pérez-Doñate, V.; Torrens, F.; Pérez-Giménez, F. Computational Identification of Chemical Compounds with Potential Activity against Leishmania amazonensis using Nonlinear Machine Learning Techniques. Curr. Top. Med. Chem. 2018, 18, 2347–2354. [Google Scholar] [CrossRef] [PubMed]
  31. Burza, S.; Croft, S.L.; Boelaert, M. Leishmaniasis. Lancet 2018, 392, 951–970. [Google Scholar] [CrossRef] [PubMed]
  32. Christensen, S.M.; Belew, A.T.; El-Sayed, N.M.; Tafuri, W.L.; Silveira, F.T.; Mosser, D.M. Host and parasite responses in human diffuse cutaneous leishmaniasis caused by L. amazonensis. PLoS Negl. Trop. Dis. 2019, 13, e0007152. [Google Scholar] [CrossRef]
  33. Patino, L.H.; Muskus, C.; Ramírez, J.D. Transcriptional responses of Leishmania (Leishmania) amazonensis in the presence of trivalent sodium stibogluconate. Parasites Vectors 2019, 12, 348. [Google Scholar] [CrossRef]
  34. Vamathevan, J.; Clark, D.; Czodrowski, P.; Dunham, I.; Ferran, E.; Lee, G.; Li, B.; Madabhushi, A.; Shah, P.; Spitzer, M.; et al. Applications of machine learning in drug discovery and development. Nat. Rev. Drug Discov. 2019, 18, 463–477. [Google Scholar] [CrossRef]
  35. Kode srl Dragon (Software for Molecular Descriptor Calculation) Version 7.0.10. 2017. Available online: https://chm.kode-solutions.net (accessed on 21 November 2021).
  36. StatSoft, Inc. STATISTICA (Data Analysis Software System); Version 8.0; StatSoft, Inc.: Tulsa, OK, USA, 2007; Available online: http://www.statsoft.com/ (accessed on 23 February 2026).
  37. Johnson, R.A.; Wichern, D.W. Applied Multivariate Statistical Analysis, 6th ed.; Pearson Prentice Hall: Upper Saddle River, NJ, USA, 2007; p. 773. [Google Scholar]
  38. Niazi, S.K.; Mariam, Z. Recent Advances in Machine-Learning-Based Chemoinformatics: A Comprehensive Review. Int. J. Mol. Sci. 2023, 24, 11488. [Google Scholar] [CrossRef] [PubMed]
  39. Flores-Balmaseda, N. Identificación de Nuevos Compuestos con Potencial Actividad Antileishmaniásica Mediante Estudios In Silico. Master’s Thesis, Universidad Central “Marta Abreu” de Las Villas, Santa Clara, Cuba, 2015. [Google Scholar]
  40. OECD. (Q)SAR Assessment Framework: Guidance for the Regulatory Assessment of (Quantitative) Structure Activity Relationship Models and Predictions, 2nd ed.; OECD Publishing: Paris, France, 2024. [Google Scholar]
  41. Victor Wandera, L.; Dennis, K.; Mary Lemasulani, M.; Njoka Grace, M.; Musyimi Daniel, K. Comparative Analysis of Cross-Validation Techniques: LOOCV, K-folds Cross-Validation, and Repeated K-folds Cross-Validation in Machine Learning Models. Am. J. Theor. Appl. Stat. 2024, 13, 127–137. [Google Scholar] [CrossRef]
  42. Qiu, J. An Analysis of Model Evaluation with Cross-Validation: Techniques, Applications, and Recent Advances. Adv. Econ. Manag. Political Sci. 2024, 99, 69–72. [Google Scholar] [CrossRef]
  43. Gaulton, A.; Bellis, L.J.; Bento, A.P.; Chambers, J.; Davies, M.; Hersey, A.; Light, Y.; McGlinchey, S.; Michalovich, D.; Al-Lazikani, B.; et al. ChEMBL: A large-scale bioactivity database for drug discovery. Nucleic Acids Res. 2012, 40, D1100–D1107. [Google Scholar] [CrossRef]
  44. Liu, T.; Lin, Y.; Wen, X.; Jorissen, R.N.; Gilson, M.K. BindingDB: A web-accessible database of experimentally determined protein-ligand binding affinities. Nucleic Acids Res. 2007, 35, D198–D201. [Google Scholar] [CrossRef] [PubMed]
  45. Kim, S.; Thiessen, P.A.; Bolton, E.E.; Chen, J.; Fu, G.; Gindulyte, A.; Han, L.; He, J.; He, S.; Shoemaker, B.A.; et al. PubChem Substance and Compound databases. Nucleic Acids Res. 2016, 44, D1202–D1213. [Google Scholar] [CrossRef] [PubMed]
  46. Martin, M.B.; Grimley, J.S.; Lewis, J.C.; Heath, H.T.; Bailey, B.N.; Kendrick, H.; Yardley, V.; Caldera, A.; Lira, R.; Urbina, J.A.; et al. Bisphosphonates Inhibit the Growth of Trypanosoma brucei, Trypanosoma cruzi, Leishmania donovani, Toxoplasma gondii, and Plasmodium falciparum:  A Potential Route to Chemotherapy. J. Med. Chem. 2001, 44, 909–916. [Google Scholar] [CrossRef] [PubMed]
  47. Picconi, P.; Hind, C.; Jamshidi, S.; Nahar, K.; Clifford, M.; Wand, M.E.; Sutton, J.M.; Rahman, K.M. Triaryl Benzimidazoles as a New Class of Antibacterial Agents against Resistant Pathogenic Microorganisms. J. Med. Chem. 2017, 60, 6045–6059. [Google Scholar] [CrossRef]
  48. Ricco, C.; Abdmouleh, F.; Riccobono, C.; Guenineche, L.; Martin, F.; Goya-Jorge, E.; Lagarde, N.; Liagre, B.; Ali, M.B.; Ferroud, C.; et al. Pegylated triarylmethanes: Synthesis, antimicrobial activity, anti-proliferative behavior and in silico studies. Bioorg. Chem. 2020, 96, 103591. [Google Scholar] [CrossRef]
  49. Gangliang, H.; Meijiao, L.; Jinchuan, H.; Kunlin, H.; Hong, X. Glycosylation and Activities of Natural Products. Mini Rev. Med. Chem. 2016, 16, 1013–1016. [Google Scholar] [CrossRef]
  50. Kim, S.; Chen, J.; Cheng, T.; Gindulyte, A.; He, J.; He, S.; Li, Q.; Shoemaker, B.A.; Thiessen, P.A.; Yu, B.; et al. PubChem 2025 update. Nucleic Acids Res. 2025, 53, D1516–D1525. [Google Scholar] [CrossRef]
  51. Paul, N.M.; Yoder, R.J.; Callam, C.S. Incorporating Chemical Structure Drawing Software throughout the Organic Laboratory Curriculum. J. Chem. Educ. 2019, 96, 2638–2642. [Google Scholar] [CrossRef]
  52. O’Boyle, N.M.; Banck, M.; James, C.A.; Morley, C.; Vandermeersch, T.; Hutchison, G.R. Open Babel: An open chemical toolbox. J. Cheminf. 2011, 3, 33. [Google Scholar] [CrossRef]
  53. Alexandre, V.; Denis, F.; Dragos, H.; Olga, K.; Cedric, G.; Philippe, V.; Vitaly, S.e.; Frank, H.; Igor, V.T.; Gilles, M. ISIDA—Platform for Virtual Screening Based on Fragment and Pharmacophoric Descriptors. Curr. Comput.-Aided Drug Des. 2008, 4, 191–198. [Google Scholar] [CrossRef]
  54. McGuire, R.; Verhoeven, S.; Vass, M.; Vriend, G.; de Esch, I.J.P.; Lusher, S.J.; Leurs, R.; Ridder, L.; Kooistra, A.J.; Ritschel, T.; et al. 3D-e-Chem-VM: Structural Cheminformatics Research Infrastructure in a Freely Available Virtual Machine. J. Chem. Inf. Model. 2017, 57, 115–121. [Google Scholar] [CrossRef] [PubMed]
  55. Bouveyron, C.; Celeux, G.; Murphy, T.B.; Raftery, A.E. Bibliography. In Model-Based Clustering and Classification for Data Science: With Applications in R; Bouveyron, C., Celeux, G., Murphy, T.B., Raftery, A.E., Eds.; Cambridge Series in Statistical and Probabilistic Mathematics; Cambridge University Press: Cambridge, UK, 2019; pp. 386–414. [Google Scholar]
  56. Golbraikh, A.; Tropsha, A. Beware of q2! J. Mol. Graph. Model. 2002, 20, 269–276. [Google Scholar] [CrossRef]
  57. Mark, H.; Eibe, F.; Geoffrey, H.; Bernhard, P.; Peter, R.; Ian, H.W. The WEKA data mining software. ACM SIGKDD Explor. Newsl. 2009, 11, 10–18. [Google Scholar] [CrossRef]
  58. Cover, T.; Hart, P. Nearest neighbor pattern classification. IEEE Trans. Inf. Theory 1967, 13, 21–27. [Google Scholar] [CrossRef]
  59. Quinlan, J.R. Induction of decision trees. Mach. Learn. 1986, 1, 81–106. [Google Scholar] [CrossRef]
  60. Rumelhart, D.E.; Hinton, G.E.; Williams, R.J. Learning representations by back-propagating errors. Nature 1986, 323, 533–536. [Google Scholar] [CrossRef]
  61. Cortes, C.; Vapnik, V. Support-vector networks. Mach. Learn. 1995, 20, 273–297. [Google Scholar] [CrossRef]
  62. Casanola-Martin, G.M.; Le-Thi-Thu, H.; Marrero-Ponce, Y.; Castillo-Garit, J.A.; Torrens, F.; Perez-Gimenez, F.; Abad, C. Analysis of Proteasome Inhibition Prediction Using Atom-Based Quadratic Indices Enhanced by Machine Learning Classification Techniques. Lett. Drug Des. Discov. 2014, 11, 705–711. [Google Scholar] [CrossRef]
  63. Chicco, D.; Jurman, G. The Matthews correlation coefficient (MCC) should replace the ROC AUC as the standard metric for assessing binary classification. BioData Min. 2023, 16, 4. [Google Scholar] [CrossRef]
  64. Wishart, D.S.; Knox, C.; Guo, A.C.; Shrivastava, S.; Hassanali, M.; Stothard, P.; Chang, Z.; Woolsey, J. DrugBank: A comprehensive resource for in silico drug discovery and exploration. Nucleic Acids Res. 2006, 34, D668–D672. [Google Scholar] [CrossRef] [PubMed]
Figure 1. Distribution of compounds across major chemical classes in the dataset.
Figure 1. Distribution of compounds across major chemical classes in the dataset.
Pharmaceuticals 19 00588 g001
Figure 2. Chemical space coverage of the dataset visualized by principal component analysis (PCA).
Figure 2. Chemical space coverage of the dataset visualized by principal component analysis (PCA).
Pharmaceuticals 19 00588 g002
Figure 3. Workflow describing the algorithm used to design the training, prediction, and external validation sets based on k-means cluster analysis (k-MCA).
Figure 3. Workflow describing the algorithm used to design the training, prediction, and external validation sets based on k-means cluster analysis (k-MCA).
Pharmaceuticals 19 00588 g003
Figure 4. Comparison of correct classification percentages for training and prediction sets across all models.
Figure 4. Comparison of correct classification percentages for training and prediction sets across all models.
Pharmaceuticals 19 00588 g004
Figure 5. Comparison of correct classification percentages in the external validation set.
Figure 5. Comparison of correct classification percentages in the external validation set.
Pharmaceuticals 19 00588 g005
Figure 6. Number of predicted active and inactive compounds for each classification model in the virtual screening.
Figure 6. Number of predicted active and inactive compounds for each classification model in the virtual screening.
Pharmaceuticals 19 00588 g006
Figure 7. Distribution of compounds according to the number of machine learning models that predict them as active during the virtual screening process.
Figure 7. Distribution of compounds according to the number of machine learning models that predict them as active during the virtual screening process.
Pharmaceuticals 19 00588 g007
Table 1. Statistical parameters of the IBk model for training and prediction sets.
Table 1. Statistical parameters of the IBk model for training and prediction sets.
SetCQ (%)Sensitivity (%)Specificity (%)FPR (%)
Training0.7788.8186.9285.713.01
Prediction0.7085.0580.3973.337.84
C: Mathews correlation coefficient, Q: accuracy, FPR: false-positive rate.
Table 2. Statistical parameters for the J48 model (classification tree).
Table 2. Statistical parameters for the J48 model (classification tree).
SetCQ (%)Sensitivity (%)Specificity (%)FPR (%)
Training0.7687.7686.9285.713.01
Prediction0.6884.1180.3973.337.84
C: Mathews correlation coefficient, Q: accuracy, FPR: false-positive rate.
Table 3. Statistical parameters obtained for the MLP (multilayer perceptron) model.
Table 3. Statistical parameters obtained for the MLP (multilayer perceptron) model.
SetCQ (%)Sensitivity (%)Specificity (%)FPR (%)
Training0.7687.7686.9285.713.01
Prediction0.6884.1180.3973.337.84
C: Mathews correlation coefficient, Q: accuracy, FPR: false-positive rate.
Table 4. Statistical parameters obtained for the SVM (support vector machine) model.
Table 4. Statistical parameters obtained for the SVM (support vector machine) model.
SetCQ (%)Sensitivity (%)Specificity (%)FPR (%)
Training0.6581.4763.8593.263.85
Prediction0.5978.5060.7891.185.36
C: Mathews correlation coefficient, Q: accuracy, FPR: false-positive rate.
Table 5. External validation statistical parameters for the four classification models.
Table 5. External validation statistical parameters for the four classification models.
SetCQ (%)Sensitivity (%)Specificity (%)FPR (%)
IBk0.5979.55 71.4383.3313.04
J480.5477.2771.4378.9517.39
MLP0.6481.8271.4388.248.7
SVM0.6479.5557.141000
C: Mathews correlation coefficient, Q: accuracy, FPR: false-positive rate.
Table 6. Internal validation statistical parameters for the four classification models.
Table 6. Internal validation statistical parameters for the four classification models.
SetCQ (%)Sensitivity (%)Specificity (%)FPR (%)
IBk0.5476.9276.9273.5323.08
J480.5778.6770.7780.0014.74
MLP0.5075.3573.4472.3123.08
SVM0.5878.3260.7787.787.05
C: Mathews correlation coefficient, Q: accuracy, FPR: false-positive rate.
Table 7. Internal validation statistical parameters for the four classification models.
Table 7. Internal validation statistical parameters for the four classification models.
ModelsAccuracy (%)Sensitivity (%)Specificity (%)FPR (%)
IBKTraining set88.8186.9288.289.62
10-fold-cv76.9276.9273.5323.08
J48Training set87.7680.7791.306.41
10-fold-cv78.6770.7780.0014.74
MLPTraining set83.2279.2383.0613.46
10-fold-cv75.3573.4472.3123.08
SVMTraining set81.4763.8593.263.85
10-fold-cv78.3260.7787.787.05
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Flores-Balmaseda, N.; Rojas-Vargas, J.A.; Rojas-Socarrás, S.; Pérez-Giménez, F.; Torrens, F.; Castillo-Garit, J.A. Machine Learning-Based QSAR Models for Discovery of Inhibitors Targeting Leishmania infantum Amastigotes. Pharmaceuticals 2026, 19, 588. https://doi.org/10.3390/ph19040588

AMA Style

Flores-Balmaseda N, Rojas-Vargas JA, Rojas-Socarrás S, Pérez-Giménez F, Torrens F, Castillo-Garit JA. Machine Learning-Based QSAR Models for Discovery of Inhibitors Targeting Leishmania infantum Amastigotes. Pharmaceuticals. 2026; 19(4):588. https://doi.org/10.3390/ph19040588

Chicago/Turabian Style

Flores-Balmaseda, Naivi, Julio A. Rojas-Vargas, Susana Rojas-Socarrás, Facundo Pérez-Giménez, Francisco Torrens, and Juan A. Castillo-Garit. 2026. "Machine Learning-Based QSAR Models for Discovery of Inhibitors Targeting Leishmania infantum Amastigotes" Pharmaceuticals 19, no. 4: 588. https://doi.org/10.3390/ph19040588

APA Style

Flores-Balmaseda, N., Rojas-Vargas, J. A., Rojas-Socarrás, S., Pérez-Giménez, F., Torrens, F., & Castillo-Garit, J. A. (2026). Machine Learning-Based QSAR Models for Discovery of Inhibitors Targeting Leishmania infantum Amastigotes. Pharmaceuticals, 19(4), 588. https://doi.org/10.3390/ph19040588

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop