2.2. Model Selection and Validation
In the context of data-driven decision-making, the choice of molecular descriptors is a critical first step. Although 3D descriptors can offer detailed information, their calculation is time-consuming and can be prone to inaccuracies, making them less suitable for rapid large-scale analysis. The majority of literature on QSAR models for inhibitors focuses on 2D descriptors, which are simpler and more directly related to chemical constitution. The computational efficiency and greater interpretability of 2D descriptors are a significant advantage for decision-makers who need to quickly understand the structural features driving a prediction. For this reason, our study utilized 1613 2D molecular descriptors from the Mordred package version 1.2.0.
Quantum-chemical descriptors are inherently interpretable and can readily guide (pharmaceutical) chemists in rational compound design. Furthermore, although they are traditionally considered computationally expensive, recent advances in machine learning allow their accurate estimation at significantly reduced computational cost. Therefore, two quantum-chemical descriptors from a previous study of [
16] were added to the 1613 2D molecular descriptors from the Mordred package version 1.2.0. From the initial set of descriptors, 969 descriptors remained after excluding those with missing or non-numerical values.
The choice of regression model plays a crucial role in data-driven decision-making. Although complex “black box” models can sometimes offer marginal improvements in predictive accuracy, interpretability often takes precedence when clear insights and transparency are essential. Models that reveal how descriptors influence activity are particularly valuable in guiding practical applications.
The case when OLS performs well suggests that a strong, simple linear relationship underlies the data, making the more interpretable model the preferred choice for this application. OLS coefficients directly indicate how each descriptor relates to activity, offering explainability by design. For example, a positive coefficient immediately signals that increasing the corresponding descriptor tends to increase predicted activity—providing drug designers with straightforward and actionable insights.
To generate the results, a genetic algorithm was applied to construct 500 models; however, none fulfilled the QUIK rule test or the selection criteria outlined in
Section 3.5. Similarly, none of the five models developed using SequentialFeatureSelector, RFE, and SelectKBest algorithms passed the QUIK rule test, suggesting the presence of multicollinearity in all cases. In contrast, models obtained through RFE, SequentialFeatureSelector with forward selection, and SelectKBest with the f_regression function (Equations (
A1), (
A2) and (
A3), respectively) satisfied all metric-based evaluation criteria (
Table A1 and
Table A2).
Our descriptor selection process demonstrated that the FeatureWiz algorithm followed by stepwise feature selection was the most effective method. This procedure not only streamlines the modeling process but also helps to mitigate multicollinearity, ensuring that the selected descriptors provide unique and meaningful insights. Neither
nor
was selected, suggesting that these descriptors may only be effective for the specific class of molecules investigated by [
16] or that they require a more precise, time-consuming method for calculation when applied to larger datasets.
We present three selected models. Model 1 (Equation (
1)) and Model 2 (Equation (
2)) were developed using a correlation threshold of
, while Model 3 (Equation (
3)) was based on a threshold of
. In particular, Model 2 incorporates 10 of the 14 descriptors included in Model 1, raising questions regarding the influence of the descriptor count on model performance.
The metrics for these models are summarized in
Table 1 and
Table 2. The consistently high
F-values indicate that the selected descriptors collectively explain the variation in activity beyond chance. All models met the predefined criteria for robustness and predictability.
All models listed in
Table 1 obey the QUIK rule, with no descriptors having a VIF higher than 5, indicating minimal multicollinearity. Moreover, the low values of coefficients of determination for the training set,
, and leave-one-out cross-validation,
, calculated in Y-scrambling suggest the absence of chance correlation, further validating the reliability of the models for predicting inhibitory activity against the SARS-CoV-2 main protease.
The results demonstrate overall comparability across all models. Model 2 shows only slight variation compared to Model 1 (making it a good alternative), while Model 3 exhibits more significant differences, particularly in parameters related to prediction accuracy. These variations could impact model performance during external validation. It is worth noting that the models which we decided not to focus on and thus represented in the
Appendix A by Equations (
A1)–(
A3) also exhibit slightly worse performance compared to Model 1 (
Table A1 and
Table A2 in the Experimental Section of
Appendix A). Good performances in cross-validation experiments, as well as similar performances of the model when trained on the complete dataset, can indicate good stability of a model.
To thoroughly investigate the performances of Models 1–3, scatter plots comparing experimental and predicted activity values for the training and test sets (
Figure 2) were analyzed, along with the Williams plots (
Figure 3). Williams plots are primarily designed to define and visualize the applicability domain of a QSAR/ML model, rather than to directly assess the quality of a train/test split. In principle, they may indicate a problematic split in cases where the training set compounds are clustered in a narrow region around the centroid of the descriptor space, while test set compounds occupy more extreme regions (i.e., show high leverage values). Since leverage reflects the distance from the centroid of the training set in descriptor space, a predominance of high-leverage test compounds could suggest structural imbalance and potential extrapolation. In our case, both training and test compounds occupy the same descriptor space region and are well within the defined applicability domain. The leverage values do not indicate structural extremity of the test compounds relative to the training set. Furthermore, the Williams plot is inherently constructed based on the selected training set and descriptor combination, which naturally centers the training compounds in the defined space. Taken together, the leverage analysis does not suggest structural imbalance between the training and test sets. This conclusion is further supported by the PCA analysis, which shows comparable distribution patterns for both sets in the descriptor space and will be presented separately. Thus, it is evident that all models demonstrate strong fitness and predictability. Models 1 and 2 show similar results, showing more accurate predictions for molecules with higher
values compared to Model 3 (
Figure 2). Since one of the goals of the drug discovery process is to develop drugs with high
values, this represents a particular advantage.
Additionally, one molecule shows a slightly higher positive standard deviation (
Figure 3c), although it does not surpass the critical value of 3 (i.e., no Y-outliers are detected). Moreover, all compounds in both the training and test set fall below the lowest threshold of
(for Model 2, which has the lowest number of descriptors), indicating the absence of response outliers and suggesting that predictions of inhibitory activity against the SARS-CoV-2 main protease could be extrapolated by all three models (
Figure 3).
To further ensure that the model does not have multicollinearity and overfitting issues and to confirm that all descriptors within the model are relevant, several other linear regression methods were performed. Even though other models are not strictly necessary to present a novel data-driven methodology for feature selection, their comparison with Models 1–3 strengthens its validation, as it demonstrates its advantages relative to alternative model configurations. Thus, we present
Table A3 and
Table A4 in the Experimental Section of
Appendix A. As can be seen, performance decreases in the following order: OLS → Ridge → LassoLars → Bayesian Ridge → ARD → Lasso → Linear SVR. Since the OLS method unequivocally gives the best results and Model 1 and Model 2 show better performance than Model 3, in the remainder of the study, the focus of this research will be on the results obtained by Model 1 and Model 2. The fact that OLS performed optimally suggests that a strong, simple linear relationship underlies the data, making the more interpretable model the preferred choice for this application.
Given the relatively small size of the CHEMBL database dataset, questions may arise regarding the model’s capability to predict
values for compounds with more structurally diverse profiles. To address this,
values from various QSAR studies in the literature [
6,
7,
8,
9,
10,
11,
12,
13,
14,
15,
16] were compiled to form an external dataset of molecules not present in the CHEMBL database. From the plot of the second versus the first principal component calculated on Model 1 descriptors, it can be observed that, similar to the previously discussed SlogP versus MW plot, molecules from the external dataset occupy a slightly broader chemical space than molecules from the CHEMBL database (
Figure 4a), justifying its use. As a side note, it is convenient to notice that PCA divides the datasets into two distinctive subsets. The smaller subset consists of sulfur compounds and aromatic ketones, while the remaining molecules are in the larger subset.
The ability of Model 1 and Model 2 to predict
values of molecules from the external dataset was tested in two different ways: by using original models (results marked as 1e and 2e in
Table 2) and using models trained on the full ChEMBL dataset (results marked as 1f and 2f in
Table 2). As can be seen, all metrics show comparable values to those obtained on the internal test set. This is an excellent result given that, due to the heterogeneous sources of the external set chemicals, the higher variability in the results should be expected. In
Figure 4b, a scatter plot comparing predicted versus experimental
values for the training and external sets shows that even the two molecules from the external set with higher
values than those in the training set are accurately predicted. Additionally, as observed from the Williams plot, no outliers are detected in the external set (
Figure 4c). Therefore, there is a robust foundation to assert that the proposed model is suitable for predicting inhibitory activity within a broad chemical space (see
Figure 5) and that consequently descriptor selection methodology can extract useful information from limited data. It is worth mentioning that, although
Figure 4b,c present results of original Model 1, plots obtained by Model 2, as well as adequate models trained on the full ChEMBL dataset, do not differ significantly.
Due to the non-uniform distribution of data points, with a dense representation within the 4.5 to 5.5 range and sparse representation outside of it, the models’ ability to predict
values in these two sparse regions was tested. The models demonstrated similar performance to those presented in
Table 1 and
Table 2, confirming that they are not biased towards the majority range.
2.3. Comparison with Models Found in the Literature
To the best of our knowledge, Authors in [
9] developed the only model existing in the literature using the dataset from the CHEMBL database. However, as shown in
Table 3 detailing the performances of models found in the literature, this model is characterized by unsatisfactory statistics in external validation. Similarly, the models developed by [
6,
7,
8,
12,
16] exhibit low values of
and/or
. Furthermore, the model of [
13] does not meet accuracy requirements due to a low value of
, while the model of [
10] has this shortcoming along with a low value of
and a high value of
. Additionally, the model of [
14] fails to meet the criterion
.
Therefore, among the models from the literature, only those by [
11,
15] are more closely aligned with the requirements of state-of-the-art QSAR modelling. However, ref. [
15] did not calculate all relevant statistics to ensure that. Additionally, comparing these models directly with this study is challenging for several reasons. While 2D descriptors were employed in this study, the aforementioned research utilized 3D descriptors, which are not only more computationally intensive—requiring precise atom coordinates—but also often less intuitive. As discussed, this may result in less precise insights into the design of drug candidates compared to 2D descriptors. In addition, both models were developed for specific classes of inhibitors—ketone-based covalent inhibitors [
11] and unsymmetrical aromatic disulfides [
15]. Nevertheless, since the methodology presented in this study demonstrates comparable internal fitting parameters and superior external validation performance compared to that of [
15], it can be inferred that it is better suited to predict the inhibitory activity of novel compounds. On the other side, while the model of [
11] demonstrates better performance in 6 out of 9 parameters as shown in
Table 3, once again, it is important to note that their model is specifically trained and tested on 29 derivatives of a certain compound, making its applicability limited to a narrow chemical space.
2.4. Mechanistic Interpretation
Mechanistic interpretation of QSAR models is an essential part of modelling because it provides insights into the biological or chemical processes underlying compound activity, thus enhancing scientific understanding and model validation. In other words, it ensures the model’s predictions are grounded in known principles, improving predictive power and reliability. In drug design, mechanistic insights guide the creation of effective and safe molecules, while also meeting regulatory requirements [
29]. To better understand the mechanisms interpretation of Models 1 and 2,
Table 4 provides the physical meaning of the corresponding descriptors. Additionally,
Table 5 shows six molecules from the CHEMBL dataset (three with low
values and three with high
values), along with the values of all descriptors. This comparison highlights the differences in descriptor values between molecules with varying activity levels, further clarifying the relationship between molecular structure and inhibitory activity.
By reflecting its sensitivity to differences in atomic numbers among atoms at a lag of 7, the AATS7Z descriptor shows a positive correlation with activity. It favours molecules containing heavy atoms (such as Br, S, I, and Cl) and aromatic rings. These features can increase inhibitor activity by enhancing hydrophobicity and polarizability, which promotes intermolecular interactions over intramolecular interactions. EState_VSA9 descriptor is part of the EState family of descriptors and serves to quantify a specific aspect of Van der Waals surface area contributions based on electrotopological state considerations. In analysed datasets, molecules with high values of EState_VSA9 typically contain chlorine atoms (
Table 5), which supports the frequent use of chlorine and other halogens as substituents in drugs, enhancing ligand–protein interactions via halogen bonds [
30].
Molecules with higher values of GATS5c exhibit charge distribution over longer distances, allowing for more efficient interactions with amino residues. In contrast, molecules with lower values of GATS5c have a dense distribution of charged atoms (
Table 5), resulting in lower hydrophobicity and reduced activity. Since oxygen and nitrogen atoms tend to have relatively low Gasteiger charges, molecules containing keto and ether groups, as well as nitrogen-containing rings, are favoured. Conversely, the last two molecules listed in
Table 5 contain two nitrogen atoms from a pyrimidine ring at a topological distance of 5 from nitrogen in a secondary amide group, resulting in a low value for this descriptor. GATS5c was also identified in a study on 2,5-disubstituted furans as antimalarial drugs [
31]. This adds to its importance, given that antimalarials have demonstrated in vitro efficacy against SARS-CoV-2.
The GATS2 descriptor has small values when nitrogen, sulfur, and oxygen atoms are at a topological distance of 2, i.e., for molecules containing (thio)amide groups and multiple heteroatom-containing rings. Similarly, the SssNH descriptor summarizes electrotopological states of nitrogen atoms in the form of –NH–, thus having higher values for secondary amides, which are known to exhibit low inhibitor activity [
7,
8,
10].
Selection of n10FaRing is not surprising, given that molecules with large surface area, hydrophobic terminal groups, and stable aromatic substituents often show higher activities [
9,
11]. NaaaC is included in the model developed by [
13], as well as in research on drugs with inhibitor activity against Alzheimer’s disease [
32]. This descriptor has been associated with various interactions such as H-bonding, salt bridges, alkyl groups, and
-sigma,
-cation, and
-alkyl interactions.
naHRing shows a weak positive correlation with
values. While this might be surprising for any descriptor with integer values, naHRing has demonstrated particular significance in studies of activin receptor type-5 kinase inhibitors [
33] and cannot be omitted from Models 1 and 2 without significantly affecting performance. Additionally, the majority of research highlights the significance of heteroatomic rings [
6,
10,
14].
The descriptor StsC is related to the existence of a carbon atom with one triple and one single bond, i.e., cyano groups. According to findings by [
14], inhibitor activity increases with the electron-donating ability of substituents. Therefore, since the cyano group is electron-withdrawing, the negative sign of the coefficient for StsC can be easily understood. Furthermore, NssssC has been selected, representing the number of quaternary carbon atoms. As said, it is known that branching decreases inhibitory activity [
6,
7,
8]. In addition, this descriptor is included in the study of [
8].
The PEOE_VSA10 descriptor, part of the partial equalization of orbital electronegativity descriptors, quantifies the Van der Waals surface area contribution from specific molecular segments. In other words, this descriptor reflects Van der Waals interactions associated with the partial charges of certain atoms or functional groups within the molecule. In analysed datasets, molecules featuring pyrimidine, thiazole, and pyrazole rings exhibit high PEOE_VSA10 values (
Table 5). As said, the literature suggests that molecules with multiple heteroatom-containing rings tend to show lower inhibitory activity [
10]. Notably, PEOE_VSA10 was selected in the study of HIV-1 protease inhibitors [
34], some of which, like lopinavir/ritonavir, were proposed for SARS-CoV-2 treatment [
2,
3]. Similarly, the PEOE_VSA6 descriptor shows high values in molecules containing pyrimidine rings, secondary amides, and aromatic ketones. As mentioned, secondary amides are known for their lower activity, while aromatic ketones’ low inhibitory activity may be due to their keto groups’ poor electron donation capability. Additionally, the EState_VSA4 and SMR_VSA6 descriptors are selected, representing the electrotopological state and molar refractivity contribution of specific atoms or functional groups to the Van der Waals surface area. In analysed datasets, high EState_VSA4 values are observed in molecules containing pyrimidine rings and secondary amides, while SMR_VSA6 values peak when a S atom is bonded to the C2 atom of a pyrimidine ring and in molecules with chlorine. Given the similarity of molecules with high values of these descriptors, a detailed analysis is warranted to understand better their nuances (which is currently out of the scope of this study). However, PEOE_VSA6 and EState_VSA4 are only part of Model 1, which, as discussed, exhibits slightly better performance than Model 2, highlighting that these descriptors might be affecting the performance of molecular activity prediction.
A closer analysis of the dataset and the mechanistic interpretation of descriptors suggest that molecules with high inhibitory activity against SARS-CoV-2 main protease generally lack cyano and secondary amide groups, as well as quaternary carbons (i.e., StsC, NssssC, and SssNH descriptors have values 0). Eight of the top ten most active molecules exhibit relatively high AATS7Z values (associated with heavy atoms), and four of the top five most potent inhibitors contain chlorine, contributing to high AATS7Z and EState_VSA9 values. To the best of our knowledge, our approach resulted in models that are the first QSAR models to highlight these significant features. Additionally, the models favor aromatic rings (as represented by descriptors like n10FaRing, naHRing, and NaaaC), especially those with heteroatoms. However, this is under specific constraints, particularly from descriptors such as GATS5c and GATS2are, along with PEOE_VSA10, PEOE_VSA6, SMR_VSA6, and EState_VSA4. In summary, molecules featuring larger aromatic structures, chlorine or other heavy atoms, and multiple nitrogen-containing rings appear to be prominent inhibitors. Since each highly active molecule in the dataset can be categorized by these characteristics, combining these structural elements may lead to the synthesis of even more effective SARS-CoV-2 main protease inhibitors in the future.