Next Article in Journal
Does DrugCLIP Find the Right Pocket? A Systematic Evaluation of Binding-Site Identification Across 42 Drug Targets
Previous Article in Journal
Active Learning on Protein Language Model Embeddings Accelerates Rubisco Variant Discovery for Desired Traits
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

A Hybrid Machine Learning and Quantum Mechanical Strategy for Predicting Radical Scavenging Potential

by
Davide Zeppilli
1,†,
José Ferraz-Caetano
2,†,
M. Natália D. S. Cordeiro
2,* and
Laura Orian
1,*
1
Dipartimento di Scienze Chimiche, Università di Padova, Via Marzolo 1, 35131 Padova, Italy
2
LAQV-REQUIMTE, Department of Chemistry and Biochemistry, Faculty of Sciences, University of Porto, Rua do Campo Alegre, S/N, 4169-007 Porto, Portugal
*
Authors to whom correspondence should be addressed.
These authors contributed equally to this work.
AI Chem. 2026, 1(2), 8; https://doi.org/10.3390/aichem1020008
Submission received: 2 February 2026 / Revised: 19 March 2026 / Accepted: 7 May 2026 / Published: 15 May 2026

Abstract

We designed a supervised machine learning framework to predict standard Gibbs free energies, ΔG°, of formal hydrogen atom transfer (f-HAT) for phenolic antioxidants across different radicals and media, enabling rapid and chemically interpretable screening. We curated a DFT dataset of 71 molecules (phenolic compounds and anthocyanidins), with 207 reaction sites, 10 radical reactive oxygen/sulfur species, and three environments (leading to a total of 6210 ΔG° values). The models amass 106 numerical RDKit descriptors, augmented with one-hot encodings of medium, site, radical, and structural class, and were evaluated through a leave-one-molecule-out protocol. Among the tested regression algorithms, the random forest regressor provides the best balance of accuracy and robustness with both R2 test (≈0.94) and MAE (2.74 kcal mol−1; RMSE (≈5.0 kcal mol−1)), close to DFT chemical accuracy. The feature-importance analysis revealed that “electronic” and “experimental” (site/group) descriptors primarily drive predictions, with the radical’s maximum absolute partial charge being the most important descriptor in the prediction of a radical’s ΔG°. These results suggest that descriptor-driven RF (Random Forest) models can generalize across chemical space to provide interpretable ΔG° predictions, providing a path for chemists towards a scalable route to prioritize antioxidant candidates for broader molecular families.

Graphical Abstract

1. Introduction

Antioxidants play a key role in fundamental biological defense mechanisms against overproduction and accumulation of pro-oxidant species, like reactive oxygen species (ROS) and free radicals [1,2,3,4,5]. This altered redox equilibrium is called oxidative stress, and it is present in several diseases, particularly those affecting the central nervous system (CNS), due to the associated high level of oxygen consumption [6,7,8,9,10,11]. In this context, natural substances and drugs are able to reduce these harmful species by acting as radical scavengers [12,13,14,15,16,17]. Indeed, the intake of exogenous antioxidants through the diet, combined with the action of endogenous systems, i.e., enzymes like glutathione peroxidase (GPx), helps restore and maintain the redox balance of homeostasis in the cells [18,19,20,21,22,23].
Among those, polyphenols and other phenolic compounds represent one of the most extensively investigated classes of radical scavengers due to their strong antioxidant properties [24,25,26,27,28,29]. These compounds are characterized by the presence of an aromatic scaffold with one or multiple reactive hydroxyl functions. Moreover, many other substituents may be present, but the OH group is identified as a highly reactive site to which the scavenging potential of such molecules is ascribed [30,31,32,33].
Radical scavenging activity is rooted in numerous reaction mechanisms, the role of which on the overall antioxidant activity depends on the molecular topology and charge of the antioxidant [34]. For the majority of compounds, a relevant scavenging path is formal hydrogen atom transfer (f-HAT), which involves the transfer of one proton and one electron from the antioxidant to the radical [34,35]. This latter is thus reduced, while, upon hydrogen atom loss from the former, a new radical product is formed, which is generally more stable, and it may be neutralized within safe metabolic pathways [16,36,37,38].
HnA + R → Hn−1A + RH
The thermodynamic feasibility of this process can be quantified by computing the Gibbs free reaction energy (ΔG°) for each reaction site of a specific scavenger [34,39,40]. Furthermore, multiple free radicals can be used to mimic different biologically relevant species. Finally, a comprehensive analysis also considers the influence of environment polarity, at a minimum comparing apolar and polar media, as they model distinct biological contexts. Therefore, a thorough investigation of the scavenging potential of a large number of molecules is lengthy and time-consuming, even when focusing only on one mechanism and neglecting kinetic considerations. This explains why most of the studies in the literature are dedicated to single natural or drug species and eventually include their most important metabolites [41,42,43,44,45,46,47,48]. Importantly, f-HAT usually follows the Bell–Evans–Polanyi principle [34]; thus, in general, the most thermodynamically favored processes are also the kinetically easiest, and the kinetic analysis can be omitted in the initial screening [49,50].
Besides the calculation of ΔG°, other computational strategies have been used by some authors to estimate the scavenging potential of molecules, exploiting relevant thermodynamic parameters, like the bond dissociation energy (BDE) [31,34,51,52,53,54,55]. BDEs provide a description of the donor-H bond strength, which reflects the ability of a free radical to abstract the H atom; thus, the higher the BDE, the less favored the f-HAT. Although BDEs may provide a reactivity scale of different reaction sites in molecules, a reference state is needed to assess the thermodynamic feasibility of any reaction. Moreover, no information on the radical partner is included; therefore, BDEs do not allow discrimination between exergonic and endergonic processes or the radical selectivity of the compound of interest. Lastly, different BDEs should be considered to account for solvation effects, since f-HAT reactivity is affected by solvent polarity.
To sum up, quantum mechanics (QM) calculations allow us to compute Δ G f H A T o from any reaction sites, with multiple radicals in different environments for each molecule [34]. Even though the efficiency of QM calculations is increasing, the complete description of the energetics requires a non-negligible amount of time, which exponentially increases with the number of molecules of interest and becomes challenging in large-scale analyses. Additionally, these types of calculations intrinsically require handling a huge amount of data, which is generated by including all the different variables (Figure 1).
To overcome these difficulties, an emerging strategy consists of screening a relatively big dataset to select the most reactive sites of a molecule or the most reactive molecules within an ensemble and then focusing the ΔG° calculation on the selected systems. In this context, machine learning (ML) approaches are a promising tool to speed up the process and reduce the computational cost. ML methodologies using regressor algorithms have been successfully applied to improve computational models for chemical property predictions [56,57,58,59], in general, but also in the scavenging potential predictions [60,61,62,63,64,65], thanks to the availability of a great amount of data.
While most approaches use RDKit-generated molecular fingerprints as ML model descriptors, sometimes the characteristics of the dataset demand a more diversified pool of features [66]. In this case, RDKit numerical descriptors, which are continuous features capturing physicochemical and topological properties of molecules, are particularly advantageous. They are suitable when interpretability is a central goal, handling datasets with few negative data [67]. Unlike binary fingerprints, descriptors such as polar surface area, logP, hydrogen-bond donor/acceptor counts, or molecular refractivity carry direct physicochemical meaning. Numerical features allow chemists to rationalize ML model predictions, adding chemistry-driven insights for model explainability [68]. For example, if the reactivity toward hydroxyl radicals depends on molecular polarity, descriptors related to electronegativity or polarizability can provide an explanation for this trend. By tackling the “black-box” mantra, RDKit descriptors also offer benefits when dealing with small to medium-sized datasets, where high-dimensional binary fingerprints may cause overfitting. This is relevant in cases where experimental datasets (either from wet or in silico chemistry) are often restricted by practical limitations.
In this work, we apply a regressive ML model to accurately predict the Δ G f H A T o of several phenolic molecules divided into two groups: anthocyanidins [69] and simple phenolic compounds [70]. All reaction sites (OH and CH) of anthocyanidins are considered, while only the OH sites of the simple phenolic compounds are included, due to the well-known higher reactivity of OH sites than CH sites in such compounds. Indeed, unsaturated CH sites perform the scavenging action via RAF (Radical Adduct Formation) instead of f-HAT, a mechanism which is out of the scope of the present work [34]. Two different environments are taken into account using a continuum solvation model, that is, besides the gas phase (reference in QM calculations), water (polar medium) and benzene (apolar medium), respectively. Furthermore, ten different radicals are selected from the family of ROSs and their sulfur analogs (RSS, reactive sulfur species): OH, OCH3, OOH, OOCH3, OOCH2=CH, SH, SCH3, SSH, SSCH3, OSCH3.

2. Materials and Methods

A database was built with a variety of molecules divided into two groups (anthocyanidins and phenolic compounds) with Δ G f H A T o values calculated through a Density Functional Theory (DFT) protocol by some of us [60,70] using Gaussian16 [71]. All molecules were optimized with M06-2X meta-hybrid functional [72] combined with the 6-31G(d) basis set. The most stable conformer was selected for each compound either using dedicated software (CREST 3.0) [73] or comparing a few structures based on chemical intuition (e.g., by maximizing the number of possible H bonds). To obtain Gibbs free energies, statistical mechanics formulas for ideal gases at 298 K and 1 atm were employed as implemented in Gaussian16, by conducting analytical frequency calculations at the same level of theory. Then, single-point calculations were carried out using the same functional combined with the extended 6-311+G(d,p) basis set in the gas phase, water, and benzene; solvation was included with the SMD model [74]. The overall level of theory is denoted (SMD)-M06-2X/6-311+G(d,p)//M06-2X/6-31G(d), which is consistent with state-of-the-art protocols to calculate scavenging potential of organic molecules [29,34,39,40,41,75,76,77,78].
The complete database is included in the article’s online repository (see Data Availability Statement), while the list of the molecules investigated is shown in Figures S1 and S2. Each reaction is associated with a standard Gibbs free energy (ΔG°), including a full range of molecular, radical, and reaction properties along with a selection of RDKit descriptors. These are the molecular properties: compound name (Molecule), linear structural notation in SMILES, reactive position in the molecule (Site), chemical family or subclass (Group), and net electrical charge (Charge). Additionally, the structural classification is documented at three hierarchical levels: the SMILES Superclass, which represents broad chemical classification (e.g., phenylpropanoids and polyketides); the SMILES_Class, which defines a more detailed classification of the chemical structure (e.g., flavonoids); and the SMILES_Parent_Level, which provides the molecular scaffold/subclass (e.g., 7-hydroxyflavonoids). The reaction properties convey the reaction medium (Medium, for example, gas phase), and the reactive radical species (Radical), which is denoted by its label and a distinctive radical SMILES notation (SMILES Radical). For each SMILES notation of the molecule and radical, a set of 106 RDKit descriptors is calculated. Chemical descriptors used to feed the model were generated with RDKit version 2025.03.4 [79], running on top of Python 3.12.7 [80]. Our set of descriptors is based on the method of a published ML model that successfully used them to predict ΔGsol [56]. The set amasses 106 features, divided according to their categorical groups: van der Waals surface area (VSA), electronic, structural, and experimental descriptors. The complete list of these descriptors is provided in the Supplementary Information (Table S1) [81]. The curated database comprises 71 molecules and 207 total reaction sites with ten radicals and three different media, for a total number of 6210 ΔG° values.
We started with a script that implements a supervised-learning workflow to predict reaction energetics and analyze descriptor importance under a leave-one-molecule-out evaluation scheme. Using the curated dataset, we performed categorical feature engineering by one-hot encoding the non-numeric variables (MEDIUM, SITE, RADICAL, GROUP, SMILES_Superclass, SMILES_Class, and SMILES_Parent_Level) using a OneHotEncoder. After excluding non-predictive identifiers, all features were standardized before we tested four regressors sequentially: ordinary least squares (LinearRegression), a random forest (RF), gradient boosting (GB), and an MLPRegressor (MLP). All models used scikit-learn (Version 1.7.2) default hyperparameters unless explicitly stated.
Model assessment follows a leave-one-molecule-out protocol over the unique molecule labels. For each molecule, the algorithm withholds all its rows as the test set and trains on the remaining 70 molecules. After model fitting, predictions for the held-out molecule are obtained and summarized through multiple metrics: the coefficient of determination on the training and test partition (R2), the root mean squared error (RMSE), the mean absolute error (MAE), and the standard deviation of the predicted values (STD). Per-sample diagnostics (true energy, predicted energy, and absolute error) were recorded for every test observation. Upon completion of the protocol, we exported fold-wise statistics with a long-format table of feature importances. To enrich the prediction ledger, the protocol merges the statistical results with selected metadata from the original dataset. Hyperparameter optimization was performed only for the best-performing mode, exploring different configurations to identify the setting that yields the best generalization. Hyperparameter selection was based on MAE as the primary criterion, presented in the Supplementary Information found in the article’s online repository.
M A E = i = 0 N | p r e d i c t i o n ( i ) c o m p u t e d ( i ) | N
R M S E = i = 1 N ( p r e d i c t i o n ( i ) c o m p u t e d ( i ) ) 2 N

3. Results and Discussion

We aggregated results at the molecule-level into nine molecule classes to analyze model performance. The aggregation decreases variance at the molecule-level, facilitates systematic evaluation of trends within chemically related families, and allows more straightforward evaluation on identifiable chemical classes instead of molecular structure. The distribution of each energy value per molecule class is presented in Figure 2.
Among the four tested algorithms, the linear regression yielded negative values for all classes and was thus discarded. For the remaining algorithms, we present the statistical results in Table 1. Although the three ML algorithms yielded similar R2 test values around 0.92–0.94, the RF yielded a lower MAE of 2.7 kcal mol−1. For this reason, we selected this algorithm for hyperparameter optimization. Its statistical metrics are also presented in Table 1 (the reported leave-one-molecule-out performance may be influenced by the fact that hyperparameters were selected using the full dataset).
To test whether the high performance of the leave-one-molecule-out protocol was solely due to the use of radical, medium, site, and structural class information, a purely categorical Random Forest was also trained, with one-hot encoding only of non-numeric variables (MEDIUM, SITE, RADICAL, GROUP, SMILES_Superclass, SMILES_Class, and SMILES_Parent_Level) and none of the RDKit descriptors or continuous molecular features. The purely categorical baseline resulted in R2 = 0.933 ± 0.002, MAE = 3.27 ± 0.05 kcal mol−1, RMSE = 5.50 ± 0.09 kcal mol−1 under the same protocol. This confirms that a significant proportion of the predictive structure of the data lies in the categorical labels. However, the optimal descriptor-based Random Forest is more accurate in terms of absolute prediction, lowering MAE from 3.27 to 2.74 kcal mol−1 and RMSE from 5.50 to just below 5.0 kcal mol−1. Thus, the entire descriptor set provides key additional predictive information to the non-numeric descriptor baseline.
Given that all leave-one-molecule-out folds are derived from within the same defined chemical domain, and because all folds have the identical values of the radical, medium, site, and structure-class variables, these LOMO scores represent an upper bound on out-of-molecule performance. Other molecules that may be removed from this domain, regarding their structure/mechanism, may therefore yield worse performance.
All four algorithms demonstrated a high degree of predictive accuracy. Both the MLP and gradient boosting algorithms perform well, but they present MAE meaningfully higher than Random Forest. For the optimized RF algorithm, hyperparameter optimization slightly decreased the RMSE while keeping the R2 test and MAE within similar margins of error. After this, we performed a controlled validation benchmark with a random split by rows (ignoring molecule grouping) to show what happens when the same molecule’s different sites appear in the train/test set. For the optimized algorithm, performance increased to R2 = 0.9960, MAE = 0.84 kcal mol−1. This implies that the presence of the same molecule in both the training set and testing set substantially simplifies the problem. Therefore, the leave-one-molecule-out protocol provides a stricter estimate of generalization to unseen molecules than a random row split. Although it should be viewed as an upper-bound estimate within the chemical domain here studied. The full statistical results are presented in the article’s online repository.
To bypass a potential overfitting issue, given the high leave-one-molecule-out performance, we also implemented a molecule count learning curve. Training on 20, 40, 60, and 70 molecules and evaluating on a set of unseen molecules, it resulted in pooled R2 test values of 0.862, 0.917, 0.933, and 0.943, respectively. The results indicate a trend towards better predictive performance when the number of molecules used for training increases (until reaching a stable high-performance zone). As we present the full statistical results in the article’s online repository, the scarcity of the remaining testing samples may be the reason why variance increases with large training molecule counts.
In Figure 3, we present the regression chart for the optimized Random Forest algorithm performance. It presents the correlation of predicted to true reaction energies for all molecules, colored by SMILES_Class. Most points lie around the diagonal, denoting good agreement between the model predictions and the reference value. While there is a somewhat notable spread for phenol energy values (−60 to 20 kcal mol−1) and flavonoids (in the region of 0 to 30 kcal mol−1), the model produces similar overall trends across all molecular families and did not consistently fail to capture the expected energy values. Overall, the rest of the molecules stay within the RMSE threshold line, revealing the degree of domain generalization of the model. To clarify the predictive accuracy in each class, Figure 4 presents the R2 test scores recorded for the optimized RF algorithm.
Prediction R2 test values per class show differences in prediction accuracy for different chemical classes (these results are presented in the article’s online repository). Most of the classes, such as naphthalenes (0.96), stilbenes (0.96), anthracenes (0.91), and organooxygen compounds (0.92), were modeled with very high accuracy. Flavonoids performed exceptionally well (0.98), while naphthalenes (0.80) and especially phenols (0.67) were only moderately accurate (as expected due to their distribution in Figure 3). These differences indicate that the optimized model generalizes well to unseen molecules within the studied domain, but not uniformly across all subclasses. However, certain structural classes remain problematic to predict due to the greater diversity corresponding to chemical structures. Particularly, the biggest overestimations shown in Figure 3 belong to a specific flavonoid (ARN, see below) and many different phenols; while underestimations are caused by some benzene and substituted derivatives and many phenols, as well. These deviations are mainly attributed to anionic species, whose effect is described below.
Flavonoids, the largest category, show good performance in relation to class sizes, indicating the model performs better with larger sample sizes. In contrast, in the case of phenols, which compose 17.4% of the dataset, accuracy is very moderate. Additionally, several smaller categories (e.g., stilbenes, naphthalenes, and organooxygen compounds) had strong predictive performance. Therefore, it appears that model accuracy may not be strictly attributable to the sample size of attributes used to train the model. Looking only at the errors (Figure 5), clear differences emerge between classes. The smallest errors are found in flavonoids, as small categories such as naphthalenes and stilbenes show highly consistent predictions despite their sample sizes. In contrast, naphthacenes and phenol derivatives (which together represent almost 20% of the dataset) yield the largest errors, confirming that more samples do not necessarily reduce error if the class is diverse. Heterogeneous groups remain challenging, whereas smaller but more uniform families generalize effectively under molecule leave-out validation.
RF algorithm provides the most promising results with MAE of 2.74 kcal mol−1 (chemical accuracy for DFT calculations is around 2 kcal mol−1). Particularly, the majority of ΔG° predictions are below this threshold value with two main exceptions. The first one is a specific OH site of the anthocyanidin ARN; this site, called O19, is present only in this specific molecule since none of the other eleven anthocyanidins contains a similar OH site (Figure 6). Therefore, the under-representation of this reaction site explains the higher error in predicting its ΔG°, especially in water. Indeed, high polarity tends to generate bigger energy differences, increasing the prediction error for under-represented systems. The second exception is broader, involving a considerable number of polyphenolic anions. This category contains the anions of simple phenolic compounds, whose pKa is compatible with deprotonation in physiological environments, and they mostly belong to the class with the highest errors, i.e., phenols, explaining the bad performances of this class. The presence of deprotonated oxygen allows the formation of an intramolecular H bond in the reactants as well as in the radical products, giving rise to a plethora of energy differences that are not well predicted by the ML model. This effect is also environmentally dependent, since the biggest errors are obtained in the gas phase and in benzene (i.e., low polarity). Indeed, H-bond effects are less impactful on the energy in polar environments, causing the prediction of better results for anions in water. To confirm the bad ΔG° predictions of such anionic systems, the same ML model was applied to a similar database after removing all the anions and their relative ΔG°. Indeed, focusing only on anthocyanidins and neutral phenolic compounds, the performance of the non-optimized model increases, and an MAE of 1.63 kcal mol−1 is computed.
After confirming the predictive performance of the optimized Random Forest algorithm, we evaluated which features drive model performance. Feature importance (FI) analysis provides insight into types and sources of descriptors the model uses, providing an understanding of why it generalizes across classes of molecules and where its predictive strength might originate from. The results for the FI of the optimized model are presented in Figure 7.
In Figure 7, electronic and experimental descriptors have the largest contributions to predictive performance, with strong roles for both molecule- and radical-level features (more than 60% of FI). Descriptors associated with experimental meaning are related to the reaction site or the group (anthocyanidins or phenolic compounds), and all provide substantial contributions. Structural and VSA-type descriptors have smaller contributions, amassing about 30% of FI. The importance of experimental and electronic information shows that the molecule type (with site labeling) and the intrinsic reactivity are fundamental factors in accurate energy predictions. These descriptors allow the model to capture general chemical patterns, which support its strong performance across different molecule classes. By decomposing contributions by source, we find that radical-level features are predominant overall, showing that the specific site properties are one of the main drivers of the best predictions. Molecular-level features are also significant drivers, especially in the experimental and electronic categories, highlighting the utility of locality in reactivity. Site-level features contribute through experimental descriptors, providing chemically grounded corrections with 20% of overall importance.
Particularly, the top 10 descriptors account for over 60% of the FI contribution and are reported in Figure 8. The main ones are the maximum absolute partial charge of the radical (SMILES_RADICAL_MaxAbsPartialCharge), the group of phenolic compounds, and HallKierAlpha. The most important feature refers to a radical property; indeed, the inclusion of a specific radical is significantly impactful in the overall reaction energy, since the radical represents one of two reactants. Although this piece of information is fundamental to accurately predict the desired ΔG°, the choice of a different radical causes a simple energy shift with no insight into the scavenging activity of a specific molecule. More insightful are the following two features, highlighting the importance of reference systems with similar moieties (group phenols) and the molecular structural fingerprint (HallKierAlpha). Therefore, energy predictions greatly depend on the overall structure of the molecule of interest, but the model also benefits from analogous molecules of the same family, suggesting the intrinsic limit of the model to predict the reactivity of non-phenolic molecules with the current employed database.
Overall, the analysis of feature importance demonstrates that the model’s predictive capacity is rooted in a balance of descriptors at different levels of molecular information. This balance provides the model with the ability to predict broad chemical trends while also capturing local effects of reactivity, which helps explain why the model was able to generalize so well to different classes of molecules.

4. Conclusions

In this work, we have applied a data-driven ML model to accurately predict Gibbs free reaction energies of f-HAT involving several phenolic compounds. Particularly, the direct reduction of different harmful free radicals has been evaluated, in order to estimate the scavenging potential of selected molecules from a thermodynamic point of view. The model was validated by predicting ΔG° associated with a database built with energy values calculated with a state-of-the-art DFT protocol. The RF algorithm performed with high accuracy and a 94% prediction test score, while the mean absolute error is 2.74 kcal mol−1, close to the chemical accuracy of DFT. MAE was further decreased by removing highly noisy data due to the presence of a plethora of possible intramolecular H bonds, which are not fully represented by the current employed database.
Overall, the optimized RF model has good predictive capabilities on leave-one-molecule-out validation. It can estimate reaction energies for molecules excluded from training, if they fall in the same overall chemical domain as the molecules represented in the training set. The lower accuracy observed for phenolic anions and rare structural motifs indicates that the model does not generalize uniformly across all subclasses, which is in favor of domain-specific molecule generalization.
Specific chemical descriptors are particularly significant to the model’s predictions. Features dealing with electronic properties, reaction sites, family group, and structural information are crucial for accurate predictions. Overall, the model can predict the scavenging potential of phenolic compounds, suggesting the possibility of extending the database to include other molecular families. These results pave the way for the application of ML models in molecular screening for antioxidant drug design by testing the thermodynamic feasibility of scavenging activity of unknown organic molecules.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/aichem1020008/s1, Figure S1: List of the molecules included in the anthocyanidins group, Figure S2: List of the molecules included in the phenolic compounds group, Table S1: List of RDKit Descriptors used in ML model development (n = 106 descriptors).

Author Contributions

D.Z. and J.F.-C.: writing—original draft preparation, writing—review and editing, visualization, validation, software, methodology, investigation, formal analysis, data curation, conceptualization. M.N.D.S.C. and L.O.: writing—review and editing, methodology, conceptualization, supervision. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the Università di Padova and Portuguese national funds through the FCT/MECI (Fundação para a Ciência e Tecnologia and Ministério da Educação, Ciência e Inovação) through the project UID/50006/2025 to LAQV-REQUIMTE (Laboratório Associado para a Química Verde—Tecnologias e Processos Limpos), DOI: https://doi.org/10.54499/UID/50006/2025.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

All datasets, code, and trained models supporting the findings of this study are openly available free of charge in an online GitHub repository. The complete workflow, including input data and intermediate results, can be accessed at the article’s online repository page (https://github.com/jfcaetano/ScavML) (accessed on 1 February 2026).

Acknowledgments

Davide Zeppilli and Laura Orian are grateful to CINECA for the generous allocation of computational resources (Project ISCRA C SIM-2, HP10C9J8UJ P.I. Laura Orian). José Ferraz-Caetano’s PhD fellowship is supported by a doctoral grant (SFRH/BD/151159/2021) financed by the Fundação para a Ciência e a Tecnologia (FCT/MECI) with funds from the Portuguese state and the European Union budget through the European Social Fund and Programa Por_Centro, under the MIT Portugal Program, DOI: https://doi.org/10.54499/SFRH/BD/151159/2021.

Conflicts of Interest

The authors declare no conflicts of interest. The funding sponsors had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript, and in the decision to publish the results.

Abbreviations

The following abbreviations are used in this manuscript:
BDEBond Dissociation Energy
CNSCentral Nervous System
DFTDensity Functional Theory
f-HATFormal Hydrogen Atom Transfer
FIFeature Importance
GBGradient Boosting
GPxGlutathione Peroxidases
HATHydrogen Atom Transfer
MAEMean Absolute Error
MLMachine Learning
MLPMultilayer Perceptron
QMQuantum Mechanics
RAFRadical Adduct Formation
RFRandom Forest
RMSERoot Mean Squared Error
ROSReactive Oxygen Species
STDStandard Deviation
VSAVan der Walls Surface Area

References

  1. Umeno, A.; Biju, V.; Yoshida, Y. In Vivo ROS Production and Use of Oxidative Stress-Derived Biomarkers to Detect the Onset of Diseases Such as Alzheimer’s Disease, Parkinson’s Disease, and Diabetes. Free Radic. Res. 2017, 51, 413–427. [Google Scholar] [CrossRef]
  2. Juan, C.A.; Pérez de la Lastra, J.M.; Plou, F.J.; Pérez-Lebeña, E. The Chemistry of Reactive Oxygen Species (ROS) Revisited: Outlining Their Role in Biological Macromolecules (DNA, Lipids and Proteins) and Induced Pathologies. Int. J. Mol. Sci. 2021, 22, 4642. [Google Scholar] [CrossRef]
  3. Nimse, S.B.; Pal, D. Free Radicals, Natural Antioxidants, and Their Reaction Mechanisms. RSC Adv. 2015, 5, 27986–28006. [Google Scholar] [CrossRef]
  4. Brewer, M.S. Natural Antioxidants: Sources, Compounds, Mechanisms of Action, and Potential Applications. Compr. Rev. Food Sci. Food Saf. 2011, 10, 221–247. [Google Scholar] [CrossRef]
  5. Silva, J.P.; Coutinho, O.P. Free Radicals in the Regulation of Damage and Cell Death—Basic Mechanisms and Prevention. Drug Discov. Ther. 2010, 4, 144–167. [Google Scholar] [PubMed]
  6. Reuter, S.; Gupta, S.C.; Chaturvedi, M.M.; Aggarwal, B.B. Oxidative Stress, Inflammation, and Cancer: How Are They Linked? Free Radic. Biol. Med. 2010, 49, 1603–1616. [Google Scholar] [CrossRef]
  7. Sies, H. On the History of Oxidative Stress: Concept and Some Aspects of Current Development. Curr. Opin. Toxicol. 2018, 7, 122–126. [Google Scholar] [CrossRef]
  8. Sies, H. Oxidative Stress: A Concept in Redox Biology and Medicine. Redox Biol. 2015, 4, 180–183. [Google Scholar] [CrossRef]
  9. Gingrich, J.A. Oxidative Stress Is the New Stress. Nat. Med. 2005, 11, 1281–1282. [Google Scholar] [CrossRef]
  10. Sies, H.; Berndt, C.; Jones, D.P. Oxidative Stress. Annu. Rev. Biochem. 2017, 86, 715–748. [Google Scholar] [CrossRef] [PubMed]
  11. Aldini, G.; Altomare, A.; Baron, G.; Vistoli, G.; Carini, M.; Borsani, L.; Sergio, F. N-Acetylcysteine as an Antioxidant and Disulphide Breaking Agent: The Reasons Why. Free Radic. Res. 2018, 52, 751–762. [Google Scholar] [CrossRef]
  12. Gassen, M.; Youdim, M.B.H. Free Radical Scavengers: Chemical Concepts and Clinical Relevance. In Diagnosis and Treatment of Parkinson’s Disease—State of the Art; Journal of Neural Transmission. Supplementum; Springer: Wien, Austria, 1999; Volume 56, pp. 193–210. [Google Scholar]
  13. Kruk, I.; Aboul-Enein, H.Y.; Michalska, T.; Lichszteld, K.; Kładna, A. Scavenging of Reactive Oxygen Species by the Plant Phenols Genistein and Oleuropein. Luminescence 2005, 20, 81–89. [Google Scholar] [CrossRef]
  14. Bors, W.; Saran, M. Radical Scavenging by Flavonoid Antioxidants. Free Radic. Res. Commun. 1987, 2, 289–294. [Google Scholar] [CrossRef]
  15. Bors, W.; Michel, C. Chemistry of the Antioxidant Effect of Polyphenols. Ann. N. Y. Acad. Sci. 2002, 957, 57–69. [Google Scholar] [CrossRef]
  16. Neha, K.; Haider, M.R.; Pathak, A.; Yar, M.S. Medicinal Prospects of Antioxidants: A Review. Eur. J. Med. Chem. 2019, 178, 687–704. [Google Scholar] [CrossRef]
  17. Ribaudo, G.; Bortoli, M.; Pavan, C.; Zagotto, G.; Orian, L. Antioxidant Potential of Psychotropic Drugs: From Clinical Evidence to In Vitro and In Vivo Assessment and Toward a New Challenge for In Silico Molecular Design. Antioxidants 2020, 9, 714. [Google Scholar] [CrossRef] [PubMed]
  18. Forman, H.J.; Davies, K.J.A.; Ursini, F. How Do Nutritional Antioxidants Really Work: Nucleophilic Tone and Para-Hormesis versus Free Radical Scavenging In Vivo. Free Radic. Biol. Med. 2014, 66, 24–35. [Google Scholar] [CrossRef]
  19. Maiorino, M.; Conrad, M.; Ursini, F. GPx4, Lipid Peroxidation, and Cell Death: Discoveries, Rediscoveries, and Open Issues. Antioxid. Redox Signal. 2018, 29, 61–74. [Google Scholar] [CrossRef] [PubMed]
  20. Flohé, L. Glutathione Peroxidase: Fact and Fiction. Ciba Found. Symp. 1978, 65, 95–122. [Google Scholar]
  21. Eleutherio, E.C.A.; Silva Magalhães, R.S.; de Araújo Brasil, A.; Monteiro Neto, J.R.; de Holanda Paranhos, L. SOD1, More than Just an Antioxidant. Arch. Biochem. Biophys. 2021, 697, 108701. [Google Scholar] [CrossRef]
  22. Deisseroth, A.; Dounce, A.L. Catalase: Physical and Chemical Properties, Mechanism of Catalysis, and Physiological Role. Physiol. Rev. 1970, 50, 319–375. [Google Scholar] [CrossRef] [PubMed]
  23. Madabeni, A.; Bortoli, M.; Nogara, P.A.; Ribaudo, G.; Dalla Tiezza, M.; Flohé, L.; Rocha, J.B.T.; Orian, L. 50 Years of Organoselenium Chemistry, Biochemistry and Reactivity: Mechanistic Understanding, Successful and Controversial Stories. Chem.—Eur. J. 2024, 30, e202403003. [Google Scholar] [CrossRef] [PubMed]
  24. Wang, J.; Mazza, G. Inhibitory Effects of Anthocyanins and Other Phenolic Compounds on Nitric Oxide Production in LPS/IFN-γ-Activated RAW 264.7 Macrophages. J. Agric. Food Chem. 2002, 50, 850–857. [Google Scholar] [CrossRef]
  25. Wang, H.; Nair, M.G.; Strasburg, G.M.; Chang, Y.-C.; Booren, A.M.; Gray, J.I.; DeWitt, D.L. Antioxidant and Antiinflammatory Activities of Anthocyanins and Their Aglycon, Cyanidin, from Tart Cherries. J. Nat. Prod. 1999, 62, 294–296. [Google Scholar] [CrossRef] [PubMed]
  26. Foti, M.C. Antioxidant Properties of Phenols. J. Pharm. Pharmacol. 2007, 59, 1673–1685. [Google Scholar] [CrossRef]
  27. Beconcini, D.; Felice, F.; Fabiano, A.; Sarmento, B.; Zambito, Y.; Di Stefano, R. Antioxidant and Anti-Inflammatory Properties of Cherry Extract: Nanosystems-Based Strategies to Improve Endothelial Function and Intestinal Absorption. Foods 2020, 9, 207. [Google Scholar] [CrossRef]
  28. Martínez, V.; Mitjans, M.; Vinardell, M.P. Cytoprotective Effects of Polyphenols Against Oxidative Damage. In Polyphenols in Human Health and Disease; Elsevier: Amsterdam, The Netherlands, 2014; Volume 1, pp. 275–288. [Google Scholar]
  29. Solorzano, E.R.; Roverso, M.; Bogialli, S.; Bortoli, M.; Orian, L.; Badocco, D.; Pettenuzzo, S.; Favaro, G.; Pastore, P. Antioxidant Activity of Zuccagnia-Type Propolis: A Combined Approach Based on LC-HRMS Analysis of Bioanalytical-Guided Fractions and Computational Investigation. Food Chem. 2024, 461, 140827. [Google Scholar] [CrossRef]
  30. Spiegel, M.; Cel, K.; Sroka, Z. The Mechanistic Insights into the Role of PH and Solvent on Antiradical and Prooxidant Properties of Polyphenols—Nine Compounds Case Study. Food Chem. 2023, 407, 134677. [Google Scholar] [CrossRef]
  31. Fu, Y.-H.; Zhang, Y.; Wang, F.; Zhao, L.; Shen, G.-B.; Zhu, X.-Q. Quantitative Evaluation of the Actual Hydrogen Atom Donating Activities of O–H Bonds in Phenols: Structure–Activity Relationship. RSC Adv. 2023, 13, 3295–3305. [Google Scholar] [CrossRef] [PubMed]
  32. Platzer, M.; Kiese, S.; Tybussek, T.; Herfellner, T.; Schneider, F.; Schweiggert-Weisz, U.; Eisner, P. Radical Scavenging Mechanisms of Phenolic Compounds: A Quantitative Structure-Property Relationship (QSPR) Study. Front. Nutr. 2022, 9, 882458. [Google Scholar] [CrossRef]
  33. Navarrete, M.; Rangel, C.; Espinosa-García, J.; Corchado, J.C. Theoretical Study of the Antioxidant Activity of Vitamin E: Reactions of α-Tocopherol with the Hydroperoxy Radical. J. Chem. Theory Comput. 2005, 1, 337–344. [Google Scholar] [CrossRef] [PubMed]
  34. Galano, A.; Raúl Alvarez-Idaboy, J. Computational Strategies for Predicting Free Radical Scavengers’ Protection Against Oxidative Stress: Where Are We and What Might Follow? Int. J. Quantum Chem. 2019, 119, e25665. [Google Scholar] [CrossRef]
  35. Zeppilli, D.; Orian, L. Concerted Proton Electron Transfer or Hydrogen Atom Transfer? An Unequivocal Strategy to Discriminate These Mechanisms in Model Systems. Phys. Chem. Chem. Phys. 2025, 27, 6312–6324. [Google Scholar] [CrossRef]
  36. Coassin, M.; Tomasi, A.; Vannini, V.; Ursini, F. Enzymatic Recycling of Oxidized Ascorbate in Pig Heart: One-Electron vs Two-Electron Pathway. Arch. Biochem. Biophys. 1991, 290, 458–462. [Google Scholar] [CrossRef]
  37. Bowry, V.W.; Mohr, D.; Cleary, J.; Stocker, R. Prevention of Tocopherol-Mediated Peroxidation in Ubiquinol-10-Free Human Low Density Lipoprotein. J. Biol. Chem. 1995, 270, 5756–5763. [Google Scholar] [CrossRef] [PubMed]
  38. Villalba, J.M.; Navarro, F.; Gómez-Díaz, C.; Arroyo, A.; Bello, R.I.; Navas, P. Role of Cytochrome B5 Reductase on the Antioxidant Function of Coenzyme Q in the Plasma Membrane. Mol. Asp. Med. 1997, 18, 7–13. [Google Scholar] [CrossRef] [PubMed]
  39. Galano, A.; Alvarez-Idaboy, J.R. A Computational Methodology for Accurate Predictions of Rate Constants in Solution: Application to the Assessment of Primary Antioxidant Activity. J. Comput. Chem. 2013, 34, 2430–2445. [Google Scholar] [CrossRef]
  40. Spiegel, M. Current Trends in Computational Quantum Chemistry Studies on Antioxidant Radical Scavenging Activity. J. Chem. Inf. Model. 2022, 62, 2639–2658. [Google Scholar] [CrossRef]
  41. Zeppilli, D.; Ribaudo, G.; Pompermaier, N.; Madabeni, A.; Bortoli, M.; Orian, L. Radical Scavenging Potential of Ginkgolides and Bilobalide: Insight from Molecular Modeling. Antioxidants 2023, 12, 525. [Google Scholar] [CrossRef]
  42. Bortoli, M.; Dalla Tiezza, M.; Muraro, C.; Pavan, C.; Ribaudo, G.; Rodighiero, A.; Tubaro, C.; Zagotto, G.; Orian, L. Psychiatric Disorders and Oxidative Injury: Antioxidant Effects of Zolpidem Therapy Disclosed In Silico. Comput. Struct. Biotechnol. J. 2019, 17, 311–318. [Google Scholar] [CrossRef]
  43. Galano, A.; Reiter, R.J. Melatonin and Its Metabolites vs Oxidative Stress: From Individual Actions to Collective Protection. J. Pineal Res. 2018, 65, e12514. [Google Scholar] [CrossRef]
  44. Galano, A.; Vargas, R.; Martínez, A. Carotenoids Can Act as Antioxidants by Oxidizing the Superoxideradical Anion. Phys. Chem. Chem. Phys. 2010, 12, 193–200. [Google Scholar] [CrossRef]
  45. Le On-Carmona, J.R.; Galano, A. Is Caffeine a Good Scavenger of Oxygenated Free Radicals? J. Phys. Chem. B 2011, 115, 4538–4546. [Google Scholar] [CrossRef]
  46. Castañeda-Arriaga, R.; Marino, T.; Russo, N.; Alvarez-Idaboy, J.R.; Galano, A. Chalcogen Effects on the Primary Antioxidant Activity of Chrysin and Quercetin. New J. Chem. 2020, 44, 9073–9082. [Google Scholar] [CrossRef]
  47. Galano, A.; Álvarez-Diduk, R.; Ramírez-Silva, M.T.; Alarcón-Ángeles, G.; Rojas-Hernández, A. Role of the Reacting Free Radicals on the Antioxidant Mechanism of Curcumin. Chem. Phys. 2009, 363, 13–23. [Google Scholar] [CrossRef]
  48. Alberto, M.E.; Russo, N.; Grand, A.; Galano, A. A Physicochemical Examination of the Free Radical Scavenging Activity of Trolox: Mechanism, Kinetics and Influence of the Environment. Phys. Chem. Chem. Phys. 2013, 15, 4642. [Google Scholar] [CrossRef] [PubMed]
  49. Dalla Tiezza, M.; Hamlin, T.A.; Bickelhaupt, F.M.; Orian, L. Radical Scavenging Potential of the Phenothiazine Scaffold: A Computational Analysis. ChemMedChem 2021, 16, 3763–3771. [Google Scholar] [CrossRef] [PubMed]
  50. Zeppilli, D.; Grolla, G.; Di Marco, V.; Ribaudo, G.; Orian, L. Radical Scavenging and Anti-Ferroptotic Molecular Mechanism of Olanzapine: Insight from a Computational Analysis. Inorg. Chem. 2024, 63, 21856–21867. [Google Scholar] [CrossRef]
  51. Al-Sehemi, A.G.; Irfan, A. Effect of Donor and Acceptor Groups on Radical Scavenging Activity of Phenol by Density Functional Theory. Arab. J. Chem. 2017, 10, S1703–S1710. [Google Scholar] [CrossRef]
  52. Inami, K.; Iizuka, Y.; Furukawa, M.; Nakanishi, I.; Ohkubo, K.; Fukuhara, K.; Fukuzumi, S.; Mochizuki, M. Chlorine Atom Substitution Influences Radical Scavenging Activity of 6-Chromanol. Bioorg. Med. Chem. 2012, 20, 4049–4055. [Google Scholar] [CrossRef] [PubMed]
  53. Škorňa, P.; Poliak, P.; Klein, E.; Lukeš, V. Theoretical Study of the Substituent Effect on the Hydrogen Atom Transfer Mechanism of Meta- and Para-Substituted Benzenetellurols. Comput. Theor. Chem. 2016, 1079, 64–69. [Google Scholar] [CrossRef]
  54. Pratt, D.A.; DiLabio, G.A.; Brigati, G.; Pedulli, G.F.; Valgimigli, L. 5-Pyrimidinols: Novel Chain-Breaking Antioxidants More Effective than Phenols. J. Am. Chem. Soc. 2001, 123, 4625–4626. [Google Scholar] [CrossRef]
  55. Isborn, C.; Hrovat, D.A.; Borden, W.T.; Mayer, J.M.; Carpenter, B.K. Factors Controlling the Barriers to Degenerate Hydrogen Atom Transfers. J. Am. Chem. Soc. 2005, 127, 5794–5795. [Google Scholar] [CrossRef] [PubMed]
  56. Ferraz-Caetano, J.; Teixeira, F.; Cordeiro, M.N.D.S. Navigating epoxidation complexity: Building a data science toolbox to design vanadium catalysts. New J. Chem. 2024, 48, 5097–5100. [Google Scholar] [CrossRef]
  57. Ferraz-Caetano, J.; Teixeira, F.; Cordeiro, M.N.D.S. Data-Driven, Explainable Machine Learning Model for Predicting Volatile Organic Compounds’ Standard Vaporization Enthalpy. Chemosphere 2024, 359, 142257. [Google Scholar] [CrossRef] [PubMed]
  58. Li, Y.-P.; Han, K.; Grambow, C.A.; Green, W.H. Self-Evolving Machine: A Continuously Improving Model for Molecular Thermochemistry. J. Phys. Chem. A 2019, 123, 2142–2152. [Google Scholar] [CrossRef]
  59. Butler, K.T.; Davies, D.W.; Cartwright, H.; Isayev, O.; Walsh, A. Machine Learning for Molecular and Materials Science. Nature 2018, 559, 547–555. [Google Scholar] [CrossRef] [PubMed]
  60. Muraro, C.; Polato, M.; Bortoli, M.; Aiolli, F.; Orian, L. Radical Scavenging Activity of Natural Antioxidants and Drugs: Development of a Combined Machine Learning and Quantum Chemistry Protocol. J. Chem. Phys. 2020, 153, 114117. [Google Scholar] [CrossRef]
  61. Fujimoto, T.; Gotoh, H. Prediction and Chemical Interpretation of Singlet-Oxygen-Scavenging Activity of Small Molecule Compounds by Using Machine Learning. Antioxidants 2021, 10, 1751. [Google Scholar] [CrossRef]
  62. Zhong, S.; Zhang, K.; Wang, D.; Zhang, H. Shedding Light on “Black Box” Machine Learning Models for Predicting the Reactivity of HO Radicals toward Organic Compounds. Chem. Eng. J. 2021, 405, 126627. [Google Scholar] [CrossRef]
  63. Shen, Y.; Liu, C.; Chi, K.; Gao, Q.; Bai, X.; Xu, Y.; Guo, N. Development of a Machine Learning-Based Predictor for Identifying and Discovering Antioxidant Peptides Based on a New Strategy. Food Control 2022, 131, 108439. [Google Scholar] [CrossRef]
  64. Musa, K.H.; Abdullah, A.; Al-Haiqi, A. Determination of DPPH Free Radical Scavenging Activity: Application of Artificial Neural Networks. Food Chem. 2016, 194, 705–711. [Google Scholar] [CrossRef] [PubMed]
  65. Shang, Y.; Li, X.; Le, T.N.; Zhou, J.; Zhou, P.; Leong, L.P.; Li, W. Machine Learning-Based Screening of Antioxidant Activity in Resveratrol Dimers. Chem. Phys. Lett. 2025, 879, 142386. [Google Scholar] [CrossRef]
  66. Bento, A.P.; Hersey, A.; Félix, E.; Landrum, G.; Gaulton, A.; Atkinson, F.; Bellis, L.J.; De Veij, M.; Leach, A.R. An Open Source Chemical Structure Curation Pipeline Using RDKit. J. Cheminform. 2020, 12, 51. [Google Scholar] [CrossRef] [PubMed]
  67. Ferraz-Caetano, J.; Teixeira, F.; Cordeiro, M.N.D.S. Optimising Materials Properties with Minimal Data: Lessons from Vanadium Catalyst Modelling. In Challenges and Advances in Computational Chemistry and Physics; Springer Science and Business Media B.V.: Cham, Switzerland, 2025; Volume 39, pp. 117–138. [Google Scholar]
  68. Hamakawa, Y.; Miyao, T. Understanding Conformation Importance in Data-Driven Property Prediction Models. J. Chem. Inf. Model. 2025, 65, 3388–3404. [Google Scholar] [CrossRef]
  69. Bortoli, M.; Orian, L. Antioxidant Potential of Anthocyanidins: A Healthy Computational Activity for High School and Undergraduate Students. J. Chem. Educ. 2023, 100, 2591–2600. [Google Scholar] [CrossRef]
  70. Filippi, M. Topologia Molecolare e Attività Di Scavenging Di Radicali: Uno Studio Computazionale Sistematico Su Fenoli e Polifenoli. Bachelor’s Thesis, Universià di Padova, Padova, Italy, 2025. [Google Scholar]
  71. Frisch, M.J.; Trucks, G.W.; Schlegel, H.B.; Scuseria, G.E.; Robb, M.A.; Cheeseman, J.R.; Scalmani, G.; Barone, V.; Petersson, G.A.; Nakatsuji, H.; et al. Gaussian 16, Revision C.01; Gaussian Inc.: Wallingford, CT, USA, 2016. [Google Scholar]
  72. Zhao, Y.; Truhlar, D.G. The M06 Suite of Density Functionals for Main Group Thermochemistry, Thermochemical Kinetics, Noncovalent Interactions, Excited States, and Transition Elements: Two New Functionals and Systematic Testing of Four M06-Class Functionals and 12 Other Functionals. Theor. Chem. Acc. 2008, 120, 215–241. [Google Scholar] [CrossRef]
  73. Pracht, P.; Grimme, S.; Bannwarth, C.; Bohle, F.; Ehlert, S.; Feldmann, G.; Gorges, J.; Müller, M.; Neudecker, T.; Plett, C.; et al. CREST—A Program for the Exploration of Low-Energy Molecular Chemical Space. J. Chem. Phys. 2024, 160, 114110. [Google Scholar] [CrossRef]
  74. Marenich, A.V.; Cramer, C.J.; Truhlar, D.G. Universal Solvation Model Based on Solute Electron Density and on a Continuum Model of the Solvent Defined by the Bulk Dielectric Constant and Atomic Surface Tensions. J. Phys. Chem. B 2009, 113, 6378–6396. [Google Scholar] [CrossRef]
  75. Galano, A.; Medina, M.E.; Tan, D.X.; Reiter, R.J. Melatonin and Its Metabolites as Copper Chelating Agents and Their Role in Inhibiting Oxidative Stress: A Physicochemical Analysis. J. Pineal Res. 2015, 58, 107–116. [Google Scholar] [CrossRef] [PubMed]
  76. Martínez, A.; Galano, A.; Vargas, R. Free Radical Scavenger Properties of α-Mangostin: Thermodynamics and Kinetics of HAT and RAF Mechanisms. J. Phys. Chem. B 2011, 115, 12591–12598. [Google Scholar] [CrossRef] [PubMed]
  77. Galano, A. Antioxidants: The Chemical Complexity Behind a Simple Word. Acc. Chem. Res. 2025, 58, 3481–3493. [Google Scholar] [CrossRef] [PubMed]
  78. Zeppilli, D.; Aldinio-Colbachini, A.; Ribaudo, G.; Tubaro, C.; Dalla Tiezza, M.; Bortoli, M.; Zagotto, G.; Orian, L. Antioxidant Chimeric Molecules: Are Chemical Motifs Additive? The Case of a Selenium-Based Ligand. Int. J. Mol. Sci. 2023, 24, 11797. [Google Scholar] [CrossRef] [PubMed]
  79. Landrum, G. RDKit: Open-Source Cheminformatics 2025_03_4 (Q1 2025). Available online: http://www.rdkit.org/ (accessed on 18 July 2025).
  80. Python Software Foundation—Python Language Reference, version 3.12.7; Python Software Foundation: Beaverton, OR, USA, 2024; Available online: http://www.python.org (accessed on 18 July 2025).
  81. RDKit: Open-Source Cheminformatics—Descriptor Guide Online Webbook. Available online: https://www.rdkit.org/docs/GettingStartedInPython.html#list-of-available-descriptors (accessed on 10 December 2025).
Figure 1. Schematic representation of QM calculations to compute Δ G f H A T o of N molecules with M reaction sites, each with J radicals in 3 different media.
Figure 1. Schematic representation of QM calculations to compute Δ G f H A T o of N molecules with M reaction sites, each with J radicals in 3 different media.
Aichem 01 00008 g001
Figure 2. Distribution of ΔG° energy values in the experimental dataset into molecule classes.
Figure 2. Distribution of ΔG° energy values in the experimental dataset into molecule classes.
Aichem 01 00008 g002
Figure 3. Regression model prediction chart for the Random Forest algorithm in the optimized model, distributed across molecule classes. The green band represents the RMSE threshold.
Figure 3. Regression model prediction chart for the Random Forest algorithm in the optimized model, distributed across molecule classes. The green band represents the RMSE threshold.
Aichem 01 00008 g003
Figure 4. Test set R2 values of the optimized Random Forest model across molecule classes, with error bars showing standard deviations.
Figure 4. Test set R2 values of the optimized Random Forest model across molecule classes, with error bars showing standard deviations.
Aichem 01 00008 g004
Figure 5. Mean absolute error (MAE) and root mean squared error (RMSE) of the optimized Random Forest model across molecule classes, with error bars representing standard deviations.
Figure 5. Mean absolute error (MAE) and root mean squared error (RMSE) of the optimized Random Forest model across molecule classes, with error bars representing standard deviations.
Aichem 01 00008 g005
Figure 6. General structure of anthocyanidins with numbering of the main reaction sites. Site 19 is highlighted in red, being different only for the molecule ARN.
Figure 6. General structure of anthocyanidins with numbering of the main reaction sites. Site 19 is highlighted in red, being different only for the molecule ARN.
Aichem 01 00008 g006
Figure 7. Overall feature importance of the optimized Random Forest model grouped by descriptor type (electronic, experimental, structural, VSA) and source (molecule, radical, site).
Figure 7. Overall feature importance of the optimized Random Forest model grouped by descriptor type (electronic, experimental, structural, VSA) and source (molecule, radical, site).
Aichem 01 00008 g007
Figure 8. Normalized global importance of the top 10 features of the optimized Random Forest model, colored by descriptor type (electronic, experimental, structural, VSA).
Figure 8. Normalized global importance of the top 10 features of the optimized Random Forest model, colored by descriptor type (electronic, experimental, structural, VSA).
Aichem 01 00008 g008
Table 1. Model performance statistical results using different algorithms and descriptor sets.
Table 1. Model performance statistical results using different algorithms and descriptor sets.
Model
Algorithm
Descriptors
Used
R2 TestMAE/
kcal mol−1
RMSE/
kcal mol−1
MLPFull set0.944 ± 0.0023.57 ± 0.055.06 ± 0.08
Gradient BoostingFull set0.923 ± 0.0024.36 ± 0.055.93 ± 0.07
Random ForestFull set0.944 ± 0.0022.71 ± 0.055.1 ± 0.1
Optimized Random ForestOne-hot encoding non-numeric0.933 ± 0.0023.27 ± 0.055.50 ± 0.09
Optimized Random ForestFull set0.944 ± 0.0022.74 ± 0.055.0 ± 0.1
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Zeppilli, D.; Ferraz-Caetano, J.; Cordeiro, M.N.D.S.; Orian, L. A Hybrid Machine Learning and Quantum Mechanical Strategy for Predicting Radical Scavenging Potential. AI Chem. 2026, 1, 8. https://doi.org/10.3390/aichem1020008

AMA Style

Zeppilli D, Ferraz-Caetano J, Cordeiro MNDS, Orian L. A Hybrid Machine Learning and Quantum Mechanical Strategy for Predicting Radical Scavenging Potential. AI Chemistry. 2026; 1(2):8. https://doi.org/10.3390/aichem1020008

Chicago/Turabian Style

Zeppilli, Davide, José Ferraz-Caetano, M. Natália D. S. Cordeiro, and Laura Orian. 2026. "A Hybrid Machine Learning and Quantum Mechanical Strategy for Predicting Radical Scavenging Potential" AI Chemistry 1, no. 2: 8. https://doi.org/10.3390/aichem1020008

APA Style

Zeppilli, D., Ferraz-Caetano, J., Cordeiro, M. N. D. S., & Orian, L. (2026). A Hybrid Machine Learning and Quantum Mechanical Strategy for Predicting Radical Scavenging Potential. AI Chemistry, 1(2), 8. https://doi.org/10.3390/aichem1020008

Article Metrics

Back to TopTop