Next Issue
Volume 1, September
Previous Issue
Volume 1, March
 
 

AI Chem., Volume 1, Issue 2 (June 2026) – 4 articles

  • Issues are regarded as officially published after their release is announced to the table of contents alert mailing list.
  • You may sign up for e-mail alerts to receive table of contents of newly released issues.
  • PDF is the official format for papers published in both, html and pdf forms. To view the papers in pdf format, click on the "PDF Full-text" link, and use the free Adobe Reader to open them.
Order results
Result details
Select all
Export citation of selected articles as:
18 pages, 7505 KB  
Article
Does DrugCLIP Find the Right Pocket? A Systematic Evaluation of Binding-Site Identification Across 42 Drug Targets
by Bocheng Xie, Xiaokang Guo, Pengwei Xiao and Chao Yang
AI Chem. 2026, 1(2), 9; https://doi.org/10.3390/aichem1020009 - 16 May 2026
Viewed by 1200
Abstract
Contrastive learning-based models such as DrugCLIP have recently emerged as scalable tools for structure-based virtual screening by embedding protein structures and small molecules into a shared representation space. While these approaches demonstrate high throughput and competitive screening performance in ligand retrieval tasks, their [...] Read more.
Contrastive learning-based models such as DrugCLIP have recently emerged as scalable tools for structure-based virtual screening by embedding protein structures and small molecules into a shared representation space. While these approaches demonstrate high throughput and competitive screening performance in ligand retrieval tasks, their ability to correctly identify biologically relevant ligand-binding pockets has not been systematically evaluated. Here, we construct a benchmarking dataset comprising 42 pharmacologically diverse human protein targets with experimentally validated drug-bound structures spanning multiple target families. Using this dataset, we evaluate the pocket recognition capability of DrugCLIP and compare its performance with a traditional structure-based workflow (Fpocket combined with ESSA) and a machine learning-based method (P2Rank). DrugCLIP shows robust performance for well-characterized target classes, including kinases (10/10) and nuclear receptors (5/5), but exhibits markedly reduced accuracy for ion channels (1/4), GPCRs (3/5), and transporters (3/5). Notably, pocket prediction accuracy does not strongly correlate with structural data availability, suggesting that intrinsic pocket characteristics rather than training data abundance primarily affect model performance. Across the benchmark, DrugCLIP achieves an overall success rate of 71% (95% CI: 56–83%), compared with 79% (95% CI: 64–88%) for Fpocket+ESSA, and 93% (95% CI: 81–98%) for P2Rank. McNemar’s test showed no significant difference between DrugCLIP and Fpocket+ESSA (p = 0.508), whereas P2Rank significantly outperformed DrugCLIP (p = 0.012). Together, these results provide a quantitative evaluation of pocket recognition by contrastive learning-based models and highlight key limitations of embedding-based approaches for pocket localization. Full article
Show Figures

Figure 1

15 pages, 2213 KB  
Article
A Hybrid Machine Learning and Quantum Mechanical Strategy for Predicting Radical Scavenging Potential
by Davide Zeppilli, José Ferraz-Caetano, M. Natália D. S. Cordeiro and Laura Orian
AI Chem. 2026, 1(2), 8; https://doi.org/10.3390/aichem1020008 - 15 May 2026
Viewed by 875
Abstract
We designed a supervised machine learning framework to predict standard Gibbs free energies, ΔG°, of formal hydrogen atom transfer (f-HAT) for phenolic antioxidants across different radicals and media, enabling rapid and chemically interpretable screening. We curated a DFT dataset of 71 molecules (phenolic [...] Read more.
We designed a supervised machine learning framework to predict standard Gibbs free energies, ΔG°, of formal hydrogen atom transfer (f-HAT) for phenolic antioxidants across different radicals and media, enabling rapid and chemically interpretable screening. We curated a DFT dataset of 71 molecules (phenolic compounds and anthocyanidins), with 207 reaction sites, 10 radical reactive oxygen/sulfur species, and three environments (leading to a total of 6210 ΔG° values). The models amass 106 numerical RDKit descriptors, augmented with one-hot encodings of medium, site, radical, and structural class, and were evaluated through a leave-one-molecule-out protocol. Among the tested regression algorithms, the random forest regressor provides the best balance of accuracy and robustness with both R2 test (≈0.94) and MAE (2.74 kcal mol−1; RMSE (≈5.0 kcal mol−1)), close to DFT chemical accuracy. The feature-importance analysis revealed that “electronic” and “experimental” (site/group) descriptors primarily drive predictions, with the radical’s maximum absolute partial charge being the most important descriptor in the prediction of a radical’s ΔG°. These results suggest that descriptor-driven RF (Random Forest) models can generalize across chemical space to provide interpretable ΔG° predictions, providing a path for chemists towards a scalable route to prioritize antioxidant candidates for broader molecular families. Full article
Show Figures

Graphical abstract

20 pages, 2522 KB  
Article
Active Learning on Protein Language Model Embeddings Accelerates Rubisco Variant Discovery for Desired Traits
by James Young, Dillon Nelson, Liping Gu and Ruanbao Zhou
AI Chem. 2026, 1(2), 7; https://doi.org/10.3390/aichem1020007 - 15 Apr 2026
Cited by 1 | Viewed by 1448
Abstract
Ribulose-1,5-bisphosphate carboxylase/oxygenase (Rubisco) efficiency constrains carbon fixation, making it a high-value target in biotechnology. The core task of this work is a supervised regression and ranking problem on Rubisco: given a numerical representation of a protein sequence (a PLM embedding), we predict continuous [...] Read more.
Ribulose-1,5-bisphosphate carboxylase/oxygenase (Rubisco) efficiency constrains carbon fixation, making it a high-value target in biotechnology. The core task of this work is a supervised regression and ranking problem on Rubisco: given a numerical representation of a protein sequence (a PLM embedding), we predict continuous phenotypic scores such as an enzyme kinetic proxy or fitness value. The predictions then guide which variants to test next. Engineering Rubisco is a point of focus but remains challenging due to selection forces in vivo and the combinatorial space of potential mutants for ex vivo uses. We combine protein language model (PLM) embeddings with tabular learning to model Rubisco variant landscapes in two regimes. First, we analyze deep mutational scanning data providing inferred kinetic proxies, including Km for CO2 and Vmax. Second, we model a cyanobacterial screening dataset measuring mutant fitness under differing oxygen and nitrogen regimes, enabling an oxygen tolerance objective. Across tasks, a tabular foundation (TabPFN-2.5) model outperforms gradient-boosted trees on rank-based criteria for variant prioritization, including Spearman correlation and top 5% hit recovery. We then simulate active-learning campaigns initialized with 200 measured variants and iteratively acquiring batches of 48. Model-guided selection recovers more top-performing mutants than random sampling at fixed experimental budgets, even with a conservative XGBoost surrogate. We also demonstrate that Rubisco large-subunit embeddings predict cyanobacterial doubling time and cross-species kinetic parameters, suggesting that Rubisco representation remains meaningful across organisms even with multi-objective cellular constraints. Together, these results support a practical, data-efficient workflow for enzyme engineering and motivate objective-aware design strategies that complement directed evolution. Full article
Show Figures

Graphical abstract

18 pages, 4367 KB  
Article
Leveraging Bag Dissimilarity Regularized Multi-Instance Learning for Analyzing Infrared Spectra of Heterogeneous Objects
by Shiluo Huang and Zheyu Zou
AI Chem. 2026, 1(2), 6; https://doi.org/10.3390/aichem1020006 - 27 Mar 2026
Viewed by 748
Abstract
Infrared (IR) spectroscopy is a powerful tool for characterizing molecular structures and chemical groups, offering advantages such as low cost, rapid analysis, and non-destructive testing. When analyzing heterogeneous objects, spectra are typically measured from different regions to capture the local variations, presenting a [...] Read more.
Infrared (IR) spectroscopy is a powerful tool for characterizing molecular structures and chemical groups, offering advantages such as low cost, rapid analysis, and non-destructive testing. When analyzing heterogeneous objects, spectra are typically measured from different regions to capture the local variations, presenting a multi-instance learning (MIL) problem. However, existing methods primarily rely on multi-instance assumptions or explicit bag representations, often failing to fully capture the intrinsic information from implicit representations. We introduce a bag dissimilarity regularized MIL framework for analyzing IR spectra of heterogeneous objects, which integrates both explicit and implicit representations to effectively learn the MIL bags. Specifically, a bag dissimilarity regularization term is utilized to extract implicit representations, which subsequently guide the classifier based on explicit representations to enhance generalization performance. The proposed method was validated on two heterogeneous detection tasks: polydimethylsiloxane (PDMS) block assessment and polyethylene terephthalate (PET) fiber inspection. Experimental results demonstrate that our approach significantly outperforms existing methods on both datasets with a considerable margin. Full article
Show Figures

Figure 1

Previous Issue
Next Issue
Back to TopTop