Next Article in Journal
AI Chemistry: Advancing Chemistry and Its Applications Through Artificial Intelligence
Previous Article in Journal
BioRamanNet: A Neural Network Framework for Biological Raman Spectroscopy Classification
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Assessing Modern AI-Driven Protein-Ligand Modeling with Phenethylamine and Tryptamine Psychedelics

by
Benjamin R. Cummins
* and
Charles D. Nichols
Department of Pharmacology and Experimental Therapeutics, Louisiana State University Health Sciences Center, New Orleans, LA 70112, USA
*
Author to whom correspondence should be addressed.
AI Chem. 2026, 1(1), 4; https://doi.org/10.3390/aichem1010004
Submission received: 30 December 2025 / Revised: 27 January 2026 / Accepted: 2 February 2026 / Published: 10 February 2026

Abstract

Modern advances in artificial intelligence have accelerated the development of computational tools for protein–ligand structure prediction, yet their real-world performance remains uneven across receptor classes and ligand chemotypes. Recently published cryo-EM structures of several different psychedelics bound to the serotonin 5HT2A receptor provide a unique opportunity to explore how modern AI-based modeling performs in a pharmacologically important GPCR system. Here, we compare three major approaches: AI-based protein–ligand cofolding (Boltz-2), a leading AI-driven docking module (Uni-Mol Docking v2), and a widely used classical physics-based docking pipeline (AutoDock Vina) across a series of tryptamine and phenethylamine psychedelics. Predicted binding poses were comparatively assessed through structural alignment with these newly available cryo-EM complexes. Additionally, calcium-mobilization assays were performed to provide a coarse functional readout for comparison with computationally predicted binding affinities. This study integrates methodological review with exploratory benchmarking to illustrate how different modeling paradigms behave on a shared receptor–ligand test set. Our results highlight substantial variation between modeling strategies, with AI-based cofolding often producing global binding orientations more closely resembling experimental structures, and classical docking showing greater variability across ligands, while still outperforming AI-driven docking on average. These observations underscore both the growing utility and current limitations of AI-assisted structure prediction in serotonergic drug discovery, and emphasize the importance of careful, experimentally anchored evaluation as such tools continue to advance.

1. Introduction

Artificial intelligence (AI) has begun reshaping the landscape of computational chemistry and drug discovery. Over the past several years, deep learning-based methods have expanded from sequence and structure prediction into tasks traditionally dominated by physics-based modeling, including docking, free-energy estimation, and protein–ligand interaction prediction. Models trained on large structural and biochemical datasets can now infer binding poses and approximate affinities in ways that complement, and in some cases rival, conventional molecular docking workflows. Yet despite rapid progress, systematic evaluations anchored to experimentally resolved complexes are still relatively rare. The recent resolution of new cryo-EM structures for several psychedelics bound to the serotonin 5-HT2A receptor provides a well-suited testbed for exploring both the capabilities and limitations of emerging AI-based structure-prediction methods [1].
The 5-HT2A receptor is a class A G protein-coupled receptor (GPCR) central to the pharmacology of classical psychedelics. Certain tryptamines such as N,N-dimethyltryptamine (DMT) and phenethylamines such as mescaline produce their characteristic psychoactive effects primarily through agonism at this receptor [2]. Because GPCRs possess substantial conformational flexibility and complex activation pathways, they have historically been challenging for structure-based modeling, making them an informative system for evaluating newer AI-driven approaches. The cryo-EM structures of psilocin:5-HT2A, DMT:5-HT2A, mescaline:5-HT2A and other complexes define high-confidence binding orientations for these two major psychedelic chemotypes, enabling direct comparison between predicted poses and experimentally observed ligand geometries [1].
To use these structures as benchmarks, we selected a representative panel of closely related tryptamine and phenethylamine psychedelics, including DiPT, 4-HO-DiPT, 2C-B, and DOB(–), to test alongside psilocin, DMT, and mescaline (Figure 1). These additional four ligands share core scaffolds with the cryo-EM–resolved compounds, allowing predicted orientations to be assessed relative to well-defined empirical binding modes. Their extensive pharmacological literature also makes them suitable for integrating computational predictions with functional readouts.
To examine how modern computational tools perform on this receptor–ligand system, we focused on three widely used modeling paradigms that span the current methodological landscape: AI-based protein–ligand cofolding, AI-driven docking, and classical physics-based docking.
Boltz-2 (developed by MIT Jameel Clinic & Recursion) represents the newest of these approaches, extending the Boltz framework by modeling protein–ligand binding as a joint generative process, in which the receptor pocket and ligand are simultaneously optimized toward a thermodynamically plausible bound state. Architecturally, it uses SE(3)-equivariant neural networks to represent geometric relationships between ligand atoms and protein residues, enabling the model to reason over 3D structure without losing rotational or translational invariance. Boltz-2 is trained with a denoising diffusion objective, in which corrupted protein–ligand geometries are iteratively refined toward native conformations; this implicitly teaches the model how ligand binding shapes the local protein environment and vice versa [3]. The model is additionally trained on large datasets of experimental affinity measurements (Kd, Ki, IC50), which are collapsed into a single qualitative “affinity” target: this allows Boltz-2 to output μM-range estimates that should be interpreted comparatively rather than as absolute binding constants.
One distinguishing feature of Boltz-2 is its explicit capacity for induced-fit modeling. Classical docking typically assumes a static receptor backbone, but Boltz-2 dynamically adjusts side-chain and local backbone conformations as part of its cofolding trajectory. This is particularly relevant for GPCRs, where binding pockets are partially plastic and frequently undergo subtle rearrangements upon agonist binding. Boltz-2’s training incorporates tens of thousands of protein–ligand complexes, giving it a broad structural prior that complements its learned energy function [4]. Its rising popularity stems both from its claimed improvements in pose accuracy relative to docking benchmarks and from its accessibility, offering high-level predictions without requiring force fields, search algorithms, or explicit scoring-function tuning. Its relevance to the present study stems from its growing visibility in AI-driven drug design and its explicit design goal of capturing protein flexibility, a key challenge for GPCR modeling.
Uni-Mol Docking v2 (developed by DP Technology) represents an alternative AI-centric strategy. Rather than cofolding receptor and ligand, it belongs to the emerging class of 3D molecular representation learners built on transformer-like attention mechanisms adapted for spatial data. Uni-Mol Docking v2 uses a dual-encoder architecture: one encoder processes ligand conformers, and the other encodes the protein binding pocket, both using geometric tokenization and pair-distance features. The model is trained end-to-end to predict ligand placements via learned distributions over interatomic distances and angles, effectively internalizing docking heuristics from large-scale structural datasets rather than relying on physical scoring terms [5].
A key innovation in Uni-Mol is its use of pretrained molecular backbones. The ligand encoder is pretrained on millions of small molecules using self-supervised tasks such as masked atom reconstruction and 3D conformation prediction. The protein encoder is pretrained on fragmented pocket structures, enabling the model to develop a statistical prior for steric compatibility and hydrogen-bonding geometries. During docking, Uni-Mol predicts a ligand pose by sampling from a generative head that outputs coordinates or coordinate deltas conditioned on the pocket representation [5]. The creators have reported strong performance on standard docking benchmarks such as CrossDocked and PDBbind, often outperforming classical docking baselines, particularly for ligands and pockets similar to those seen during training. Unlike Boltz-2, it does not explicitly model protein flexibility, but it leverages data-driven priors that sometimes allow it to “guess” plausible induced-fit states within a fixed pocket. In this way, it occupies a methodological space between diffusion-based cofolding (Boltz-2) and physics-based search (AutoDock Vina).
AutoDock Vina (developed by Scripps CCSB), in contrast, represents the classical, physics-inspired approach to modeling molecular recognition. It relies on a scoring function combining steric, hydrophobic, hydrogen-bonding, and empirical affinity terms, optimized through a quasi-Newton local search embedded within a global stochastic sampling algorithm. Unlike neural models, which learn geometric and energetic priors from large datasets, Vina defines its scoring function analytically and evaluates candidate poses deterministically during the search [6,7]. Mechanistically, Vina explores the ligand’s conformational and rotational space by sampling rigid-body orientations and rotatable bonds, inserting these candidate geometries into a predefined protein pocket, and minimizing the scoring function to estimate a low-energy pose.
Because the receptor is treated as largely rigid aside from user-specified side-chain flexibility, Vina can struggle with targets requiring substantial induced fit, an issue particularly relevant for GPCRs. Nevertheless, its strengths include speed, interpretability, reproducibility, and over twenty years of accumulated validation across diverse targets. AutoDock Vina remains one of the most widely used docking tools in academic and industrial settings, forming a natural baseline for comparison against AI-driven models. Although it lacks the data-driven priors of AI-based methods, its many years of use and extensive benchmarking make it an essential reference point in a comparative study, allowing evaluation of whether AI methods meaningfully outperform classical workflows on a GPCR system where ligand-specific micro-adjustments of the binding pocket are known to influence activity.
Notably, an intermediate class of docking tools has emerged that augments classical physics-based search with machine learning–based scoring functions. A prominent example is Gnina (developed by Koes lab, University of Pittsburgh), an open-source molecular docking platform that extends the AutoDock Vina framework by incorporating deep learning into the scoring stage. Unlike Vina’s empirical scoring function, Gnina uses an ensemble of three-dimensional convolutional neural networks (CNNs) trained on large datasets of protein–ligand complexes to reassess and rank candidate poses generated during docking. On standard benchmarks, this approach significantly improves the likelihood that the top-ranked pose is close to the experimentally observed binding mode [8]. Intermediate implements such as this were not included as a primary method in this study, as several recent publications have already benchmarked Gnina’s deep learning assisted docking relative to classical Vina and other methods across large datasets, and our goal is to establish a simple and contrasting baseline of classical physics-based docking (AutoDock Vina) against modern AI methods with fundamentally different modeling philosophies [9]. In this context, Vina serves as an accessible representative of established physics-based workflows, acknowledging that AI-augmented variants such as Gnina can improve pose ranking but sit conceptually between the paradigms we aim to compare.
Together, these three tools (Boltz, Uni-Mol, Vina) illustrate three major methodological categories now available for modeling protein–ligand interactions: learned structural cofolding, learned docking, and physics-based docking. Evaluating their performance on a shared set of psychedelics, anchored by recently resolved 5-HT2A structures, allows both a practical and conceptual examination of how AI is reshaping structure-based neuroscience drug discovery.

2. Materials and Methods

2.1. Ligand Preparation

SMILES strings for the seven ligands (tryptamines psilocin, DMT, DiPT & 4-HO-DiPT, and phenethylamines 2C-B, DOB(-) & mescaline) were obtained from PubChem, and are available for download (Supplementary Materials). Three-dimensional structures were generated in Avogadro 1.2 (developed by Open Chemistry), followed by MMFF94 force-field geometry optimization to ensure low-energy conformations suitable for docking and cofolding [10]. Each ligand structure was exported in PDB and SDF formats for compatibility with downstream software.

2.2. Receptor Structure Preparation

Cryo-EM structures of 5-HT2A bound to psilocin, DMT and mescaline were downloaded from the recent Gumpper et al. publication (PDB codes 9AS7, 9AS1 and 9AS5, respectively) [1]. Receptors were prepared in UCSF ChimeraX 1.11 (developed by RBVI, UCSF) using the Dock Prep tool with standard preprocessing:
  • Removal of co-crystallized ligands, water molecules, stabilizing nanobodies, and detergents
  • Identification and correction of incomplete side chains using Dunbrack rotamer library
  • Addition of all hydrogens with protonation states assigned at physiological pH
  • Optimization of hydrogen-bonding networks
  • Charge assignment using AMBER ff14SB residue templates
  • Removal of alternate conformers and insertion codes
These prepared receptor models were used across all three computational platforms to ensure methodological comparability [11,12]. All ligand protonation states were also assigned at physiological pH to enable relevant interactions. For example, ligands with an amine group had a protonated nitrogen to facilitate expected contact with D3.32.

2.3. Definition of the Docking Region

Binding boxes were defined in PyMOL 1.8.5 (developed by Schrödinger) by centering a 28 Å cubic region around the orthosteric ligand position in each cryo-EM structure (Figure 2) [13]. This space was chosen to fully encompass the canonical binding pocket and the known extended binding trajectories of phenethylamines and tryptamines. The same box coordinates were used for AutoDock Vina and Uni-Mol Docking v2.

2.4. AutoDock Vina Docking via AMDock

Classical physics-based docking was performed using AutoDock Vina 1.2.1 within the AMDock 1.5.2 graphical interface (developed by Valdés-Tresanco et al.) (Figure 3) [14]. Receptor and ligand files were converted to PDBQT format via AutoDockTools, applying Gasteiger charges and rotatable-bond definitions.
Vina outputs ten poses ranked by predicted binding free energy (Figure 4). Each set of poses was visually aligned to the cryo-EM ligand using ChimeraX. For each ligand, the pose with the closest structural alignment (typically but not always Vina’s top-ranked prediction) was selected, and its predicted affinity value (kcal/mol) and estimated Ki (μM) was used for comparison.

2.5. Uni-Mol Docking v2

The same receptor PDB file, ligand SDF files, and PyMOL-defined docking coordinates were used as input for Uni-Mol Docking v2 (stable model), accessed via Bohrium, as recommended by the developers (Figure 5). Output structures were downloaded as SDF files and directly loaded into ChimeraX for alignment with both the cryo-EM structures and the Vina/Boltz-2 predictions.

2.6. Boltz-2 Protein–Ligand Cofolding

Boltz-2 simulations were run locally on version 2.2.1 using both command-line inference and the ChimeraX Boltz tool [12]. Human 5HT2A receptor FASTA sequence was obtained from NCBI, and multiple-sequence alignments (MSAs) generated using the Boltz-recommended remote MSA server [15]. Structures for each ligand-receptor complex were predicted using default settings, returning five structural samples (almost always nearly identical to each other) and one predicted affinity value (µM). Predicted structures were aligned to the cryo-EM complexes in ChimeraX. The predicted µM affinity is qualitative, as Boltz-2 was trained on mixed Kd/Ki/IC50 datasets without distinguishing definitions.

2.7. Structural Comparison and Pose Evaluation

All predicted ligand poses were loaded simultaneously in ChimeraX and assessed through RMSD (root mean square deviation) measurement of heavy atoms relative to the cryo-EM ligand, visual evaluation of global orientation (indole/phenethylamine disposition, pose inversion, depth of insertion) and local interaction analysis (S242/S5.46 H-bonding, D155/D3.32 salt bridge, hydrophobic pocket engagement). These structural comparisons were the primary basis for evaluating model accuracy.

2.8. Calcium Mobilization Assays

To provide a coarse experimental readout of ligand activity, Fluo-2 AM calcium assays (ION Biosciences, San Marcos, TX, USA) were conducted using HEK293T cells transiently transfected with human 5-HT2A receptor using Lipofectamine 3000 (Waltham, MA, USA). Cells were plated in black, flat clear bottom, chimney-style 96-well CELLSTAR plates and assayed in a Molecular Devices FlexStation plate reader (San Jose, CA, USA).
All ligands except mescaline were tested across a 10−4.5 to 10−10 M dilution series, covering the range over which 5-HT2A agonists typically exert measurable physiological or clinically relevant signaling responses, which were normalized to serotonin (5-HT), shown in Figure 6 [16]. Although binding assays would provide a more direct comparison to computational affinity predictions, calcium mobilization was used as an accessible functional screen to loosely contextualize ligand activity relative to predicted binding strengths.

3. Results

3.1. RMSD and Pose Evaluation Framework

Direct whole-ligand RMSD comparisons were not appropriate because several of the tested analogues were referenced against cryo-EM structures containing related but non-identical ligands. To provide a consistent metric, RMSD values were instead computed between fixed anchor atoms on each ligand’s core scaffold: for tryptamines, the upper carbon bridging the indole 5- and 6-membered rings; for mescaline, the carbon bearing the longest aliphatic substituent. These anchors were selected as positions close to the center of the molecule that would best reflect the positional differences seen in overall placement. Selecting peripheral atoms at the end of the extended chains, by contrast, output a more dramatic quantitative difference for inverted ligands that still had their core roughly in the expected region. Although simplified, this method allowed a uniform numerical comparison across structurally diverse analogues and reliably reflected the magnitude of spatial displacement observed visually. Visual inspection was used in parallel to identify whether a ligand’s predicted orientation was plausible in the context of the experimentally observed poses.
It was expected that because cryo-EM structures with separate ligands represent slightly different activation states, choice of receptor structure could influence computational docking outcomes. However, early tests revealed that using the different receptor files prepared from the DMT-bound, psilocin-bound, and mescaline-bound cryo-EM structures resulted in negligible differences in predicted orientations, so all docking comparisons were performed using models run in the psilocin-bound receptor (PDB 9AS7) to standardize ligand-to-receptor geometry. This selection was not relevant to Boltz-2, as it folds a protein from a provided amino acid sequence, rather than docking in a provided 3D structure [3]. Importantly, this may result in a notable advantage in accuracy for Boltz-2, as the receptor is co-folded with the ligand in each experiment, rather than docking repetitively in the same prepared structure.

3.2. Overall Pose Accuracy Across Methods

Accuracies for all models are quantitatively summarized in Table 1. Across all seven ligands, Boltz-2 consistently produced the most accurate and most stable binding poses. In nearly every case, all five Boltz-2 outputs superimposed closely with one another and fell near the experimentally observed binding mode (Figure 7). RMSD values (0.38–1.05 Å) were uniformly low. No catastrophic pose failures were observed, even for the more bulky or flexible tryptamines (DiPT, 4-HO-DiPT), where both other methods struggled significantly. This suggests that co-folding provides a substantial advantage in flexible GPCR pockets where classical rigid-receptor docking can often fail [17].
AutoDock Vina, run through AMDock, showed much greater variance across its ten output poses. For most ligands, the top-ranked prediction was also the closest to the cryo-EM reference, with the majority producing a few reasonable poses in the upper half of the output ranking list (Figure 8). For more rigid molecules like 2C-B or DOB(-), Vina slightly outperformed Uni-Mol and occasionally approached Boltz-2’s accuracy. However, the tertiary tryptamines DiPT and 4-HO-DiPT were notable exceptions. For these two ligands, the correct pose appeared only among lower-ranked predictions (#10 for DiPT; #8 for 4-HO-DiPT). When these corrected poses were used, RMSDs improved to 1.236 Å (DiPT) and 2.74 Å (4-HO-DiPT), but the majority of Vina-generated poses for these ligands were misplaced, typically clustered above the orthosteric pocket rather than entering it, suggesting steric trapping in the vestibule.
This behavior aligns with known limitations: sometimes Vina’s scoring function does not reliably penalize incorrect macro-orientations, and its sampling efficiency decreases sharply as ligand flexibility increases [18]. Physics-based docking in general can struggle with a lack of receptor flexibility and an inability to correctly sample large rotatable groups and accommodate steric crowding [19].
Uni-Mol Docking v2 often produced ligand placements within the correct pocket but with incorrect orientations, sometimes inverted or rotated relative to the expected pose. This likely arises from orientation ambiguity in Uni-Mol’s SE(3)-equivariant model architecture: without explicit energetic penalties, the model often scores reversed frames similarly to correct frames [20,21]. Although RMSDs for smaller ligands (mescaline, DMT, psilocin, 2C-B, DOB(-)) were within a reasonable range (1–3 Å), Uni-Mol also substantially mispositioned the highly flexible tryptamines. DiPT, in particular, failed to enter the pocket entirely and remained trapped higher in the receptor vestibule, with 4-HO-DiPT landing in the pocket upside-down. Generally, the alignments were less accurate than Boltz-2, and similar to, or slightly worse than, the top predicted poses from AutoDock Vina.

3.3. Per-Ligand Performance Summaries

3.3.1. Psilocin (Cryo-EM Ligand)

All three methods produced plausible poses with small deviations. Boltz-2 displayed the narrowest variance and closest overlay, but all fell within a similar spatial window (Figure 9).

3.3.2. DMT (Cryo-EM Ligand)

All methods performed well. Interestingly, the Vina and Uni-Mol poses converged to almost the same location despite being generated through unrelated scoring functions. Boltz-2 again remained closest to the experimental structure.

3.3.3. Mescaline (Cryo-EM Ligand)

All three methods placed the ligand correctly, with evenly distributed small deviations. This mirror’s mescaline’s relatively rigid scaffold and simple binding geometry.

3.3.4. 2C-B

All programs placed 2C-B within the correct binding region. Uni-Mol placed the ligand backwards relative to the cryo-EM-consistent orientation, whereas Vina and Boltz-2 retained the correct facing.

3.3.5. DOB(-)

All methods correctly located the general pose with minor spatial variance. Boltz-2 again produced the most consistent overlays.

3.3.6. DiPT

Boltz-2 placed DiPT accurately with tightly clustered outputs. Both Vina and Uni-Mol struggled significantly. Most Vina poses failed to fully enter the pocket, and Uni-Mol positioned the ligand too high in the vestibule. Only Vina’s lowest-ranked pose (#10) approximated the correct orientation (Figure 9).

3.3.7. 4-HO-DiPT

Uni-Mol produced an upside-down orientation, while Vina required deep ranking (#8) to locate the correct pose. Boltz-2 again produced a coherent and accurate placement. These failures for DiPT and 4-HO-DiPT likely reflect steric restrictions, which classical docking and AI-guided rigid docking cannot fully resolve, whereas co-folding (Boltz-2) naturally adapts receptor conformation to the ligand surface (Figure 9).

3.4. Affinity Predictions

AutoDock Vina predictions showed substantially greater variability than Boltz-2, both in pose accuracy and in predicted Ki values. Although AMDock’s conversion of Vina binding energies into predicted Ki offered a more interpretable scale, the resulting values still failed to track with either structural plausibility or experimental EC50 potencies. AutoDock generally placed the ligand in the correct pocket, but these lowest-energy poses produced predicted Ki values that collapsed into a narrow band that did not meaningfully correlate with EC50 measurements or with Boltz-2 binding predictions.
Boltz-2 produced affinity estimates as single predicted µM values per ligand. While these values reflect a model trained on mixed Ki/Kd/IC50 data and are therefore inherently qualitative, rank ordering did appear to partially align with the hierarchy of known functional potency. 2C-B and DOB(-) were predicted as the strongest binders, consistent with known high potency at 5-HT2A [22]. Tertiary tryptamines (DiPT, 4-HO-DiPT) were predicted in the mid nanomolar–low micromolar range, and mescaline was predicted weakest, which is also consistent with pharmacological expectations [23]. At 21 µM, this prediction was unusually weak, and could indicate limitations of the affinity model for small, highly polar ligands.

3.5. Experimental Comparison

All ligands except mescaline were screened experimentally in calcium assay, with the strongest predicted binders (2C-B, DOB(-)) corresponding to the most potent calcium signaling responses (Emax 70–84%, EC50 ~1–1.5 nM). Tertiary tryptamines DiPT and 4-HO-DiPT demonstrated intermediate potencies (Emax 71–76%, EC50 ~50–100 nM) weaker than the phenethylamines, with DMT and psilocin landing in the 10–50 nM range. Though these assays screened potency rather than measuring affinity, a loose linear trend could be observed between Boltz-2 affinity estimates and measured EC50. Given the data heterogeneity, this should not be overinterpreted, and reflects coarse ligand potency categories rather than true quantitative prediction.

3.6. Trend Summary

With the lowest RMSD across 7 ligands, and no catastrophic failures, Boltz-2’s co-folding approach most reliably recovers cryo-EM binding orientations, even for more sterically bulky or highly flexible ligands. Classical docking (AutoDock Vina) was reasonably accurate for rigid ligands but exhibiting sampling failure in the sterically constrained GPCR pocket, showing pose instability for flexible tertiary tryptamines. Uni-Mol Docking v2 performs similarly to Vina for smaller ligands but occasionally showed a tendency toward pose inversion or displacement, especially for asymmetric ligands. Boltz-2 affinity estimates appear to qualitatively track experimental calcium potencies better than classical docking scores, despite limitations as a fused Ki/Kd/IC50 model.

4. Discussion

This exploratory study offers a side-by-side comparison of three fundamentally different computational paradigms applied to a pharmacologically important and structurally complex GPCR system. Although not designed as a strict benchmarking exercise, the patterns that emerged across the ligand series reveal both the promise and current limitations of modern AI-augmented modeling in serotonergic drug discovery.
The superior pose accuracy of Boltz-2 likely reflects its architectural advantage: it performs joint cofolding of the receptor and ligand using deep geometric diffusion models trained on large multimodal structural corpora [3]. Rather than threading a rigid ligand into a mostly rigid protein, Boltz-2 optimizes both partners simultaneously, learning physically plausible conformational couplings. This is particularly advantageous in GPCRs, where the binding pocket is deep, hydrophobic, and conformationally flexible. This global, end-to-end learned representation likely explains Boltz-2’s capacity to place larger or more substituted tryptamines (e.g., DiPT, 4-HO-DiPT) into the correct orthosteric region even when steric congestion or pocket narrowing misled the other methods. The near-identical alignment of its five output poses further supports the stability and consistency of its internal learned scoring objective.
In contrast, AutoDock Vina relies on an empirical scoring function and stochastic search procedure. Its success for smaller, well-studied ligands such as psilocin, DMT, and mescaline reflects the robustness of classical docking when ligand flexibility and steric bulk are limited [24]. However, its difficulty with DiPT and 4-HO-DiPT highlight well-known limitations: rigid or semi-rigid receptor models underestimate GPCR plasticity, torsional sampling becomes inefficient for branched ligands, and the scoring function is not equipped to accurately evaluate long-range polar interactions or conformational accommodation within the pocket [25,26]. The need to rely on lower-ranked poses underscores the stochastic nature of Vina’s search procedure, where viable solutions may exist but are not prioritized by the scoring model.
Uni-Mol Docking v2, being an AI-based docking engine, occupies a conceptual middle ground between the two other approaches. It does not cofold receptor and ligand but instead uses a transformer-like architecture with 3D equivariant attention to generate ligand conformers and protein representations [5]. This gives it higher expressive capacity than classical docking, but it remains susceptible to local misorientations. The model appears to correctly identify the binding region but sometimes misclassifies the optimal orientation, a limitation likely tied to its reliance on statistical patterns from training data rather than a physically grounded energy landscape. Nonetheless, for several ligands Uni-Mol matched or outperformed Vina, suggesting that generative geometric reasoning is beginning to bridge the gap between traditional and end-to-end learned approaches.
For the ligands in the solved cryo-EM structures (psilocin, DMT, mescaline), all models performed well, with small and roughly even RMSD spreads. This could indicate that when the target–ligand system closely resembles training data and structural precedents, AI-based models and classical docking converge on similar solutions. For the ligands with increased bulk or unusual substitutions, Boltz-2 maintained accuracy, with less reliable outputs from Uni-Mol Docking v2 and AutoDock Vina. This pattern strongly supports the interpretation that ligand-specific steric accommodation and receptor flexibility, processes learned implicitly by Boltz-2, are critical for accurate GPCR modeling and are not well captured by rigid docking.
Among the phenethylamines (2C-B, DOB(-), mescaline), all methods placed the ligands in the correct pocket, though Uni-Mol occasionally reversed orientation (Figure 10). These observations suggest that even relatively simple symmetric scaffolds can challenge models lacking explicit energetic evaluation.
Molecular docking originated as a physics-based search and scoring problem, with the goal of sampling ligand conformations in a protein binding site and evaluating their complementarity using empirical or semiempirical scoring functions. These traditional methods directly explore the geometric search space of ligand–protein interactions and compute energies based on steric contacts and approximate force fields. They benefit from interpretability and a firm physical grounding, and they do not require training data, but they can suffer from limited sampling efficiency and accuracy when the search space is large or when receptor flexibility plays a significant role. This limitation is particularly acute in blind docking scenarios, where binding sites are unknown and the entire protein surface must be explored, leading to high computational cost and reduced reliability due to the expansive conformational search space.
To address these shortcomings, machine learning-based docking approaches have been developed. Data-driven models, such as convolutional neural networks or graph neural networks, predict binding modes or identify likely binding pockets based on patterns learned from large structural databases of protein–ligand complexes. Machine learning-based blind docking techniques can dramatically improve speed and early binding site identification relative to classical blind docking, especially when scanning entire protein surfaces without prior site knowledge. However, these approaches can exhibit overfitting to training distributions and reduced performance on novel targets not well represented in their training sets, highlighting a trade-off between learned generalization and physical interpretability.
The field is now moving toward hybrid physics–AI strategies that integrate the strengths of both paradigms. In hybrid models, physics-based sampling and scoring are augmented with machine learning components that guide search, refine scoring, or detect pockets, aiming to improve accuracy and efficiency simultaneously. Early results suggest that these integrated methods can outperform purely classical or purely data-driven workflows by combining physical realism and learned structural priors, although they still face challenges in balancing computational cost, generalizability, and interpretability [27].

4.1. Relationship Between Computational Predictions and Experimental Data

Predicted Boltz-2 affinity scores showed a weak but noticeably more coherent trend with experimental EC50 values than AutoDock’s Ki predictions, although neither metric should be overinterpreted. The EC50 data reflect downstream signaling efficacy rather than binding affinity, and the sample size here was small. Nevertheless, the partial alignment between Boltz-2 scores and biological potency may indicate that learned energy landscapes capture subtle structure–activity features that classical scoring functions overlook.

4.2. Limitations of This Exploratory Study

Several constraints should be emphasized:
  • RMSD measurements were based on core-atom alignment rather than whole-ligand RMSD due to scaffold heterogeneity.
  • Only one functional assay was used as a general screen (calcium mobilization), which does not directly report affinity.
  • For Uni-Mol Docking v2, the Bohrium web interface limited direct control over sampling depth. It’s possible that with more structures outputted in a rank order, as was the case for AutoDock, there would have been more accurate predictions to select from.
  • For Boltz-2, the affinity score’s biochemical interpretation remains unclear given its multi-objective training.
These constraints reflect the intended exploratory, technology-survey nature of this work, rather than a controlled benchmarking experiment.

5. Conclusions

This study provides an early comparative snapshot of three distinct modeling paradigms applied to a clinically relevant GPCR system. Boltz-2’s consistent accuracy across both canonical and structurally challenging ligands illustrates the growing strength of end-to-end AI cofolding methods, particularly for deep, flexible binding pockets where classical approaches struggle. Uni-Mol Docking v2, while not achieving Boltz-2’s consistency, demonstrates the promise of generative equivariant architectures for docking tasks and may improve with expanded training or increased sampling. AutoDock Vina, despite limitations, remains competitive for simple ligands and provides a valuable classical contrast against which AI-based methods can be contextualized.
These findings reinforce a broader message regarding the integration of AI in biochemistry and drug discovery: as structural prediction models become increasingly sophisticated, careful experimental anchoring (here via cryo-EM structures and coarse functional assays) remains essential. This work serves as a small but instructive probe into how modern AI tools behave on a real pharmacological problem, illustrating both their transformative potential and the continued need for rigorous, mechanistically informed evaluation.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/aichem1010004/s1, File S1: SMILES strings for all ligands, File S2: links to the programs used.

Author Contributions

Conceptualization, B.R.C. and C.D.N.; methodology, B.R.C.; formal analysis, B.R.C.; investigation, B.R.C.; resources, C.D.N.; data curation, B.R.C.; writing—original draft preparation, B.R.C.; writing—review and editing, B.R.C. and C.D.N.; visualization, B.R.C.; supervision, C.D.N.; funding support, C.D.N. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by 2A Biosciences, New York, USA.

Data Availability Statement

Raw data such as calcium assay readout was omitted for brevity, and is available from the authors upon request.

Acknowledgments

Molecular graphics and analyses performed with UCSF ChimeraX, developed by the Resource for Biocomputing, Visualization, and Informatics at the University of California, San Francisco, with support from National Institutes of Health R01-GM129325 and the Office of Cyber Infrastructure and Computational Biology, National Institute of Allergy and Infectious Diseases. All software used is provided free of charge for academic use from the developers listed. Boltz-2 and Uni-Mol Docking v2 are available under MIT license, Autodock Vina under Apache license, AMDock under GPLv3. The authors used Chat GPT-5 Mini for early formatting of references and recommendation of relevant modern literature.

Conflicts of Interest

C.D.N. has a sponsored research agreement with 2A Biosciences and is on its board of directors. The funders had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results.

Abbreviations

The following abbreviations are used in this manuscript:
2CB4-bromo-2,5-dimethoxyphenethylamine
AIArtificial Intelligence
GPCRG Protein-Coupled Receptor
DMTN,N-Dimethyltryptamine
CNNConvolutional Neural Network
DOB2,5-Dimethoxy-4-bromoamphetamine
DiPTN,N-Diisopropyltryptamine
RMSDRoot mean square deviation

References

  1. Gumpper, R.H.; Jain, M.K.; Kim, K.; Sun, R.; Sun, N.; Xu, Z.; DiBerto, J.F.; Krumm, B.E.; Kapolka, N.J.; Kaniskan, H.Ü.; et al. The structural diversity of psychedelic drug actions revealed. Nat. Commun. 2025, 16, 2734. [Google Scholar] [CrossRef] [PubMed]
  2. López-Giménez, J.F.; González-Maeso, J. Hallucinogens and serotonin 5-HT2A receptor-mediated signaling pathways. Curr. Top. Behav. Neurosci. 2018, 36, 45–73. [Google Scholar] [CrossRef] [PubMed]
  3. Passaro, S.; Corso, G.; Wohlwend, J.; Reveiz, M.; Thaler, S.; Somnath, V.R.; Getz, N.; Portnoi, T.; Roy, J.; Stark, H.; et al. Boltz-2: Towards accurate and efficient binding affinity prediction. bioRxiv 2025. [Google Scholar] [CrossRef] [PubMed]
  4. Wohlwend, J.; Corso, G.; Passaro, S.; Getz, N.; Reveiz, M.; Leidal, K.; Swiderski, W.; Atkinson, L.; Portnoi, T.; Chinn, I.; et al. Boltz-1 democratizing biomolecular interaction modeling. bioRxiv 2025, 2024, 624167. [Google Scholar]
  5. Alcaide, E.; Gao, Z.; Ke, G.; Li, Y.; Zhang, L.; Zheng, H.; Zhou, G. Uni-mol docking v2: Towards realistic and accurate binding pose prediction. In International Conference on Artificial Neural Networks; Springer Nature: Cham, Switzerland, 2025; pp. 34–41. [Google Scholar]
  6. Eberhardt, J.; Santos-Martins, D.; Tillack, A.F.; Forli, S. AutoDock Vina 1.2.0: New docking methods, expanded force field, and python bindings. J. Chem. Inf. Model. 2021, 61, 3891–3898. [Google Scholar] [CrossRef]
  7. Trott, O.; Olson, A.J. AutoDock Vina: Improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading. J. Comput. Chem. 2010, 31, 455–461. [Google Scholar] [CrossRef]
  8. McNutt, A.T.; Francoeur, P.; Aggarwal, R.; Masuda, T.; Meli, R.; Ragoza, M.; Sunseri, J.; Koes, D.R. GNINA 1.0: Molecular docking with deep learning. J. Cheminform. 2021, 13, 43. [Google Scholar] [CrossRef]
  9. Tripathi, A.; Suri, K.; Murugan, N.A. Assessing the accuracy of binding pose prediction for kinase proteins and 7-azaindole inhibitors: A study with AutoDock4, Vina, DOCK 6, and GNINA 1.0. RSC Adv. 2025, 15, 47051–47065. [Google Scholar] [CrossRef]
  10. Hanwell, M.D.; Curtis, D.E.; Lonie, D.C.; Vandermeersch, T.; Zurek, E.; Hutchison, G.R. Avogadro: An advanced semantic chemical editor, visualization, and analysis platform. J. Cheminform. 2012, 4, 17. [Google Scholar] [CrossRef]
  11. Pettersen, E.F.; Goddard, T.D.; Huang, C.C.; Meng, E.C.; Couch, G.S.; Croll, T.I.; Morris, J.H.; Ferrin, T.E. UCSF ChimeraX: Structure visualization for researchers, educators, and developers. Protein Sci. 2021, 30, 70–82. [Google Scholar] [CrossRef]
  12. Meng, E.C.; Goddard, T.D.; Pettersen, E.F.; Couch, G.S.; Pearson, Z.J.; Morris, J.H.; Ferrin, T.E. UCSF ChimeraX: Tools for structure building and analysis. Protein Sci. 2023, 32, e4792. [Google Scholar] [CrossRef] [PubMed]
  13. The PyMOL Molecular Graphics System, version 2.1; Schrodinger, LLC.: New York, NY, USA, 2018.
  14. Valdes-Tresanco, M.S.; Valdes-Tresanco, M.E.; Valiente, P.A.; Moreno, E. AMDock: A versatile graphical tool for assisting molecular docking with Autodock Vina and Autodock4. Biol. Direct 2020, 15, 12. [Google Scholar] [CrossRef] [PubMed]
  15. Mirdita, M.; Schütze, K.; Moriwaki, Y.; Heo, L.; Ovchinnikov, S.; Steinegger, M. ColabFold: Making protein folding accessible to all. Nat. Methods 2022, 19, 679–682. [Google Scholar] [CrossRef] [PubMed]
  16. Nichols, D.E. Structure–activity relationships of serotonin 5-HT2A agonists. Wiley Interdiscip. Rev. Membr. Transp. Signal. 2012, 1, 559–579. [Google Scholar] [CrossRef]
  17. Jain, A.N. Effects of protein conformation in docking: Improved pose prediction through protein pocket adaptation. J. Comput.-Aided Mol. Des. 2009, 23, 355–374. [Google Scholar] [CrossRef]
  18. Xu, M.; Shen, C.; Yang, J.; Wang, Q.; Huang, N. Systematic investigation of docking failures in large-scale structure-based virtual screening. ACS Omega 2022, 7, 39417–39428. [Google Scholar] [CrossRef]
  19. Totrov, M.; Abagyan, R. Flexible ligand docking to multiple receptor conformations: A practical alternative. Curr. Opin. Struct. Biol. 2008, 18, 178–184. [Google Scholar] [CrossRef]
  20. Zhou, G.; Gao, Z.; Ding, Q.; Zheng, H.; Xu, H.; Wei, Z.; Zhang, L.; Ke, G. Uni-mol: A universal 3D molecular representation learning framework. In Proceedings of the Eleventh International Conference on Learning Representations, Kigali, Rwanda, 1–5 May 2023. [Google Scholar]
  21. Buttenschoen, M.; Morris, G.M.; Deane, C.M. PoseBusters: AI-based docking methods fail to generate physically valid poses or generalise to novel sequences. Chem. Sci. 2024, 15, 3130–3139. [Google Scholar] [CrossRef]
  22. Ray, T.S. Psychedelics and the human receptorome. PLoS ONE 2010, 5, e9019. [Google Scholar] [CrossRef]
  23. Rickli, A.; Moning, O.D.; Hoener, M.C.; Liechti, M.E. Receptor interaction profiles of novel psychoactive tryptamines compared with classic hallucinogens. Eur. Neuropsychopharmacol. 2016, 26, 1327–1337. [Google Scholar] [CrossRef]
  24. Huang, S.Y. Comprehensive assessment of flexible-ligand docking algorithms: Current effectiveness and challenges. Brief. Bioinform. 2018, 19, 982–994. [Google Scholar] [CrossRef]
  25. Kufareva, I.; Katritch, V.; Stevens, R.C.; Abagyan, R. Advances in GPCR modeling evaluated by the GPCR Dock 2013 assessment: Meeting new challenges. Structure 2014, 22, 1120–1139. [Google Scholar] [CrossRef]
  26. Warren, G.L.; Andrews, C.W.; Capelli, A.M.; Clarke, B.; LaLonde, J.; Lambert, M.H.; Head, M.S. A critical assessment of docking programs and scoring functions. J. Med. Chem. 2006, 49, 5912–5931. [Google Scholar] [CrossRef]
  27. Roomi, M.S.; Culletta, G.; Longo, L.; de Azevedo, W.F., Jr.; Perricone, U.; Tutone, M. Docking in the Dark: Insights into Protein–Protein and Protein–Ligand Blind Docking. Pharmaceuticals 2025, 18, 1777. [Google Scholar] [CrossRef]
Figure 1. Test series includes three resolved structures (psilocin, DMT, mescaline) and four closely related analogues.
Figure 1. Test series includes three resolved structures (psilocin, DMT, mescaline) and four closely related analogues.
Aichem 01 00004 g001
Figure 2. The docking region (purple outline shown here in PyMOL) was defined generously, centered on cryo-EM ligand position in the binding pocket.
Figure 2. The docking region (purple outline shown here in PyMOL) was defined generously, centered on cryo-EM ligand position in the binding pocket.
Aichem 01 00004 g002
Figure 3. AMDock is a GUI that streamlines ligand docking with AutoDock Vina.
Figure 3. AMDock is a GUI that streamlines ligand docking with AutoDock Vina.
Aichem 01 00004 g003
Figure 4. AutoDock Vina outputs 10 rank-ordered predictions by default via AMDock.
Figure 4. AutoDock Vina outputs 10 rank-ordered predictions by default via AMDock.
Aichem 01 00004 g004
Figure 5. The Bohrium app offers Uni-Mol Docking service in a web browser.
Figure 5. The Bohrium app offers Uni-Mol Docking service in a web browser.
Aichem 01 00004 g005
Figure 6. Representative outputs from a calcium assay. All ligands were tested in duplicate in two separate experiments, normalized to serotonin, and fit with non-linear regression.
Figure 6. Representative outputs from a calcium assay. All ligands were tested in duplicate in two separate experiments, normalized to serotonin, and fit with non-linear regression.
Aichem 01 00004 g006
Figure 7. (a) All 5 boltz predictions for 2CB overlaid with one another are nearly indistinguishable. (b) Comparing these to the resolved structure for mescaline (green) indicates accuracy, not just precision. Models prepared in UCSF ChimeraX.
Figure 7. (a) All 5 boltz predictions for 2CB overlaid with one another are nearly indistinguishable. (b) Comparing these to the resolved structure for mescaline (green) indicates accuracy, not just precision. Models prepared in UCSF ChimeraX.
Aichem 01 00004 g007
Figure 8. Autodock outputs 10 rank-ordered predictions, with the first generally being the most accurate, and the top half generally landing somewhere in the expected binding pocket. Shown above are the 10 predictions for psilocin (yellow), with the resolved structure (green) for comparison. Model prepared in UCSF ChimeraX.
Figure 8. Autodock outputs 10 rank-ordered predictions, with the first generally being the most accurate, and the top half generally landing somewhere in the expected binding pocket. Shown above are the 10 predictions for psilocin (yellow), with the resolved structure (green) for comparison. Model prepared in UCSF ChimeraX.
Aichem 01 00004 g008
Figure 9. Resolved structure of psilocin in green, Autodock outputs in yellow, Uni-Mol in orange, Boltz in red. (a): all models predicted psilocin with relative accuracy, showing an even spread. (b): DiPT was problematic for Uni-Mol and Vina alike, both getting stuck outside of the binding pocket. Only pose 10 for Vina made it to the expected orientation, with pose 1 landing nearly identical to Uni-Mol’s top prediction, ~10 Angstroms away. (c): 4-HO-DiPT also saw Vina (pose 1) getting stuck in the top of the receptor, with Uni-Mol making it into the binding pocket, but upside down. Models prepared in UCSF ChimeraX.
Figure 9. Resolved structure of psilocin in green, Autodock outputs in yellow, Uni-Mol in orange, Boltz in red. (a): all models predicted psilocin with relative accuracy, showing an even spread. (b): DiPT was problematic for Uni-Mol and Vina alike, both getting stuck outside of the binding pocket. Only pose 10 for Vina made it to the expected orientation, with pose 1 landing nearly identical to Uni-Mol’s top prediction, ~10 Angstroms away. (c): 4-HO-DiPT also saw Vina (pose 1) getting stuck in the top of the receptor, with Uni-Mol making it into the binding pocket, but upside down. Models prepared in UCSF ChimeraX.
Aichem 01 00004 g009
Figure 10. DOB(-) predictions aligned with the resolved structure of mescaline (green). While Boltz-2 (red) is closest, and Autodock (yellow) slightly tilted, Uni-Mol (orange) also arrived in the correct region, but with the extended chain flipped to the other direction, resulting in an RMSD of 2.791, shown between the anchor atom of this Uni-mol prediction and that of the resolved structure. Receptor structures are shown for the resolved receptor (green, background) and the Boltz-2 folded receptor (red, background). Model prepared in UCSF ChimeraX.
Figure 10. DOB(-) predictions aligned with the resolved structure of mescaline (green). While Boltz-2 (red) is closest, and Autodock (yellow) slightly tilted, Uni-Mol (orange) also arrived in the correct region, but with the extended chain flipped to the other direction, resulting in an RMSD of 2.791, shown between the anchor atom of this Uni-mol prediction and that of the resolved structure. Receptor structures are shown for the resolved receptor (green, background) and the Boltz-2 folded receptor (red, background). Model prepared in UCSF ChimeraX.
Aichem 01 00004 g010
Table 1. Summarized values for RMSD, affinity predictions and experimentally determined potencies.
Table 1. Summarized values for RMSD, affinity predictions and experimentally determined potencies.
CompoundRMSD [Boltz-2]RMSD [Vina]RMSD [Uni-Mol]Boltz-2 Affinity Prediction [μM]AutoDock Vina Ki Prediction [μM]Normalized LogEC50EC50Emax
Psilocin0.6910.9960.9740.1520.36−7.987510.3 nM50.835
DMT0.5531.6741.5250.6617.2−7.313548.7 nM39.61
4-HO-DiPT1.0542.741.9143.817.2−7.303549.8 nM76.61
DiPT0.9481.23610.2434.320.36−7.0197.7 nM71.66
DOB(-)0.8241.6672.7910.02820.36−8.92451.19 nM83.845
2C-B0.6480.9053.0840.01656.05−8.82751.49 nM69.165
Mescaline0.3761.2831.2172139.99
Serotonin0.9151.9518.0020.009617.2−8.5053.13 nM100
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Cummins, B.R.; Nichols, C.D. Assessing Modern AI-Driven Protein-Ligand Modeling with Phenethylamine and Tryptamine Psychedelics. AI Chem. 2026, 1, 4. https://doi.org/10.3390/aichem1010004

AMA Style

Cummins BR, Nichols CD. Assessing Modern AI-Driven Protein-Ligand Modeling with Phenethylamine and Tryptamine Psychedelics. AI Chemistry. 2026; 1(1):4. https://doi.org/10.3390/aichem1010004

Chicago/Turabian Style

Cummins, Benjamin R., and Charles D. Nichols. 2026. "Assessing Modern AI-Driven Protein-Ligand Modeling with Phenethylamine and Tryptamine Psychedelics" AI Chemistry 1, no. 1: 4. https://doi.org/10.3390/aichem1010004

APA Style

Cummins, B. R., & Nichols, C. D. (2026). Assessing Modern AI-Driven Protein-Ligand Modeling with Phenethylamine and Tryptamine Psychedelics. AI Chemistry, 1(1), 4. https://doi.org/10.3390/aichem1010004

Article Metrics

Back to TopTop