Abstract
High-content screening (HCS) is a useful phenotypic drug discovery technology that combines automated microscopy, image analysis, and high-throughput experimentation to comprehensively characterize biological responses to diverse perturbations. This review summarizes the methodological fundamentals of high-content analysis, including image preprocessing, cell segmentation, feature processing, and downstream analysis, as well as the diverse phenotypic datasets generated from different biological models, perturbation strategies, and staining approaches. Recent advances in artificial intelligence, particularly deep learning, have improved cell segmentation, image representation learning, and phenotypic profiling, enabling more accurate and scalable analysis of HCS data. We further highlight emerging applications of AI-powered HCS in pharmaceutical research, with a particular focus on the discovery of bioactive compounds from natural sources. Finally, we discuss current challenges and future perspectives, including the construction of large-scale phenotypic databases, the integration of AI throughout the screening workflow, and the development of intelligent screening platforms. These advances are expected to accelerate phenotype-driven drug discovery and promote innovation in natural product research.
1. Introduction
High-content screening is an image-based phenotypic screening technology that combines automated microscopy, fluorescent labeling, and quantitative image analysis to systematically characterize cellular responses to genetic, chemical, or environmental perturbations [1,2,3,4,5]. By simultaneously measuring multiple morphological, spatial, and functional features, HCS provides a multidimensional representation of cellular phenotypes that extends far beyond conventional endpoint assays. Modern HCS platforms integrate high-throughput imaging systems with advanced computational pipelines, enabling the acquisition and analysis of millions of cellular images across diverse biological models [6,7,8]. As a result, HCS has become an indispensable tool in numerous fields, including drug discovery, functional genomics, toxicology, and rare disease research [9,10,11,12,13,14,15].
Unlike traditional target-based screening approaches, which typically focus on a single molecular target or biochemical readout, HCS can capture holistic cellular responses and allow the identification of bioactive compounds through phenotypic changes. This capability is particularly valuable for natural product research. Natural products and herbal medicines contain structurally diverse molecules that often exert their biological activities through multiple targets and pathways. Such complexity presents significant challenges for conventional screening methods, which frequently fail to capture the full spectrum of biological activities [16,17,18,19,20]. In contrast, HCS enables unbiased phenotypic profiling and provides rich cellular information, including changes in morphology, organelle organization, signaling pathways, and cellular functions, thereby facilitating the discovery of bioactive compounds.
The development of HCS can be traced back to the emergence of automated fluorescence microscopy and image analysis technologies in the late 1990s and early 2000s. Early HCS studies relied heavily on manually designed image features and classical machine learning algorithms for phenotype classification. While these approaches significantly improved the throughput of phenotypic screening, their performance was constrained by limited feature representation and the complexity of biological images. Over the past decade, rapid advances in artificial intelligence (AI), particularly deep learning, have transformed the field. Convolutional neural networks, vision transformers, self-supervised learning frameworks, and generalist models have demonstrated unprecedented capabilities in image segmentation, feature extraction, and phenotype classification [21,22,23,24,25]. These advances have enabled the extraction of biologically meaningful information directly from raw microscopy images, reducing the dependence on handcrafted features and improving analytical accuracy and scalability.
The integration of AI with HCS has been especially transformative for the discovery of bioactive compounds from natural sources. Natural product libraries and herbal extracts often contain thousands of chemically diverse constituents with unknown biological activities. AI-driven HCS provides a useful platform for systematically interrogating these complex resources, prioritizing active fractions, predicting mechanisms of action, identifying potential targets, and uncovering novel therapeutic opportunities [6,7,8]. Despite these advances, several challenges remain. The high dimensionality and heterogeneity of imaging data, batch effects [26,27] across experimental platforms, limited availability of annotated datasets, and the interpretability of AI models continue to hinder broader adoption. In addition, the complexity of natural product mixtures presents unique difficulties for data analysis, activity attribution, and experimental standardization. Currently, HCS has attracted increasing attention in natural product research. A recent review [28] has summarized the applications of HCS technology in traditional Chinese medicine (TCM) research, with a particular focus on the implementation of HCS studies, optimization of experimental parameters, and major application areas. However, it does not provide a comprehensive discussion of HCA from a computational perspective or cover recent advances in AI-based HCA. In this review, we therefore focus on the methodological advances in HCA and the applications of AI-based HCA in bioactive compound discovery. Otherwise, we also primarily focus on traditional Chinese medicine (TCM) and plant-derived natural products, which represent important sources of bioactive compounds. Although natural sources also include microbial, fungal, marine, and animal sources, these areas are beyond the scope of our expertise and are therefore not covered in this review.
This review provides a comprehensive overview of recent advances in AI-driven high-content analysis (HCA). Section 2 outlines the methodological fundamentals of HCA, diverse phenotypic data generated by HCS experiments, and emerging AI-powered analytical approaches. Section 3 highlights recent advances in deep learning-based intelligent HCS and its applications in the discovery of bioactive compounds from natural sources. Finally, Section 4 discusses the current challenges and future perspectives of AI-powered high-content screening in natural product research, with a focus on phenotypic database construction, AI integration across the screening workflow, and the development of intelligent screening platforms.
2. Methodological Fundamentals of HCA
2.1. Typical High-Content Analysis
High-content analysis refers to the process of converting digital microscopy images into quantitative measurements for each experimental perturbation and performing downstream analyses on these measurements. The resulting data can be organized into a matrix in which each row represents a perturbation and each column corresponds to a feature extracted from the observed phenotype. A typical HCA workflow (Figure 1) consists of four major steps: image preprocessing, cell segmentation, feature processing, and downstream analysis.
Figure 1.
Overview of the high-content analysis workflow.
The first step is image preprocessing. Due to variations in illumination sources and vignetting effects commonly observed at image edges, acquired images often exhibit non-uniform illumination, resulting in intensity fluctuations of approximately 10–30%, which can affect downstream analyses [29]. Therefore, illumination correction is routinely performed in traditional image analysis pipelines. Illumination correction methods can generally be categorized into three groups. Prospective methods construct correction models using reference images acquired during microscopy imaging and therefore require careful calibration during image acquisition. Retrospective single-image methods generate an independent correction model for each image based solely on the image itself. Although these methods reduce the burden of image acquisition, they may compromise comparability across images. In contrast, retrospective multi-image methods build correction models using all images collected within an experiment, resulting in more robust correction and consequently becoming the most widely adopted approach in biological laboratories. In addition to illumination correction, normalization and standardization are frequently applied. Normalization transforms pixel intensities into a predefined range, typically between 0 and 1. For example, min-max normalization rescales each pixel value by subtracting the minimum pixel intensity and dividing by the intensity range of the image. This procedure reduces variations across images and improves comparability between datasets acquired under different conditions. Standardization, on the other hand, subtracts the mean pixel value and divides by the standard deviation, producing a normalized distribution with zero mean and unit variance. This process mitigates global intensity shifts, balances feature scales, and prevents specific features from disproportionately influencing model training.
The second step is cell segmentation. By identifying and isolating individual cellular instances within an image, cell segmentation enables quantitative analysis at single-cell resolution. Fundamentally, segmentation is a pixel-level classification task. Existing segmentation methods can generally be divided into two categories. The first category comprises model-based methods, which rely on prior knowledge and manual parameter optimization based on visual inspection of segmentation results. Parameters may include expected object size and shape or image histogram distributions. Most model-based methods belong to the category of binary semantic segmentation and generate a binary mask corresponding to the original image, where pixels are classified as either foreground (cell) or background. Common approaches include threshold-based methods such as Otsu thresholding, edge-based methods such as edge detection algorithms, region-based methods including region growing and watershed segmentation, and graph-based methods such as GraphCut [30]. The second category consists of machine learning-based methods. These approaches utilize paired training datasets containing raw images and manually annotated masks to learn segmentation models automatically. Compared with model-based methods, machine learning approaches require fewer manually defined parameters and generally achieve superior performance in complex imaging scenarios [31]. However, early machine learning models often exhibited limited generalizability, requiring extensive manual annotation and programming expertise when applied to new datasets. In contrast, many model-based methods have been integrated into user-friendly software packages such as CellProfiler 3.0 [32] and ImageJ2, making them accessible to researchers without computational backgrounds. Consequently, these approaches have remained widely used in many laboratories.
The third step is feature processing. Features represent quantitative measurements extracted from images and contain information describing cellular phenotypes. Feature processing consists of two key procedures: feature extraction and feature aggregation. Feature extraction strategies can be broadly classified into three categories. The first involves the direct measurement of manually designed features. Building upon previous studies, Carpenter et al. developed CellProfiler, which enables extraction of a wide variety of image features through an intuitive graphical interface. These features can generally be divided into three groups. The first group characterizes full images, including image quality metrics, overall fluorescence intensity, and cell counts. The second group quantifies individual cells, including measurements such as fluorescence intensity, area, size, and circularity. The third group captures microenvironmental information, including neighboring-cell relationships and organelle colocalization patterns. The second strategy integrates feature extraction directly into downstream tasks through end-to-end deep learning frameworks. For example, Chen et al. [33] trained a classification model directly on raw microscopy images to distinguish damaged, healthy, and irrelevant cells, with feature extraction implicitly learned within a VGG-based architecture. The third strategy is based on representation learning [34], in which feature extraction is decoupled from downstream tasks. Rather than relying on manually designed descriptors, representation learning approaches learn generalizable feature representations directly from large-scale datasets. Following feature extraction, the features obtained from individual images must be aggregated. For example, in a typical HCS experiment, a single perturbation may correspond to one well, with four fields of view imaged per well. Under a single-cell analysis framework, features extracted from individual cells are first aggregated into image-level representations, which are then combined across multiple fields of view to generate perturbation-level profiles. Common aggregation strategies include mean profiling, median profiling, and Kolmogorov–Smirnov profiling.
The final step is downstream analysis. Through visualization techniques and machine-learning approaches, extracted phenotypic features can be interpreted and validated to reveal biologically meaningful patterns and relationships. Three major downstream analysis strategies are commonly employed. The first is clustering, which groups perturbations with similar phenotypic profiles in an unsupervised manner. Clustering can not only validate known biological relationships but also uncover previously unrecognized associations among treatments. A widely used method is hierarchical clustering [35], which visualizes similarity matrices as heatmaps and facilitates the identification of patterns across dozens or even hundreds of samples. The second strategy involves dimensionality reduction and visualization. High-dimensional phenotypic features are projected into two- or three-dimensional spaces to facilitate interpretation of the underlying feature distributions. Common dimensionality-reduction methods include principal component analysis (PCA), t-distributed stochastic neighbor embedding (t-SNE) [36], and uniform manifold approximation and projection (UMAP) [37]. The third strategy is classification. Supervised machine-learning models are trained using annotated datasets and subsequently applied to predict phenotypic outcomes for previously unseen perturbations [5]. Commonly used classifiers include support vector machines, random forests, and neural networks.
It should be noted that these steps are not universally required in every HCA workflow and can be modified according to specific research objectives. For example, improvements in imaging quality and algorithmic performance have reduced the necessity of illumination correction in some applications. Similarly, when datasets become extremely large and cell segmentation imposes substantial computational and economic costs, feature extraction can be performed directly on whole images without explicit segmentation. Moreover, morphological similarity may often indicate similar biological effects or phenotypic responses; however, the underlying molecular mechanisms require further experimental validation.
2.2. Diverse Phenotypic Data Generated by HCS
Due to the complexity of the screening workflow, HCS generates highly diverse and complex phenotypic datasets. This diversity is mainly reflected in three aspects: biological models, perturbation reagents, and staining strategies (Figure 2). First, a wide variety of screening models can be employed, spanning multiple biological scales from microscopic cellular systems to mesoscopic tissues and even whole organisms. The most commonly used model is the conventional two-dimensional (2D) cell culture system because of its simplicity and convenience. However, accumulating evidence suggests that cellular responses observed in 2D cultures may differ from those observed in native tissues in vivo [38]. To address this limitation, more physiologically relevant biomimetic models have been developed. For example, micropatterning technologies can control the alignment of cardiomyocytes on engineered substrates, promoting coordinated mechanical contraction-relaxation cycles and electrophysiological activity that more closely resemble native cardiac tissues [39]. More recently, organoid technologies [40,41,42,43] have emerged as useful experimental platforms. Compared with traditional cell culture models, organoids possess unique three-dimensional architectures and cell-type-specific organization, enabling them to recapitulate physiological functions and responses that closely resemble those observed in vivo. Consequently, organoid-based models have been increasingly incorporated into HCS workflows. In addition, small model organisms such as C. elegans [44,45] and zebrafish [46,47] can be utilized for phenotypic screening, allowing biological responses to be monitored at the level of an intact organism.
Figure 2.
Diversity of phenotypic data in HCA.
Second, HCS can employ a diverse range of perturbation reagents, which may be selected either randomly or specifically designed to interrogate particular biological functions. Depending on the experimental design, different types of perturbations can be compared to explore their phenotypic effects. Perturbation libraries may consist of small molecules, including FDA-approved drug collections, natural product libraries, or compound libraries targeting specific diseases or signaling pathways. Alternatively, genetic perturbations can be introduced using siRNA libraries, which reduce gene expression through mRNA degradation or gene silencing, or CRISPR libraries [48], which exploit CRISPR-Cas9-mediated genome editing to induce targeted genetic modifications. Other perturbation resources, such as TCM component libraries and peptide libraries, have also been incorporated into HCS studies, further expanding the scope of phenotypic discovery.
Third, staining strategies are highly diverse and are largely determined by the objectives of the screening campaign. On one hand, fluorescence labeling can be specifically designed to monitor particular biological mechanisms. For example, the DNA damage response protein 53BP1 exhibits a diffuse nuclear distribution under normal conditions but accumulates at sites of DNA double-strand breaks following genotoxic stress. Based on this phenomenon, fluorescent probes targeting 53BP1 have been developed to quantitatively evaluate DNA damage induced by different perturbations [33]. On the other hand, broadly applicable staining panels can be employed to monitor a wide range of cellular processes. Such assays enable the characterization of diverse biological activities, including cell cycle progression, cell adhesion and migration, cellular metabolism, and organelle-specific phenotypic alterations. By combining multiple fluorescent markers within a single experiment, HCS can generate comprehensive phenotypic profiles that capture diverse aspects of cellular behavior.
Together, the diversity of biological models, perturbation reagents, and staining strategies enables HCS to generate rich and multidimensional phenotypic datasets. While this complexity greatly enhances the biological information content of HCS experiments, it also poses substantial challenges for data analysis and interpretation, thereby driving the development of increasingly sophisticated computational and AI-based analytical methods.
2.3. Artificial Intelligence-Powered High-Content Analysis
Deep learning employs multi-layer neural network architectures composed of successive layers of trainable parameters to learn hierarchical representations from data. Such deep neural networks (DNNs) typically use differentiable nonlinear activation functions and are trained end-to-end using backpropagation and gradient-based optimization algorithms. Through successive nonlinear transformations, DNNs can automatically extract increasingly abstract feature representations with varying degrees of selectivity and invariance, enabling them to capture complex patterns in input data while reducing reliance on manually engineered features. In recent years, deep learning has transformed HCA, particularly in two key areas: cell segmentation and microscopy image feature extraction.
In the deep learning era, the development of cell segmentation algorithms (Figure 3) has largely followed two directions. The first focuses on achieving state-of-the-art performance on specific datasets. In 2015, Ronneberger et al. [21] introduced U-Net, a novel convolutional neural network architecture specifically designed for biomedical image segmentation. Evaluations on electron microscopy and transmitted-light microscopy datasets demonstrated that U-Net significantly outperformed previous sliding-window-based convolutional neural network approaches. Building upon U-Net, Schmidt et al. developed StarDist [49], which represents nuclei as star-convex polygons (see [49] for methodological details) and effectively addresses segmentation errors caused by overlapping or touching cells. Compared with conventional U-Net architectures, StarDist exhibited superior performance in densely packed cellular environments. Subsequently, Graham et al. proposed Hover-Net [50], a deep neural network that encodes instance information using horizontal and vertical distances from nuclear pixels to their corresponding centroids (see [50] for methodological details). This strategy enabled accurate segmentation of clustered nuclei with highly heterogeneous morphologies in hematoxylin and eosin (H&E)-stained histopathology images. The second direction aims to develop general-purpose segmentation methods capable of performing robustly across diverse datasets. In 2021, Stringer et al. developed Cellpose [25], a generalist segmentation framework extending beyond nuclei to more morphologically complex cellular structures. Cellpose was trained on a highly diverse collection of microscopy images and introduced a novel instance representation based on simulated diffusion dynamics (see [25] for methodological details). Specifically, the model was trained to predict the horizontal and vertical gradients of a flow field that encodes cellular topology. Experimental results demonstrated that Cellpose substantially outperformed StarDist while maintaining strong performance across diverse imaging modalities and cell types. Many excellent studies have built upon Cellpose and adapted it to HCS image datasets [51,52]. Otherwise, researchers developed general-purpose models for phase-contrast microscopy and tissue imaging [53,54]. Building upon these advances, the Stringer group continuously refined the Cellpose series [55,56], making substantial contributions to the cell segmentation community. These developments have greatly improved the robustness and accessibility of cell segmentation and have facilitated the widespread adoption of deep learning in HCA workflows.
Figure 3.
AI-powered cell segmentation methods.
In addition to segmentation, image feature representation has become another major focus of AI-driven HCA (Figure 4). Early studies primarily focused on dataset-specific representation learning. Pawlowski et al. [57] investigated whether convolutional neural networks pretrained on ImageNet could be directly applied to microscopy image feature extraction. By evaluating multiple architectures, they demonstrated that Inception V3 achieved the best balance between accuracy and computational efficiency on the BBBC021 benchmark dataset. Subsequently, Perakis et al. [58] utilized contrastive learning to generate single-cell representations without requiring manual annotations. Their approach improved mechanism-of-action prediction on the BBBC021 dataset and, for the first time, enabled an unsupervised representation learning method to match the performance of the best supervised approaches available at that time. Around the same period, Hua et al. [59] constructed a microscopy image classification task analogous to ImageNet to learn representations of microscopic images, termed CytoImageNet, which contains approximately 890,000 images spanning 894 categories. However, this approach did not achieve superior performance. Currently, state-of-the-art approaches have increasingly adopted self-supervised learning. In 2022, Kobayashi et al. introduced Cytoself [60], a self-supervised framework built upon the assumption that proteins with similar localization patterns should exhibit similar image representations. By pretraining on 24,382 protein images from the OpenCell database without manual annotation, Cytoself successfully generated high-resolution protein localization maps and demonstrated the power of self-supervised learning for biological image analysis. Another important challenge in HCA is the presence of batch effects. Technical variations arising from differences in experimental batches often degrade model performance. To address this issue, Lin et al. [27] proposed Batch Effects Normalization (BEN), a training strategy that aligns the concept of biological experimental batches with mini-batches used in deep learning optimization. Under this framework, each training mini-batch is sampled exclusively from a single experimental batch. BEN significantly improved representation learning performance and achieved state-of-the-art results on the RxRx1-Wilds benchmark dataset, highlighting the importance of explicitly accounting for experimental variation during model training. In 2024, our team proposed Microsnoop [61], a model trained on 10,458 high-quality microscopy images using a masked self-supervised training strategy (see [61] for methodological details). We evaluated it on ten high-quality datasets comprising more than 2.23 million images, covering diverse downstream representation tasks, including protein localization classification, mechanism-of-action prediction, and cell-cycle classification. Using four evaluation metrics (accuracy, Matthews correlation coefficient, balanced accuracy, and the F1-score), Microsnoop demonstrated robust performance and significantly outperformed CytoImageNet and Cytoself. These advances provide useful tools for extracting valuable information from HCS data.
Figure 4.
AI-powered microscopy image feature extraction methods.
3. Applications in Pharmaceutical Research
3.1. Deep Learning-Based Intelligent High-Content Screening
The rapid development of deep learning has created new opportunities for HCS. These advances are primarily reflected in two aspects (Figure 5). First, HCS images can be directly used to construct end-to-end deep learning models for specific downstream tasks. Second, deep learning-powered cell segmentation and image representation learning can be incorporated into HCS workflows to facilitate single-cell analysis and downstream phenotypic profiling.
Figure 5.
Applications of artificial intelligence in high-content screening.
One important application is the direct prediction of biological outcomes from HCS images. For example, Grafton et al. [62] established a training dataset using eight compounds with well-characterized mechanisms of action, including bortezomib (a proteasome inhibitor), doxorubicin (a topoisomerase inhibitor), cisapride (a 5-HT4 receptor agonist), and sorafenib (a tyrosine kinase inhibitor). Based on these data, they developed a deep learning model capable of automatically identifying compound-induced cardiotoxicity from cellular images. The proposed approach outperformed conventional methods based solely on fluorescence intensity measurements, demonstrating the ability of deep learning to capture complex phenotypic patterns associated with drug toxicity.
Deep learning has also enabled more sophisticated HCA through accurate cell segmentation and advanced image representation learning. Li et al. [63] developed π-PhenoDrug, a deep learning-based HCS pipeline for phenotype-driven drug discovery. The framework employs a customized UNet++-based nucleus segmentation model (NUSeg), incorporating an Xception encoder and scSE attention modules, to accurately identify individual cells from fluorescence microscopy images. Following segmentation, 106 single-cell morphological features, including intensity, morphology, and texture descriptors, were extracted to construct comprehensive phenotypic profiles. These profiles were subsequently analyzed using both random forest-based supervised classification and unsupervised clustering to evaluate drug-induced phenotypic perturbations and predict compound activity. Applied to melanoma cell lines, π-PhenoDrug successfully identified compounds with potential anti-melanoma effects from an 80-compound library and demonstrated higher sensitivity than conventional single-readout assays in detecting weak but biologically relevant phenotypes. Drug development for disorders such as Parkinson’s disease is often hindered by the lack of robust and screenable cellular phenotypes. To address this challenge, Schiff et al. [64] developed an automated drug discovery platform based on Cell Painting high-content imaging. Using an ImageNet-pretrained Inception network to generate image representations, they analyzed more than one million skin-cell images obtained from 91 individuals, including patients with Parkinson’s disease and healthy controls. Remarkably, this unbiased phenotypic screening strategy, which did not rely on any disease-specific biomarkers or prior biological assumptions, successfully distinguished patients from healthy controls and was even capable of differentiating disease subtypes.
These findings highlight the potential of deep learning-powered HCS to uncover subtle disease-associated phenotypes and facilitate drug discovery in areas where conventional screening approaches are limited. Collectively, these studies demonstrate that deep learning has expanded the scope of HCS from traditional image analysis to comprehensive phenotypic discovery. By enabling end-to-end prediction, accurate single-cell analysis, and biologically meaningful representation learning, deep learning has significantly enhanced the sensitivity and scalability of HCS.
3.2. Application in Bioactive Compounds Discovery from Natural Sources
Over the years, HCS has emerged as an important tool in natural product research (Table 1). Xu et al. [65] established an HCS-based phenotypic screening platform to identify TCM compounds capable of inhibiting transforming growth factor-β1-induced epithelial-mesenchymal transition (EMT). A549 cells were stained with TRITC-phalloidin and Hoechst 33342, and high-content imaging was used to quantify multiple morphological features, including cell area, roundness, and length. PCA was subsequently applied to integrate these phenotypic parameters and evaluate EMT-associated morphological changes. Using this strategy, 306 TCM-derived monomeric compounds were screened, leading to the identification of five anti-EMT candidates, namely camptothecin, dimethyl curcumin, artesunate, sinapine, and berberine. Further molecular and functional validation confirmed that these compounds increased E-cadherin expression, reduced vimentin and α-SMA levels, enhanced cell adhesion, and suppressed cell migration, demonstrating the utility of HCS-based phenotypic profiling for natural product discovery.
Table 1.
Representative applications of HCA in natural product research.
HCS can even be applied to investigate synergistic effects of different bioactive compounds from natural sources. Li et al. [66] combined HCS with high-resolution mass spectrometry to investigate the antithrombotic mechanisms of Guanxinning Tablet (GXNT), a TCM formula composed of Salvia miltiorrhiza and Ligusticum striatum. Major compounds were characterized using UPLC-HRMS and molecular networking, followed by functional screening in a PHZ-induced zebrafish thrombosis model. The study identified cryptotanshinone and senkyunolide I as two key bioactive compounds that synergistically restored blood circulation and suppressed thrombosis. Mechanistic analyses revealed that cryptotanshinone primarily alleviated oxidative stress and platelet activation, whereas senkyunolide I regulated the coagulation cascade, together producing enhanced antithrombotic efficacy. Chen et al. [67] integrated an angiogenesis-defective zebrafish model with HUVECs to identify pro-angiogenic components from GXNT. Through systematic screening of major herbal components, salvianolic acid B (Sal B) and ferulic acid (FA) were identified as a synergistic compound pair that significantly promoted endothelial cell migration, proliferation, and vascular sprouting. In addition to bioactive compound discovery, HCS has also been applied to safety assessment. For example, Wang et al. [68] established a multiparametric HCS platform for hepatotoxicity assessment of TCM injections. By simultaneously monitoring cellular proliferation, nuclear morphology, mitochondrial function, and membrane permeability in HepG2 cells, the assay rapidly identified potential hepatotoxic formulations and showed strong concordance with subsequent animal toxicity studies.
Recently, an increasing number of researchers have begun to incorporate deep learning into HCS-based natural product research. Chen et al. [33] employed a U-Net-based nuclear segmentation model to isolate individual cells from whole-field microscopy images. Based on prior biological knowledge that the DNA damage response protein 53BP1 forms distinct nuclear foci following ionizing radiation-induced DNA damage, single-cell images were categorized into three classes: damaged, normal, and nonsignaling. Approximately 2000 images were manually annotated for each category and subsequently expanded to 8000 images per class using data augmentation techniques. A VGG-19-based classification model was then trained to predict the damage status of previously unseen compounds. Applying this workflow to a library of 315 natural compounds led to the identification of isoliquiritigenin as a potential DNA damage-protective compound. Xing et al. [69] developed Deep-DPC, which integrates label-free time-series digital phase contrast (DPC) imaging with cellular morphology analysis for anti-fibrotic drug discovery. DPC images were first processed to extract 31 dynamic morphological features, and cells were categorized into six fibrosis-related phenotypic states based on TGF-β-induced morphological changes. Approximately 2000 images were manually annotated for each type and expanded to 48,000 images through data augmentation. An Inception V4-based classifier was subsequently trained to distinguish resting fibroblasts from activated myofibroblasts, achieving 92.3% accuracy on independent test images. Applied to a library of 1400 natural products, Deep-DPC identified 29 anti-fibrotic hit compounds and led to the discovery of Neo-Przewaquinone A, a previously unreported anti-fibrotic molecule that inhibits TGF-β receptor I signaling and alleviates cardiac fibrosis in vivo. Sun et al. [70] developed CPHNet, a deep learning-based Cell Painting screening pipeline for the discovery of anti-high-altitude pulmonary edema (HAPE) agents. The framework first generated more than 100,000 multichannel Cell Painting images from hypoxia-treated alveolar epithelial and pulmonary endothelial cells. A customized YOLOv8-based segmentation network (SegNet) was trained on manually annotated images to accurately identify and segment individual cells and subcellular structures. Subsequently, a six-channel ResNet-50-derived hypoxia scoring network (HypoNet) was trained using over 200,000 single-cell images to quantitatively assess cellular hypoxic states based on morphology. The resulting platform enabled automated phenotypic screening of candidate compounds and successfully identified ferulic acid and resveratrol as promising anti-HAPE agents, whose efficacy was further validated in both a 3D alveolus-on-a-chip model and an in vivo mouse model. Fang et al. [71] developed a high-content phenotypic screening platform based on a single-copy transgenic Caenorhabditis elegans strain expressing COL-12::GFP, enabling quantitative visualization of collagen biosynthesis and secretion in vivo. To facilitate automated image analysis, worm images acquired using a high-throughput imaging system were processed with the deep learning-based segmentation tool Scellseg, allowing accurate isolation of individual worms and fluorescence quantification. Using this platform, the authors screened 614 natural small molecules and identified 26 preliminary hits that enhanced collagen-associated fluorescence signals. Subsequent validation demonstrated that Danshensu, Lawsone, and Sanguinarine significantly increased COL-12 protein levels, primarily through post-transcriptional regulation of collagen biogenesis and secretion.
Overall, the integration of deep learning and high-content screening has opened new opportunities for natural product research, and many customized pipelines have been developed for bioactive compound discovery. Some studies [69,71] have begun to explore the use of general-purpose models for data analysis, thereby reducing the need for extensive data annotation and providing new possibilities for the rapid development of discovery pipelines applicable to emerging research scenarios. However, there is still insufficient evidence to demonstrate that performance improvements in segmentation and image representation can directly translate into substantial improvements in hit identification. More research is needed to determine whether advances in AI performance can translate into improved hit identification, biological validation, mechanism-of-action elucidation, and downstream drug discovery outcomes. Otherwise, existing studies have employed multiscale biological models, including whole-organism systems such as C. elegans, and combined HCS with cutting-edge technologies such as Cell Painting [72] to improve phenotypic analysis and compound discovery. Nevertheless, most efforts have been limited to the evaluation of single compounds. Given that the therapeutic effects of many natural products arise from multi-component synergy, extending deep learning-enabled HCS to characterize synergistic interactions and combinatorial effects remains a major challenge and a promising avenue for future investigation. Last but not least, because many natural products can induce strong stress or cytotoxic phenotypes, considering the safety of natural products during the screening process, such as using HCA to distinguish specific pharmacological responses from nonspecific cellular stress or cell death, is also an important research direction.
4. Conclusions and Perspectives
Artificial intelligence has transformed high-content screening, enabling more accurate image analysis, robust phenotypic profiling, and efficient bioactive compound discovery. By integrating advances in deep learning, cell segmentation, and feature extraction, HCS has evolved from a conventional image-analysis technique into a useful platform for phenotype-driven drug discovery. The complexity of chemical structures and biological mechanisms often limits the effectiveness of traditional target-based screening approaches. These approaches include not only single-target screening but also multi-target strategies, which commonly rely on computational methods to estimate the potential interactions or binding affinities between compounds and predefined targets [73,74]. Such approaches generally require prior knowledge of the targets of interest or a drug–target network database, making it challenging to identify unexpected or previously unknown targets. In contrast, AI-HCS provides a complementary phenotype-driven strategy that can rapidly capture complex cellular phenotypes (such as Cell Painting) resulting from the unknown modulation of multiple targets and signaling pathways, without requiring the targets to be specified in advance. This feature is particularly relevant to natural products, whose biological activities often arise from multi-target and multi-pathway interactions. Compared with existing multi-target approaches outside the HCS framework, HCS therefore offers a complementary strategy for investigating the phenotypic consequences of complex biological mechanisms and may facilitate the discovery of previously unrecognized mechanisms of action. Despite these advances, several challenges remain before AI-powered HCS can realize its full potential in natural product-based drug discovery.
One of the major limitations currently hindering the development of AI-driven HCS is the lack of large-scale, standardized datasets specifically designed for natural product research. In contrast to the computer vision field, where publicly available datasets such as ImageNet [75,76] have played a pivotal role in advancing deep learning, phenotypic datasets for natural products remain fragmented across individual laboratories and research projects. Most published studies focus on relatively small collections of compounds, making it difficult to train robust and generalizable AI models. Future efforts should focus on establishing comprehensive natural product phenotypic databases that integrate high-content imaging data, chemical structures, biological activities, molecular targets, and experimental metadata. This can help mitigate the risk of data leakage during model training, such as when data from the same batch, or even from the same plate or well, are included in both the training and evaluation sets. Such leakage may introduce model bias or overfitting and compromise the reliability of performance evaluation. Additionally, metadata such as drug concentration and treatment duration can also be incorporated to explore concentration-dependent and time-dependent phenotypes. Establishing standardized datasets would therefore help promote more reliable benchmark evaluation and facilitate meaningful cross-study comparisons. Furthermore, standardized data formats [77] and annotation protocols will be essential for improving data interoperability and facilitating the development and dissemination of reusable analytical pipelines.
Current AI applications in HCS are primarily concentrated in image analysis tasks, such as cell segmentation, feature extraction, and phenotype classification. However, the entire HCS workflow contains numerous stages that could benefit from artificial intelligence, such as image acquisition, quality control, hit identification, and lead prioritization. Recent advances in multimodal learning [78] and large language models [79] provide new opportunities to develop end-to-end intelligent screening systems. In the future, AI is expected to evolve from a data-analysis tool into a comprehensive decision-support system that assists researchers throughout the entire drug discovery process.
Another important future direction is the construction of intelligent and digitalized HCS platforms. Current screening workflows often involve multiple disconnected software packages and substantial manual intervention, creating bottlenecks in data processing, interpretation, and knowledge extraction. The integration of automated microscopy, laboratory robotics, cloud computing, artificial intelligence, and laboratory information management systems has the potential to establish fully digitalized screening environments. Such platforms [80,81] could enable automated experiment execution, real-time image analysis, dynamic decision-making, and closed-loop optimization. For natural product research, intelligent HCS platforms could further integrate chemical profiling technologies such as LC-MS, metabolomics, and molecular networking, creating a unified framework that links chemical composition, cellular phenotypes, and biological activities.
In conclusion, the convergence of artificial intelligence, high-content screening, and natural product science is opening new avenues for phenotype-driven drug discovery. The establishment of large-scale phenotypic databases, the integration of AI throughout the screening workflow, and the development of intelligent digitalized platforms are expected to accelerate the discovery of bioactive compounds from natural sources and promote the next generation of natural product-based pharmaceutical innovation.
Author Contributions
Conceptualization, Y.W. (Yi Wang), X.F. and D.X.; data curation, D.X.; writing—original draft preparation, D.X., Z.Z., H.W. and Y.W. (Yingchao Wang); writing—review and editing, D.X., Y.W. (Yi Wang) and X.F.; visualization, D.X.; supervision, Y.W. (Yi Wang) and X.F. All authors have read and agreed to the published version of the manuscript.
Funding
This research was funded by the National Natural Science Foundation of China (No. 82505197); the “Pioneer” and “Leading Goose” R&D Program of Zhejiang (2025C01110); the China Postdoctoral Science Foundation (BX20240318); and the Hangzhou Key Scientific Research Plan Projects (2025SZD1B25).
Institutional Review Board Statement
Not applicable.
Informed Consent Statement
Not applicable.
Data Availability Statement
No new data were created or analyzed in this study. Data sharing is not applicable to this article.
Conflicts of Interest
The authors declare no conflicts of interest.
References
- Hughes, R.E.; Elliott, R.J.R.; Dawson, J.C.; Carragher, N.O. High-content phenotypic and pathway profiling to advance drug discovery in diseases of unmet need. Cell Chem. Biol. 2021, 28, 338–355. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zanella, F.; Lorens, J.B.; Link, W. High content screening: Seeing is believing. Trends Biotechnol. 2010, 28, 237–245. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Singh, S.; Carpenter, A.E.; Genovesio, A. Increasing the content of high-content screening: An overview. SLAS Discov. 2014, 19, 640–650. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Boutros, M.; Heigwer, F.; Laufer, C. Microscopy-based high-content screening. Cell 2015, 163, 1314–1325. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lin, S.; Schorpp, K.; Rothenaigner, I.; Hadian, K. Image-based high-content screening in drug discovery. Drug Discov. Today 2020, 25, 1348–1361. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chandrasekaran, S.N.; Ceulemans, H.; Boyd, J.D.; Carpenter, A.E. Image-based profiling for drug discovery: Due for a machine-learning upgrade? Nat. Rev. Drug Discov. 2020, 20, 145–159. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Pratapa, A.; Doron, M.; Caicedo, J.C. Image-based cell phenotyping with deep learning. Curr. Opin. Chem. Biol. 2021, 65, 9–17. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Krentzel, D.; Shorte, S.L.; Zimmer, C. Deep learning in image-based phenotypic drug discovery. Trends Cell Biol. 2023, 33, 538–554. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Mattiazzi Usaj, M.; Styles, E.B.; Verster, A.J.; Friesen, H.; Boone, C.; Andrews, B.J. High-content screening for quantitative cell biology. Trends Cell Biol. 2016, 26, 598–611. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Heynen-Genel, S.; Pache, L.; Chanda, S.K.; Rosen, J. Functional genomic and high-content screening for target discovery and deconvolution. Expert Opin. Drug Discov. 2012, 7, 955–968. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Li, S.; Xia, M. Review of high-content screening applications in toxicology. Arch. Toxicol. 2019, 93, 3387–3396. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Bellomo, F.; Medina, L.; Leo, E.; Panarella, A.; Emma, F. High-content drug screening for rare diseases. J. Inherit. Metab. Dis. 2017, 40, 601–607. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Arta, R.K.; Watanabe, Y.; Egawa, J.; Lemmon, V.P.; Someya, T. Linking autism risk genes to morphological and pharmaceutical screening by high-content imaging: Future directions and opinion. Psychiatry Clin. Neurosci. 2025, 79, 435–446. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hu, Y.; Xue, X.; Han, T.; Li, Y.; Zhang, T.; Lu, T.; Zhang, P. An effective system for senescence modulating drug development using quantitative high-content analysis and high-throughput screening. Commun. Biol. 2025, 8, 1316. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhang, Z.; Gao, L.; Zheng, H.; Zhong, Y.; Li, G.; Ye, Z.; Sun, Q.; Wang, B.; Weng, Z. High-content imaging and deep learning-driven detection of infectious bacteria in wounds. Bioprocess Biosyst. Eng. 2025, 48, 301–315. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Koehn, F.E.; Carter, G.T. The evolving role of natural products in drug discovery. Nat. Rev. Drug Discov. 2005, 4, 206–220. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Harvey, A.L. Natural products in drug discovery. Drug Discov. Today 2008, 13, 894–901. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wright, G.D. Unlocking the potential of natural products in drug discovery. Microb. Biotechnol. 2019, 12, 55–57. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Atanasov, A.G.; Zotchev, S.B.; Dirsch, V.M.; Orhan, I.E.; Banach, M.; Rollinger, J.M.; Barreca, D.; Weckwerth, W.; Bauer, R.; Bayer, E.A.; et al. Natural products in drug discovery: Advances and opportunities. Nat. Rev. Drug Discov. 2021, 20, 200–216. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Newman, D.J. Natural products and drug discovery. Natl. Sci. Rev. 2022, 9, nwac206. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional networks for biomedical image segmentation. In Proceedings of the Medical Image Computing and Computer Assisted Intervention, Munich, Germany, 5–9 October 2015. [Google Scholar] [CrossRef] [Scilit]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. In Proceedings of the Neural Information Processing Systems, Long Beach, CA, USA, 4–9 December 2017. [Google Scholar] [CrossRef] [Scilit]
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin Transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021. [Google Scholar] [CrossRef] [Scilit]
- He, K.; Chen, X. Masked autoencoders are scalable vision learners. In Proceedings of the Computer Vision and Pattern Recognition, Nashville, Tennessee, 19–25 June 2021. [Google Scholar] [CrossRef] [Scilit]
- Stringer, C.; Wang, T.; Michaelos, M.; Pachitariu, M. Cellpose: A generalist algorithm for cellular segmentation. Nat. Methods 2021, 18, 100–106. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Leek, J.T.; Scharpf, R.B.; Bravo, H.C.; Simcha, D.; Langmead, B.; Johnson, W.E.; Geman, D.; Baggerly, K.; Irizarry, R.A. Tackling the widespread and critical impact of batch effects in high-throughput data. Nat. Rev. Genet. 2010, 11, 733–739. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lin, A.; Lu, A.X. Incorporating knowledge of plates in batch normalization improves generalization of deep learning for microscopy images. In Proceedings of Machine Learning Research; JMLR: Cambridge, MA, USA, 2022. [Google Scholar] [CrossRef] [Scilit]
- Chen, X.; Li, L.; Zhang, M.; Yang, J.; Lyu, C.; Xu, Y.; Yang, Y.; Wang, Y. Guidelines for application of high-content screening in traditional Chinese medicine: Concept, equipment, and troubleshooting. Acupunct. Herb. Med. 2024, 4, 1–15. [Google Scholar] [CrossRef] [Scilit]
- Smith, K.; Li, Y.; Piccinini, F.; Csucs, G.; Balazs, C.; Bevilacqua, A.; Horvath, P. CIDRE: An illumination-correction method for optical microscopy. Nat. Methods 2015, 12, 404–406. [Google Scholar] [CrossRef] [Scilit]
- Yi, F.; Moon, I. Image segmentation: A survey of graph-cut methods. In Proceedings of the International Conference on Systems and Informatics, Yantai, China, 19–21 May 2012. [Google Scholar] [CrossRef] [Scilit]
- Sommer, C.; Straehle, C.; Kothe, U.; Hamprecht, F.A. Ilastik: Interactive learning and segmentation toolkit. In Proceedings of the IEEE International Symposium on Biomedical Imaging, Chicago, IL, USA, 30 March–2 April 2011. [Google Scholar] [CrossRef] [Scilit]
- Carpenter, A.E.; Jones, T.R.; Lamprecht, M.R.; Clarke, C.; Kang, I.; Friman, O.; Guertin, D.A.; Chang, J.; Lindquist, R.A.; Moffat, J.; et al. CellProfiler: Image analysis software for identifying and quantifying cell phenotypes. Genome. Biol. 2006, 7, R100. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chen, X.; Xun, D.; Zheng, R.; Zhao, L.; Lu, Y.; Huang, J.; Wang, R.; Wang, Y. Deep-learning-assisted assessment of DNA damage based on Foci images and its application in high-content screening of lead compounds. Anal. Chem. 2020, 92, 14267–14277. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Sanchez, M.; Bourriez, N.; Bendidi, I.; Cohen, E.; Svatko, I.; Del Nery, E.; Tajmouati, H.; Bollot, G.; Calzone, L.; Genovesio, A. Large scale compound selection guided by cell painting reveals activity cliffs and functional relationships. Commun. Biol. 2026, 9, 225. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Rohban, M.H.; Singh, S.; Wu, X.; Berthet, J.B.; Bray, M.-A.; Shrestha, Y.; Varelas, X.; Boehm, J.S.; Carpenter, A.E. Systematic morphological profiling of human gene and allele function via Cell Painting. eLife 2017, 6, e24060. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Maaten, L.; Geoffrey, H. Visualizing data using t-SNE. J. Mach. Learn. Res. 2008, 9, 2579–2605. [Google Scholar] [CrossRef] [Scilit]
- Becht, E.; McInnes, L.; Healy, J.; Dutertre, C.-A.; Kwok, I.W.H.; Ng, L.G.; Ginhoux, F.; Newell, E.W. Dimensionality reduction for visualizing single-cell data using UMAP. Nat. Biotechnol. 2019, 37, 38–44. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Cree, I.A.; Glaysher, S.; Harvey, A.L. Efficacy of anti-cancer agents in cell lines versus human primary tumour tissue. Curr. Opin. Pharmacol. 2010, 10, 375–379. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Pijnappels, D.A.; Schalij, M.J.; Ramkisoensing, A.A.; Van Tuyn, J.; De Vries, A.A.F.; Van Der Laarse, A.; Ypey, D.L.; Atsma, D.E. Forced alignment of mesenchymal stem cells undergoing cardiomyogenic differentiation affects functional integration with cardiomyocyte cultures. Circ. Res. 2008, 103, 167–176. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kuhn, M.R.; Wolcott, E.A.; Langer, E.M. Developments in gastrointestinal organoid cultures to recapitulate tissue environments. Front. Bioeng. Biotechnol. 2025, 13, 1521044. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lai, W.; Geliang, H.; Bin, X.; Wang, W. Effects of hydrogel stiffness and viscoelasticity on organoid culture: A comprehensive review. Mol. Med. 2025, 31, 83. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Price, F.D.; Matyas, M.N.; Gehrke, A.R.; Chen, W.; Wolin, E.A.; Holton, K.M.; Gibbs, R.M.; Lee, A.; Singu, P.S.; Sakakeeny, J.S.; et al. Organoid culture promotes dedifferentiation of mouse myoblasts into stem cells capable of complete muscle regeneration. Nat. Biotechnol. 2025, 43, 889–903. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Youhanna, S.; Kemas, A.M.; Wright, S.C.; Zhong, Y.; Klumpp, B.; Klein, K.; Motso, A.; Michel, M.; Ziegler, N.; Shang, M.; et al. Chemogenomic screening in a patient-derived 3D fatty liver disease model reveals the CHRM1-TRPM8 axis as a novel module for targeted intervention. Adv. Sci. 2025, 12, 2407572. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Torres, A.K.; Mira, R.G.; Pinto, C.; Inestrosa, N.C. Studying the mechanisms of neurodegeneration: C. elegans advantages and opportunities. Front. Cell. Neurosci. 2025, 19, 1559151. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Rani, N.; Alam, M.M.; Parvez, S. Toxicity of glyphosate accelerates neurodegeneration in Caenorhabditis elegans model of Alzheimer’s disease. Front. Toxicol. 2025, 7, 1578230. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hernández-Silva, D.; López-Abellán, M.D.; Martínez-Navarro, F.J.; García-Castillo, J.; Cayuela, M.L.; Alcaraz-Pérez, F. Development of a short telomere zebrafish model for accelerated aging research and antiaging drug screening. Aging Cell 2025, 24, e70007. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Siddiqui, S.; Siddiqui, H.; Riguene, E.; Nomikos, M. Zebrafish: A versatile and powerful model for biomedical research. BioEssays 2025, 47, e70080. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Traxler, P.; Reichl, S.; Folkman, L.; Shaw, L.; Fife, V.; Nemc, A.; Pasajlic, D.; Kusienicka, A.; Barreca, D.; Fortelny, N.; et al. Integrated time-series analysis and high-content CRISPR screening delineate the dynamics of macrophage immune regulation. Cell Syst. 2025, 16, 101346. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Schmidt, U.; Weigert, M.; Broaddus, C.; Myers, G. Cell detection with star-convex polygons. In Proceedings of the Medical Image Computing and Computer Assisted Intervention, Granada, Spain, 16–20 September 2018. [Google Scholar] [CrossRef] [Scilit]
- Graham, S.; Vu, Q.D.; Raza, S.E.A.; Azam, A.; Tsang, Y.W.; Kwak, J.T.; Rajpoot, N. Hover-Net: Simultaneous segmentation and classification of nuclei in multi-tissue histology images. Med. Image Anal. 2019, 58, 101563. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Xun, D.; Chen, D.; Zhou, Y.; Lauschke, V.M.; Wang, R.; Wang, Y. Scellseg: A style-aware deep learning tool for adaptive cell instance segmentation by contrastive fine-tuning. iScience 2022, 25, 105506. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lam, V.K.; Byers, J.M.; Robitaille, M.C.; Kaler, L.; Christodoulides, J.A.; Raphael, M.P. A self-supervised learning approach for high throughput and high content cell segmentation. Commun. Biol. 2025, 8, 780. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Edlund, C.; Jackson, T.R.; Khalid, N.; Bevan, N.; Dale, T.; Dengel, A.; Ahmed, S.; Trygg, J.; Sjögren, R. LIVECell-a large-scale dataset for label-free live cell segmentation. Nat. Methods 2021, 18, 1038–1045. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Greenwald, N.F.; Miller, G.; Moen, E.; Kong, A.; Kagel, A.; Dougherty, T.; Fullaway, C.C.; McIntosh, B.J.; Leow, K.X.; Schwartz, M.S.; et al. Whole-cell segmentation of tissue images with human-level performance using large-scale data annotation and deep learning. Nat. Biotechnol. 2022, 40, 555–565. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Pachitariu, M.; Stringer, C. Cellpose 2.0: How to train your own model. Nat. Methods 2022, 19, 1634–1641. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Stringer, C.; Pachitariu, M. Cellpose3: One-click image restoration for improved cellular segmentation. Nat. Methods 2025, 22, 592–599. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Pawlowski, N.; Caicedo, J.C.; Singh, S.; Carpenter, A.E.; Storkey, A. Automating morphological profiling with generic deep convolutional networks. bioRxiv 2016, 085118. [Google Scholar] [CrossRef] [Scilit]
- Perakis, A.; Gorji, A.; Jain, S.; Chaitanya, K.; Rizza, S.; Konukoglu, E. Contrastive learning of single-cell phenotypic representations for treatment classification. In Proceedings of the Machine Learning in Medical Imaging, Strasbourg, France, 27 September 2021. [Google Scholar] [CrossRef] [Scilit]
- Hua, S.B.Z.; Lu, A.X.; Moses, A.M. CytoImageNet: A large-scale pretraining dataset for bioimage transfer learning. In Proceedings of the Neural Information Processing Systems, Online, 6–14 December 2021. [Google Scholar] [CrossRef] [Scilit]
- Kobayashi, H.; Cheveralls, K.C.; Leonetti, M.D.; Royer, L.A. Self-supervised deep learning encodes high-resolution features of protein subcellular localization. Nat. Methods 2022, 19, 995–1003. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Xun, D.; Wang, R.; Zhang, X.; Wang, Y. Microsnoop: A generalist tool for microscopy image representation. Innovation 2024, 5, 100541. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Grafton, F.; Ho, J.; Ranjbarvaziri, S.; Farshidfar, F.; Budan, A.; Steltzer, S.; Maddah, M.; Loewke, K.E.; Green, K.; Patel, S.; et al. Deep learning detects cardiotoxicity in a high-content screen with induced pluripotent stem cell-derived cardiomyocytes. eLife 2021, 10, e68714. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Li, X.; Ouyang, Q.; Han, M.; Liu, X.; He, F.; Zhu, Y.; Leng, L.; Ma, J. Π-PhenoDrug: A comprehensive deep learning-based pipeline for phenotypic drug screening in high-content analysis. Adv. Intell. Syst. 2025, 7, 2400635. [Google Scholar] [CrossRef] [Scilit]
- Schiff, L.; Migliori, B.; Chen, Y.; Carter, D.; Bonilla, C.; Hall, J.; Fan, M.; Tam, E.; Ahadi, S.; Fischbacher, B.; et al. Integrating deep learning and unbiased automated high-content screening to identify complex disease signatures in human fibroblasts. Nat. Commun. 2022, 13, 1590. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Xu, M.; Cui, Q.; Su, W.; Zhang, D.; Pan, J.; Liu, X.; Pang, Z.; Zhu, Q. High-content screening of active components of traditional Chinese medicine inhibiting TGF-β-induced cell EMT. Heliyon 2022, 8, e10238. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Li, J.; Liu, H.; Yang, Z.; Yu, Q.; Zhao, L.; Wang, Y. Synergistic effects of cryptotanshinone and senkyunolide I in guanxinning tablet against endogenous thrombus formation in zebrafish. Front. Pharmacol. 2021, 11, 622787. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chen, J.; Wang, Y.; Wang, S.; Zhao, X.; Zhao, L.; Wang, Y. Salvianolic acid B and ferulic acid synergistically promote angiogenesis in HUVECs and zebrafish via regulating VEGF signaling. J. Ethnopharmacol. 2022, 283, 114667. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, M.; Liu, C.-X.; Dong, R.-R.; He, S.; Liu, T.-T.; Zhao, T.-C.; Wang, Z.-L.; Shen, X.-Y.; Zhang, B.-L.; Gao, X.-M.; et al. Safety evaluation of Chinese medicine injections with a cell imaging-based multiparametric assay revealed a critical involvement of mitochondrial function in hepatotoxicity. J. Evid.-Based Complement. Altern. Med. 2015, 2015, 379586. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Xing, X.; Yan, X.; Tan, Y.; Liu, Y.; Cui, Y.; Feng, C.; Cai, Y.; Dai, H.; Gao, W.; Zhou, P.; et al. Deep-DPC: Deep learning-assisted label-free temporal imaging discovery of anti-fibrotic compounds by controlling cell morphology. J. Adv. Res. 2025, 78, 703–716. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Sun, D.; Yang, X.; Huang, C.; Bai, Z.; Shen, P.; Ni, Z.; Huang-fu, C.; Hu, Y.; Wang, N.; Tang, X.; et al. CPHNet: A novel pipeline for anti-HAPE drug screening via deep learning-based cell painting scoring. Respir. Res. 2025, 26, 91. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Fang, J.; Wu, X.; Meng, X.; Xun, D.; Xu, S.; Wang, Y. Discovery of natural small molecules promoting collagen secretion by high-throughput screening in Caenorhabditis elegans. Molecules 2022, 27, 8361. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Seal, S.; Trapotsi, M.-A.; Spjuth, O.; Singh, S.; Carreras-Puigvert, J.; Greene, N.; Bender, A.; Carpenter, A.E. Cell Painting: A decade of discovery and innovation in cellular imaging. Nat. Methods 2024, 22, 254–268. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chaudhari, R.; Fong, L.W.; Tan, Z.; Huang, B.; Zhang, S. An up-to-date overview of computational polypharmacology in modern drug discovery. Expert Opin. Drug Discov. 2020, 15, 1025–1044. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Manen-Freixa, L.; Antolin, A.A. Polypharmacology prediction: The long road toward comprehensively anticipating small-molecule selectivity to de-risk drug discovery. Expert Opin. Drug Discov. 2024, 19, 1043–1069. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Liu, Z.; He, K. A decade’s battle on dataset bias: Are we there yet? In Proceedings of the International Conference on Learning Representations, Singapore, 24–28 April 2025. [Google Scholar] [CrossRef] [Scilit]
- Yüksel, B.B.; Yılmazer Metin, A. Artificial intelligence breakthroughs and data futures: A retrospective and prospective review. Acad. Platf. J. Eng. Smart Syst. 2026, 14, 1–16. [Google Scholar] [CrossRef] [Scilit]
- Yeh, F.-C. DSI studio: An integrated tractography platform and fiber data hub for accelerating brain research. Nat. Methods 2025, 22, 1617–1619. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Meijer, D.; Beniddir, M.A.; Coley, C.W.; Mejri, Y.M.; Öztürk, M.; Van Der Hooft, J.J.J.; Medema, M.H.; Skiredj, A. Empowering natural product science with AI: Leveraging multimodal data and knowledge graphs. Nat. Prod. Rep. 2025, 42, 654–662. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Feuerriegel, S.; Maarouf, A.; Bär, D.; Geissler, D.; Schweisthal, J.; Pröllochs, N.; Robertson, C.E.; Rathje, S.; Hartmann, J.; Mohammad, S.M.; et al. Using natural language processing to analyse text data in behavioural science. Nat. Rev. Psychol. 2025, 4, 96–111. [Google Scholar] [CrossRef] [Scilit]
- Allmendinger, S.; Bonenberger, L.; Endres, K.; Fetzer, D.; Gimpel, H.; Kühl, N. Multi-agent AI. Electron. Mark. 2026, 36, 18. [Google Scholar] [CrossRef] [Scilit]
- Ghareeb, A.E.; Chang, B.; Mitchener, L.; Yiu, A.; Szostkiewicz, C.J.; Shved, D.; Gyimesi, G.J.; Laurent, J.M.; Wright, S.M.; Razzak, M.T.; et al. A multi-agent system for automating scientific discovery. Nature 2026, 655, 497–505. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.




