2.1. Trends in Publications and Research Focus
The task of cell segmentation in microscopy has become a key area of research in biomedical image analysis given its essential role in quantifying cellular phenotypes, tracking dynamics, and supporting diagnostic workflows. Over the past two decades, this field has gone through a significant evolution from the use of classical image processing methods to the adoption of modern deep learning techniques. This reflects, not only technological advancements in microscopy and computation, but also an increasing demand for scalable and accurate analysis in high-throughput biology and clinical pathology.
In the early 2000s, segmentation tasks relied heavily on traditional computer vision algorithms. These included global and adaptive thresholding techniques such as Otsu’s method [
13], edge detection [
14], watershed-based segmentation [
15], and morphological operations (useful in tasks like noise removal or shape extraction). Such methods were often combined with manual feature extraction and simple machine learning classifiers, like support vector machines [
16] and random forests [
17]. While these approaches worked reasonably well on clean and high-contrast images, they struggled in cases with dense cellular clustering, diverse cell shapes, and noisy backgrounds. Such conditions are common in modalities like brightfield and phase-contrast microscopy [
4]. During this phase, tools like CellProfiler [
18] were instrumental as they provided a flexible, open-source platform that enabled experts to design image analysis workflows combining segmentation, feature extraction, and statistical quantification. It helped establish reproducible pipelines for large-scale biological experiments, such as RNAi screening [
19] and early high-content phenotypic profiling works [
20].
With the advances in microscopy and the growth of high-throughput imaging with techniques such as fluorescence labeling, live-cell imaging, and volumetric scanning, the complexity and quantity of image data increased substantially. This highlighted the limitations of previous segmentation approaches and provoked a shift towards data-driven methods. The introduction of deep learning into biomedical image analysis in the mid-2010s was a turning point. A major advancement was the U-Net architecture [
7], which was designed specifically for biomedical segmentation. Its encoder–decoder structure with skip connections enabled precise localization while leveraging multi-scale features, and it showed strong performance with relatively few labeled samples. U-Net became the basis for many other variants and extensions across microscopy, histology, and other medical imaging domains [
21,
22].
Throughout the late 2010s and early 2020s, a variety of deep learning models were proposed to address the main challenges in microscopy. These include dealing with dense cellular clustering [
23,
24], segmenting overlapping or irregularly shaped cells [
10], and adapting to multiple imaging modalities [
25]. Deep learning was also extended to 3D segmentation [
26] and cell tracking in time-lapse data [
6]. A growing number of public datasets and community benchmarks supported these developments. Challenges like the Data Science Bowl 2018 [
27], the ISBI Cell Tracking Challenge [
28] or the NeurIPS 2022 Cell Segmentation Challenge [
29] offered high-quality annotations and standardized evaluation metrics. These resources accelerated the research progress by enabling reproducible comparisons and catalyzing community innovation.
More recently, between 2023 and 2025, there has been a rise in methods using foundation models and multimodal learning to further improve segmentation accuracy and generalization across diverse datasets. The combination of vision transformers with convolutional architectures have demonstrated superior performance on complex tasks involving heterogeneous cell populations and varying imaging conditions [
30]. In addition, self-supervised and few-shot learning approaches have gained attention, addressing the problem of limited annotated training data by enabling models to learn robust representations directly from unlabeled images [
31,
32]. These innovations have established a trend towards more flexible, scalable, and generalizable segmentation methodologies, pushing the boundaries of biological image analysis.
The growing attention to cell segmentation in microscopy is evident from the steady rise in scientific publications over the past two decades. A PubMed query using the terms “microscopy” AND “cell segmentation” returned 305 works in the year 2000 compared to peaks above 600 in 2021, with publication counts remaining above 500 in subsequent years (see
Figure 1). This surge is particularly marked after 2015, coinciding with the widespread adoption of convolutional neural networks, transfer learning, and self-supervised learning in biomedical imaging [
33,
34]. The trend highlights a clear methodological shift from classical image processing approaches to deep learning-based strategies. More recently, research has begun to explore foundation models and multimodal frameworks, reflecting the field’s continuous push toward robust, generalizable, and biologically meaningful cell segmentation.
2.2. Applications of Cell Segmentation in Biomedical Research
Cell segmentation is a foundational component of biomedical image analysis. It enables multiple downstream applications in both research and clinical domains. Accurate delineation of individual cells and nuclei provides essential information about morphology, spatial organization, and temporal dynamics, which can be exploited in fields like cancer diagnostics, drug screening, and developmental biology.
In cancer, precise nuclear and cytoplasmic segmentation is utilized to extract features like size, shape, and texture that correlate with malignancy, mitotic activity, and tumor grade [
35,
36]. Concretely, histopathology workflows use instance segmentation for tumor boundary delineation, detection of tumor-infiltrating lymphocytes, and gland segmentation in prostate and breast cancer tissues. By automating these tasks, segmentation reduces inter-observer variability and significantly speeds up diagnostic processes.
In drug discovery and high-content screening, segmentation allows the quantification of phenotypic responses at a single-cell level. Tools such as CellProfiler [
18] have made it possible to perform large-scale profiling of cellular morphology in response to drug perturbations, facilitating mode-of-action prediction and toxicity assessment [
37].
Segmentation is also crucial in developmental and stem cell biology, where tracking of cells over time in 3D images allows lineage tracing and understanding of morphogenetic processes [
28]. Similarly, in neuroscience, nuclear and soma segmentation is used to analyze cell distributions, layer structures, and pathological changes in brain tissue.
Recent progress in spatial omics technologies, such as spatial transcriptomics and multiplexed imaging, depend on cell segmentation to map molecular data to individual cells, enabling spatially resolved single-cell analysis [
38]. This integration of imaging and genomics needs highly accurate segmentation, particularly in densely packed tissues with diverse cell types.
Overall, cell segmentation functions as the bridge between raw microscopy images and quantitative biological insight. Its applications are increasingly diverse and key for modern biomedical research.
2.3. Microscopy Techniques
Microscopy-based biomedical research encompasses a wide range of cell types, including cultured mammalian cells, stem cells, microbial organisms, and tissue biopsies. Each of them contains different morphological features and presents unique imaging challenges. For example, breast cancer cell lines such as MCF-7 and T47D are widely used in cancer biology to investigate tumor progression and drug responses [
39,
40]. Meanwhile, cells like Staphylococcus aureus serve as a key model in microbiology and infectious disease research [
41]. Additionally, tissue sections introduce further complexity due to their dense and diverse cellular composition [
1].
Addressing this wide variety of cellular features required the development of multiple microscopy approaches tailored to different needs. Early advances in microscopy began with transmitted light techniques such as phase-contrast microscopy [
42], polarized light microscopy [
43], and differential interference contrast (DIC) microscopy [
44]. These label-free modalities enhanced the visibility of live cells by increasing intrinsic contrast without the need for staining. Although they represented a significant improvement for cell biology research, these techniques posed some limitations for automated image analysis.
Over time, microscopy techniques have developed into two primary categories: labeled and label-free imaging. Label-free methods, such as brightfield and phase-contrast microscopy, have been key for live-cell observation due to their simplicity and non-invasive nature. However, these modalities often suffer from low contrast and less clear cellular features, which present significant challenges for computational analysis and segmentation [
45]. In contrast, the rise of labeled imaging, most notably fluorescence microscopy, revolutionized cell biology by enabling the visualization of specific biomolecular structures. This distinction between label-free and labeled approaches continues to shape modern imaging strategies, particularly in how image data is interpreted and processed by automated pipelines.
Fluorescence microscopy remains one of the most powerful and used tools in modern cell biology, offering specific and dynamic visualization of cellular structures and processes. This technique uses fluorescent probes like dyes, genetically encoded fluorescent proteins, and targeted antibodies. These bind selectively to biomolecules such as proteins, lipids, or ions [
46]. This molecular specificity has enabled major advances in understanding cellular organization, protein localization, and real-time signaling events.
To overcome the diffraction limit of conventional fluorescence microscopy, several super-resolution techniques have been developed. Methods such as Photoactivated Localization Microscopy (PALM) [
3], Stimulated Emission Depletion (STED) [
2], and Stochastic Optical Reconstruction Microscopy (STORM) [
47] allow imaging at the nanometer scale, making it possible to study subcellular architectures with higher detail.
Building on these advances, fluorescence microscopy has further evolved to address limitations in imaging depth, speed, and live-cell compatibility. Light-sheet fluorescence microscopy (LSFM) [
5] and, in particular, lattice light-sheet microscopy (LLSM) [
48] allow fast and volumetric imaging of live cells and tissues with minimal photodamage. These techniques have enabled the capture of dynamic biological processes in three dimensions over time, with recent applications such as embryogenesis [
49], neural activity mapping [
50], and immune cell dynamics [
51].
More recently, fluorescence microscopy has advanced through high-content and multiplexed imaging strategies. Techniques such as spectral imaging, molecular barcoding, and sequential fluorescence in situ hybridization (seqFISH) [
52,
53] allow for the simultaneous detection of many molecular species within the same sample. In addition, the integration of deep learning and artificial intelligence is rapidly transforming fluorescence microscopy data analysis. AI-driven methods enhance image denoising, resolution, and segmentation [
54,
55]. This facilitates the reconstruction of high-quality images from low-exposure data, reduction of phototoxicity, and even prediction of fluorescence labels from transmitted-light images. Together, these advances demonstrate how fluorescence microscopy keeps pushing the boundaries of cellular imaging.
2.4. Computer Vision Tasks
Artificial intelligence has revolutionized the analysis of microscopy data by enabling a broad spectrum of tasks, from low-level image interpretation to complex biological insight. The main Computer Vision (CV) tasks in this domain include object detection, classification, semantic and instance segmentation, anomaly detection, cell tracking, and 3D segmentation and reconstruction. Each of these plays an important role in biomedical image analysis, enabling processes like quality control, phenotyping, and the modeling of dynamic biological processes.
Object Detection consists of identifying and localizing individual cells or structures using bounding boxes. This tasks often serves as a first step for more complex tasks such as segmentation or tracking. While traditional detection relied on handcrafted features and region proposal methods [
56], deep learning-based detectors such as the Single Shot Multibox Detector (SSD) [
57] and You Only Look Once (YOLO) [
58] revolutionized the field by achieving real-time and end-to-end detection in a single network pass. Even though these models were originally developed for natural scenes, they have been adapted for microscopy and histopathology images to detect nuclei, mitotic events, and tissue abnormalities [
59]. Object detection plays a key role in applications like mitosis detection in cancer diagnostics and identifying regions of interest for downstream analysis [
33].
Image Classification is one of the most fundamental tasks, where models are trained to assign discrete labels to full images or certain regions of interest. Common applications in biomedical contexts involve classifying between cancerous and non-cancerous tissue samples, identifying different cell types, detecting stages of infection, or predicting cellular responses to treatments. In the past, most of the methodologies contained image descriptors and classical machine learning methods, such as random forests or support vector machines [
60]. However, they have been replaced by convolutional neural networks (CNNs) given their higher ability to learn feature representations. In high-content screening workflows, classification models are often used for automated phenotypic profiling, supporting large-scale drug discovery and toxicity studies [
61].
More recent trends in microscopy classification include the adoption of transformer-based architectures [
62], self-supervised learning [
63], and multimodal fusion (e.g., combining image data with metadata or gene expression) [
64]. These approaches aim to enhance generalization across datasets and experimental conditions, which is the main challenge in the field due to batch effects and biological variability.
Additionally, explainability is a growing research focus. Saliency maps, class activation maps (CAMs), and other visualization techniques are used to highlight those regions that contribute most to the model’s decision, supporting interpretability in clinical or biological contexts [
65,
66].
Semantic Segmentation provides pixel-level classification of microscopy images, assigning each pixel to a specific class, such as nucleus, cytoplasm, background, or tissue type. This task is particularly important for morphometric analyses, allowing researchers to quantify features like cell size, shape, and spatial organization. A major advancement in this area was the development of U-Net [
7], which introduced a symmetric encoder–decoder architecture with skip connections, enabling accurate localization and robust generalization from relatively small datasets. Since then, U-Net has become the baseline in the field, inspiring numerous adaptations and extensions incorporating deeper backbones, residual connections, attention modules, and adversarial refinement strategies [
67,
68]. Complementary architectures such as DeepLab [
69] have also demonstrated strong performance, particularly with their use of dilated convolutions and fully connected conditional random fields (CRFs) for improving object boundary precision. These models have been adapted for biomedical images, where capturing fine details such as cell borders is essential.
More recently, transformer-based architectures and transfer learning have shown promising results for generalization across datasets. Models like UNETR [
70] and SegFormer [
71] leverage self-attention mechanisms, improving segmentation accuracy in complex biomedical samples. These architectures, combined with transfer learning strategies, have shown strong performance even with limited annotated data. For instance, comparative studies of deep transfer learning models have demonstrated their potential for generalization and domain adaptation [
72]. Such approaches reduce the dependence on large annotated datasets while enhancing cross-domain robustness.
Semantic segmentation has now been applied to a wide range of imaging modalities, from fluorescence microscopy of cultured cells to brightfield and histological tissue sections. Ongoing challenges, like staining variability, imaging artifacts, and domain shifts across labs and instruments, have grown interest in unsupervised domain adaptation and self-supervised pretraining techniques.
Instance segmentation goes beyond semantic segmentation by not only classifying each pixel but also distinguishing individual objects within the same class (see
Figure 2). This detail is critical for analyzing densely packed or overlapping cells. In single-cell biology, instance segmentation enables accurate quantification of cell counts, spatial organization, and cellular heterogeneity.
State-of-the-art approaches in cell instance segmentation have been built upon general frameworks such as Mask R-CNN [
8], adapted to biological imaging contexts. Domain-specific tools like Cellpose [
23] and StarDist [
10] incorporated tailored strategies to accurately delineate cell boundaries even under challenging imaging conditions. The segmented instances produced by these models often provide the starting point for subsequent analyses, such as cell tracking, lineage reconstruction, and phenotypic profiling. A detailed discussion of instance segmentation methods and their applications is provided in the next section.
Anomaly Detection aims to identify rare, unexpected, or abnormal patterns within microscopy images. Such anomalies may correspond to unusual phenotypes, mitotic defects, apoptotic bodies, or imaging artifacts. Due to their rarity in most datasets, anomaly detection often relies on unsupervised or self-supervised learning techniques. Methods such as autoencoders [
73], generative adversarial networks (GANs) [
74], and contrastive learning [
75] are frequently employed to model normal data distributions, enabling us to flag deviations as anomalies. In pathology, anomaly detection has been applied to identify tumor regions in large histological images [
76] and detect poorly differentiated cells in hematological samples [
77].
Cell Tracking over time is essential for studying dynamic biological processes such as migration, proliferation, differentiation, and apoptosis. This task involves identifying and associating cells across consecutive time-lapse frames to reconstruct their temporal trajectories. Traditional approaches have relied on object detection and motion prediction algorithms, such as Kalman filters or nearest-neighbor heuristics [
78,
79]. However, deep learning-based tracking methods, including recurrent neural networks (RNNs) [
80] and graph neural networks [
81], have shown substantial improvements under conditions such as cell division, merging, or sudden changes in shape. For example, deep learning models have been applied to track breast cancer cells in migration assays, even under challenging conditions with occlusions and rapid motion [
6].
More recently, models like DeepSea [
82], Cellpose [
23], and Omnipose [
83] have integrated segmentation and tracking capabilities into unified pipelines. These models offer robustness across diverse imaging modalities and cell types. In particular, Omnipose extends Cellpose by improving segmentation of irregularly shape cells, making it especially valuable for bacterial and morphologically diverse datasets.
Three-dimensional Segmentation and Volumetric Analysis refers to the analysis of three-dimensional imaging data acquired from modalities such as confocal microscopy, light-sheet fluorescence microscopy, and electron microscopy. These techniques produce volumetric datasets where biological structures extend across multiple optical sections. Accurate 3D segmentation and reconstruction are essential for quantitative analysis of tissue architecture, subcellular organization, and organoid morphology [
84]. To address this, 2D deep learning models have been extended into 3D, with architectures like 3D U-Net [
26] and V-Net [
85]. These approaches employ volumetric convolutions to capture spatial context in all three dimensions. Due to the high computational demands of volumetric data, specialized strategies such as patch-based training, tiling, multi-scale approaches, and hybrid 2D/3D pipelines are often adopted [
22,
86]. These models have provided detailed insights of complex biological system. For instance, they have been applied to tasks like segmentation of cell nuclei in z-stacks [
87], synapse detection in electron microscopy volumes [
88], and neuronal circuit tracing [
89].
2.5. Instance Segmentation
As introduced in the previous section, instance segmentation refers to the task of identifying and delineating individual objects within an image, assigning unique labels to each detected instance.
Table 1 provides an overview of the most representative models in this area.
Mask R-CNN [
8] sets a strong foundation by combining object detection and pixel-level segmentation using a region-based approach. Since then, numerous architectures have emerged, focusing on refining mask quality, improving instance separation, or enhancing efficiency. Notably, earlier works such as adversarial and recurrent models [
90] employed convolutional LSTM structures to capture spatial and temporal features, marking some of the first deep learning attempts in biomedical instance segmentation.
In histopathology, instance segmentation is vital for identifying nuclei, glands, and tissue compartments. These are key for cancer grading, tumor analysis, and digital pathology pipelines. HoVer-Net [
91] emerged as a landmark model in this field by predicting horizontal and vertical distance maps to better separate clustered nuclei, achieving strong performance across several histological datasets. Meanwhile, Mesmer [
24], trained in multiplexed images, demonstrated strong generalizability across tissues and imaging protocols, enabling the automated extraction of key cellular characteristics, such as subcellular location of protein signal.
In microscopy, where images include diverse cell types acquired via various modalities like fluorescence, brightfield, or phase-contrast, the instance segmentation task presents unique challenges. Some of them are overlapping cells, low contrast, and highly variable shapes. Custom variants of Mask R-CNN and U-Net hybrids have been widely used to balance precise localization with segmentation accuracy. More specialized models like StarDist [
10] introduced a star-convex polygon representation for segmenting nuclei in fluorescence images, greatly improving accuracy in crowded environments. Cellpose [
23] leveraged vector flow fields to robustly delineate cells across multiple modalities and morphologies, later extended by Cellpose 2.0 [
92] and Cellpose 3.0 [
97], which offer interactive training and support for 3D segmentation. BriFiSeg [
94] addressed the complexity of gland segmentation through a multi-scale approach tailored to accommodate variable gland morphologies. CPP-Net [
95] proposed a contour proposal network to enhance instance segmentation in densely packed cell images. LACSS [
93] introduced a weakly supervised framework leveraging image-level annotations for effective segmentation. Omnipose [
83] enhanced performance on bacterial and irregularly shaped cells by modeling more flexible object contours and improving boundary localization. Cellulus [
96] recently combined self-supervised learning and multi-scale features for improved segmentation of heterogeneous microscopy datasets.
As segmentation demands continue to grow, particularly in applications with limited annotations or new imaging modalities, recent research has turned toward foundation models. These large-scale pretrained models aim to provide general-purpose segmentation capabilities with minimal fine-tuning. In the following section, we explore the emergence of foundation models in biomedical image segmentation and their adaptation to microscopy data.
2.6. Foundation Models
Foundation models represent a transformative paradigm in computer vision, defined by their large scale, versatility, and strong generalization capabilities across diverse tasks with minimal fine-tuning. These models are typically pre-trained on massive datasets using self-supervised learning objectives, enabling them to learn broad, general-purpose visual representations that can be adapted to downstream applications such as segmentation, classification, and detection through prompt-based or lightweight tuning strategies [
98]. Many of these models leverage transformer-based architectures, including the Vision Transformer (ViT), which has shown remarkable ability to capture long-range dependencies and contextual information in images, further boosting model generalization [
99]. This architectural shift plays a key role in enabling foundation models to transfer effectively across domains, even when domain-specific annotated data is scarce.
A notable example is the Segment Anything Model (SAM), developed by Meta AI [
9]. Trained on over one billion masks from 11 million images, SAM introduces a highly flexible prompting interface that accepts inputs such as points, bounding boxes, or masks. This design supports both interactive and automated segmentation across an enormous wide range of image types and domains. SAM’s impressive zero-shot generalization capabilities make it particularly appealing for biomedical applications, where labeled data is often limited and manual annotation is very time-consuming.
Building on SAM’s foundation, researchers have begun tailoring it to better address the unique challenges of biomedical images (see
Table 2). For instance, Cellpose-SAM combines SAM’s mask generation strengths with the domain-specific expertise of Cellpose [
12]. Similarly, Cell-SAM employs domain-specific training strategies to refine segmentation outputs, accommodating the dense, low-contrast, and morphologically diverse structures characteristic of cellular microscopy [
11]. MicroSAM focuses specially on micro-scale cellular details, enhancing sensitivity to detect contours and faint boundaries frequently found in brightfield and label-free imaging modalities [
100]. Vista 2D [
101] is another recent adaptation that improves segmentation of 2D microscopy images by integrating SAM with contrastive learning techniques, further boosting robustness under challenging imaging conditions. These adaptations underscore that while SAM provides a powerful generalist base, incorporating biological priors and specialized knowledge is crucial for achieving high-precision biomedical segmentation.
Despite these advances, challenges remain, including improving model robustness to noisy or low-quality images and reducing false positives in densely packed cellular environments. Nevertheless, foundation models are rapidly becoming indispensable tools in biological image analysis. As ongoing research continues to incorporate domain-specific refinements, these models are expected to surpass traditional segmentation methods in accuracy, adaptability, and efficiency.
2.7. Datasets
High-quality datasets are fundamental to the development and benchmarking of instance segmentation algorithms in biomedical research. To support effective model training, such datasets must include high-resolution images alongside instance-level annotations, where each individual cell (or nucleus) is assigned a unique and non-overlapping mask. These annotations are crucial for evaluating, not just whether the correct regions are segmented, but also whether individual cells are properly separated, especially in crowded or complex tissue environments.
However, creating reliable instance segmentation datasets in microscopy presents multiple challenges. For instance, manual annotation of individual cells requires substantial domain expertise to accurately delineate cell boundaries, particularly in cases involving overlapping structures, low contrast, or irregular shapes. This is further intensified in certain modalities like brightfield or phase-contrast microscopy, where boundaries are often poorly defined. In addition, annotation consistency across large datasets can be difficult to maintain due to inter-annotator variability and subjective interpretation of ambiguous boundaries. Such inconsistencies and label noise can reduce model generalizability, especially in those that rely heavily on clean supervision.
Early progress in dataset development was stronger in histology than microscopy, particularly for tasks involving nuclear and tissue segmentation. In 2016, a breast cancer histopathology dataset that remains influential in studies involving H&E-stained tissue classification and segmentation [
102]. A few years later, the MoNuSeg dataset [
103] extended this effort by providing manually annotated nuclear masks across diverse tissue types, serving as a benchmark for both segmentation and generalization studies. More recently, MoNuSAC [
104] and Pannuke [
105] datasets introduced multi-class instance-level annotations of nuclei, enabling evaluation of instance segmentation performance across multiple nuclear categories. Similarly, NuCLS delivered large-scale annotations of nuclei in breast cancer slides, combining crowd-sourced and expert-labeled data to improve label quality and scale [
106]. Collectively, these datasets have helped establish benchmarks for deep learning models applied to clinical and histopathological data.
Microscopy cell segmentation research has been propelled by the release of numerous publicly available datasets encompassing a wide range of imaging modalities, cell types, and annotation styles.
Table 3 summarizes representative microscopy datasets for segmentation, highlighting the imaging modalities. One of the earlier widely adopted resources was the Data Science Bowl (DSB18) dataset [
27], introduced through a Kaggle challenge and offering annotated fluorescence microscopy images of nuclei from diverse experimental settings. It remains a foundational benchmark for nuclear segmentation. The Cellpose dataset [
23] expanded the diversity of available data by including a broad set of cell types imaged using fluorescence, brightfield, and phase-contrast microscopy. The dataset features hand-annotated masks curated for generalist model development across modalities. Around the same time, LIVECell [
107] was introduced, providing high-resolution phase-contrast images across multiple live cell lines along with dense instance masks, designed to support segmentation and tracking in time-lapse imaging. In addition, TissueNet [
24] extended these efforts to tissue-scale fluorescence microscopy, comprising tens of thousands of immunofluorescence images from human tissues spanning over 60 anatomical and disease contexts, and providing high-quality nuclear and whole-cell annotations to train robust and generalizable segmentation models.
Subsequently, the NeurIPS Cell Segmentation Challenge dataset [
29] was created to evaluate segmentation models under cross-domain conditions. It includes images from various microscopy modalities, tissue types, and staining protocols, serving as a rigorous benchmark for domain generalization in instance segmentation tasks. In parallel, domain-specific datasets continued to emerge. DeepBacs [
108] further advanced bacterial segmentation by including diverse species and imaging conditions—synthetic, brightfield, and phase-contrast—and offering finely detailed instance masks. It also introduced a suite of test sets designed to evaluate generalization across biological and technical domains. Complementing these, the EVICAN dataset [
109] provides extensive brightfield images of mammalian cells, addressing challenges related to label-free segmentation with diverse morphologies and imaging conditions.
Together, these datasets form a comprehensive ecosystem for benchmarking instance segmentation in microscopy, ranging from traditional fluorescence and live-cell imaging to more challenging bacterial and tissue-level tasks. Their continued development supports the advancement of robust, generalizable segmentation algorithms for biomedical research.
Creating high-quality annotated datasets for cell segmentation remains a major challenge due to the need for expert labeling, especially with overlapping cells, heterogeneous tissues, and low-contrast modalities like brightfield or phase-contrast. Variability in annotations and label noise can hinder model performance and generalization. Limited availability of 3D and time-lapse annotated data further restricts progress. Competitions such as the Data Science Bowl [
27], ISBI Cell Tracking Challenge [
6], and NeurIPS Cell Segmentation Challenge [
29] have helped by providing standardized datasets and benchmarks, yet the gap between existing datasets and the complexity of real-world microscopy remains significant.
2.8. Segmentation Evaluation Metrics
Evaluating the performance of instance segmentation models is essential to ensure reliable and reproducible results, especially in biomedical applications where accuracy in cell boundary detection directly impacts downstream analyses. Several standard metrics are commonly used to assess segmentation quality, each capturing different aspects of instance-level accuracy.
Intersection over Union (IoU) is a fundamental metric that quantifies the overlap between a predicted mask and the corresponding ground truth mask. It is defined as:
where
A is the predicted object region and
B is the ground truth region. A higher IoU indicates a better match between predicted and true cell boundaries. IoU is often used with a threshold to determine whether a predicted object is considered a true positive.
Dice Score is another widely used metric that measures the similarity between the predicted and ground truth masks. It is defined as:
Like IoU, Dice score ranges from 0 to 1, where 1 indicates perfect overlap. Dice score is especially useful in biomedical segmentation tasks due to its sensitivity to both false positives and false negatives, making it suitable for imbalanced datasets.
F1-score at a specific IoU threshold (e.g., 0.5) is commonly used to balance precision and recall:
This metric is sensitive to both over-segmentation and under-segmentation, making it suitable for evaluating dense cellular environments.
Average Precision (AP) summarizes the precision–recall trade-off across different IoU thresholds. In instance segmentation, AP is often computed as the mean precision over a range of IoU thresholds, commonly from 0.5 to 0.95 in steps of 0.05, known as AP@[.5:.95]:
where
r denotes recall,
is the precision at a given recall under a specific threshold, and
represents the set of evaluated Intersection-over-Union (IoU) thresholds. AP provides a comprehensive view of model performance by integrating both detection quality and segmentation accuracy. It is widely used in computer vision challenges such as COCO and adapted in biomedical evaluations.
Panoptic Quality (PQ) is a comprehensive metric that jointly evaluates segmentation quality and recognition performance. It combines the effects of true positives (TPs), false positives (FPs), and false negatives (FNs). PQ is calculated as:
where
p and
g represent matched predicted and ground truth instances. PQ effectively balances object detection and segmentation quality, making it robust across complex datasets.
Aggregated Jaccard Index (AJI) is another widely used metric in biomedical image segmentation that accounts for the intersection and union of matched objects while penalizing unmatched false positives. It is defined as:
where
is the
i-th ground truth instance,
is the corresponding predicted match,
U is the set of unmatched predictions, and
is an unmatched predicted object. AJI provides a global view of segmentation performance over the entire image and is especially useful for datasets with many touching or overlapping cells.
These metrics are typically computed over datasets such as DSB18 or the NeurIPS Cell Segmentation Challenge to allow fair comparisons across models. It is crucial to interpret these scores in the context of the biological task at hand, as small segmentation errors might have negligible or significant impact depending on the downstream application.
2.9. Challenges and Future Directions
Despite major advances in cell segmentation driven by deep learning, there are still numerous challenges that limit widespread deployment and generalization of existing methods. These challenges span data availability, domain adaptation, model scalability, evaluation consistency, and practical applicability in real-world biological workflows.
Data diversity and annotation efforts remain as main obstacles. While datasets like DSB18, the Cellpose dataset, and the NeurIPS Cell Segmentation Challenge have supported the development and benchmarking of new models, they often cover a narrow range of imaging modalities, staining protocols, and biological conditions. Manual annotation of instance segmentation masks is time-consuming and requires expert knowledge. These difficulties are compounded in 3D and time-lapse data, which are essential for understanding dynamic cellular processes but are still very underrepresented in public datasets.
Generalization and domain shift pose further difficulties. Many models perform well within the distribution of their training data but degrade significantly when applied to different cell types, imaging systems, or sample preparations. Although some generalist approaches such as Cellpose aim to address this issue, even these models often require fine-tuning or manual correction in unfamiliar contexts. Building segmentation tools that are robust across biological domains remains a central goal.
Scalability to complex data is also a limitation. While many models work well on 2D fluorescence images, they often struggle with 3D volumes, temporal sequences, or densely packed tissues. Segmentation in these settings requires both increased computational resources and specialized model architectures. Recent efforts have made progress toward integrating 3D instance segmentation and tracking, but unified models that work reliably across spatial and temporal scales are still lacking.
Evaluation inconsistencies are another source of difficulty in comparing models fairly. Different datasets often employ different annotation styles and use varied metrics such as Intersection-over-Union (IoU), average precision (AP), or F1-score, which makes direct comparisons challenging. Moreover, segmentation accuracy is not always indicative of biological utility: small errors in delineating boundaries can significantly affect downstream tasks like cell counting or spatial analysis in tissues.
To address these issues, the task of cell segmentation in microscopy images is moving toward several promising directions:
Foundation models adapted to microscopy, such as microSAM or Cell-SAM, are emerging as a flexible solution to limited annotated data and domain-specific variation.
Self-supervised learning and synthetic data generation are being explored to reduce annotation dependence and improve robustness.
Uncertainty quantification and interpretability are gaining traction to support use in clinical and high-stakes biological research.
End-to-end frameworks that link segmentation to downstream tasks (e.g., cell tracking, spatial analysis, classification) offer potential for more integrated analysis pipelines.
Instance segmentation plays a pivotal role in medical imaging, as it enables the precise delineation of individual cells, nuclei, or anatomical structures—an essential step for quantitative analysis, disease diagnosis, and treatment planning. Given the fragmented landscape of models, datasets, and evaluation practices in cell segmentation, a systematic comparison of representative state-of-the-art methods under controlled conditions is urgently needed. Such a benchmark would provide insights into how models perform across cell types and imaging modalities when trained and tested consistently. In the next section, we address this gap by comparing leading segmentation models—including traditional deep learning methods and recent hybrid approaches—on a curated set of microscopy datasets. Our aim is to establish a fair and transparent evaluation that can inform future model development and help practitioners choose appropriate tools for their specific biological tasks.