Skip to Content
SensorsSensors
  • Article
  • Open Access

25 September 2026

31 Pages

FCVGAN: A Novel Generative Framework Integrating Fuzzy C-Means Clustering with Conditional VAE-GAN for Polyp Segmentation

and
Perception, Robotics, and Intelligent Machines (PRIME), Department of Computer Science, Université de Moncton, Moncton, NB E1A 3E9, Canada
*
Author to whom correspondence should be addressed.

Abstract

Early detection of colorectal cancer is of paramount importance. Polyp segmentation using deep learning models can play an important role in this process; however, these models require large annotated datasets that are difficult to obtain in medical imaging. In this work, we propose FCVGAN (Fuzzy C-Means Conditional Variational Autoencoder Generative Adversarial Network), a mask-conditional generative framework that combines Fuzzy C-Means clustering, a Conditional VAE, and a SPADE-based GAN with a multi-scale discriminator. We evaluate two augmentation strategies across four polyp datasets: per-dataset training and transfer learning from Kvasir-SEG. The segmentation experiments repeated across five independent seeds show that transfer augmentation is the strategy that achieved the highest mean Dice score on three of the four datasets, while per-dataset augmentation performed best on Kvasir-Sessile. Compared with the corresponding baselines, transfer augmentation improved the mean Dice score by 2.4%, 0.8%, 16.0%, and 6.2% on Kvasir-SEG, CVC-ClinicDB, Kvasir-Sessile, and ETIS-Larib, respectively. The statistical analysis also showed large effect sizes for transfer augmentation over the baseline across all datasets. When performing the generation quality analysis, per-dataset training achieved a better average FID. However, the images generated with transfer learning showed lower similarity to the training data, which suggests that visual fidelity alone is not a complete indicator of augmentation effectiveness. The ablation study across the five configurations also showed that the contribution of the FCVGAN components depends on the dataset, with the full FCVGAN model providing the most balanced performance across the evaluated datasets. Overall, FCVGAN provides an effective framework for synthetic polyp augmentation, while transfer learning offers a practical strategy for supporting multiple target datasets using a single trained generator.

1. Introduction

Colorectal cancer (CRC) is considered to be a major health issue and remains one of the most commonly diagnosed malignancies and the second leading cause of cancer death worldwide, with around 1.9 million new cases and about 930,000 deaths each year [1]. The majority of CRCs follow a well-known adenoma–carcinoma sequence, where benign precancerous lesions gradually develop into malignant tumors over about 10–15 years [2]. This extended progression window presents a critical opportunity for early intervention: when CRC is detected at a localized stage, the five-year survival rate exceeds 90%, whereas diagnosis at a distant (metastatic) stage reduces five-year survival to approximately 14% [3,4].
Colonoscopy is still the gold standard for detecting and removing polyps. From a sensing perspective, the endoscope can be considered as a visual sensing system, where optical components and image sensors continuously capture images and videos of the gastrointestinal tract in real time. The quality of these acquired images can vary, depending on illumination, motion, blur, noise, specular reflections, and differences between endoscopic devices. These variations can directly affect the quality of the visual information provided to automated computer-aided systems. Studies also report adenoma miss rates of 22–28%, with flat and sessile polyps being particularly prone to being missed [5].
Deep learning has been considered as a promising solution to support clinical decision-making in colonoscopy by automatically analyzing the images acquired from these endoscopic sensing systems. Convolutional neural networks (CNNs) have demonstrated strong performance in polyp detection and segmentation, with recent methods achieving Dice scores approaching 90% on benchmark datasets [6,7]. Several recent works in biomedical sensing have also investigated deep learning for automatic analysis of endoscopic images, showing the growing role of intelligent image processing in endoscopic sensing systems [8,9]. However, these models can still fail to generalize to unseen datasets acquired from different sources. This is mainly because they are often trained and evaluated on data from single centers or relatively homogeneous populations, leading to performance degradation when applied across different clinical environments and imaging conditions [10].
Poor generalization is caused by the limited size and diversity of annotated colonoscopy datasets. In contrast to natural-image tasks that provide millions of labeled samples, medical imaging datasets typically include only hundreds to a few thousand annotated images [11]. The most widely used polyp segmentation benchmarks are Kvasir-SEG (1000 images) [12], CVC-ClinicDB (612 images) [13], and ETIS-Larib (196 images) [14]. Collectively, they provide fewer than 2000 annotated samples, which is only a small fraction of what deep learning models typically need. Additionally, getting more data is neither cheap nor simple: expert annotation is time-consuming and expensive, patient privacy must be carefully protected, and the polyps that matter most clinically, flat and serrated lesions with higher malignant potential, are precisely the ones that appear least often in available datasets [15]. Traditional data augmentation techniques, including geometric transformations (rotation, flipping, scaling) and photometric adjustments (brightness, contrast), have been widely used for data augmentation. While being effective to some extent, as they can create variations of the data, these techniques cannot generate new samples with novel pathological or anatomical information [16]. Consequently, traditional augmentation fails to address the core challenge: the absence of diverse polyp morphologies, textures, and clinical presentations in the original dataset.
Variational Autoencoders (VAEs), introduced by Kingma [17], provide a probabilistic approach to learn structured latent representations, which enables controllable generation and smooth interpolation between samples. By introducing probabilistic latent variables, VAEs offer control over the generation process; however, their reliance on reconstruction-based objectives often leads to blurrier outputs. Generative Adversarial Networks (GANs), on the other hand, are generative models that can produce sharp and realistic images through adversarial training; however, they lack an explicit latent structure, which makes controlled generation more challenging [18]. Hybrid architectures that combine VAEs with GANs have shown promise in balancing sample diversity with image quality [19]. In medical imaging, both approaches have been successfully applied to synthesize brain MRIs, chest X-rays, and histopathology slides, with studies demonstrating improved downstream task performance when synthetic data augments real training sets [20,21,22]. For colonoscopy specifically, several works have explored GAN-based polyp synthesis. Adjei et al. [23] modified Pix2Pix to generate synthetic polyp images. As a result, they achieved improvements in segmentation performance. More recently, conditional GANs with semantic guidance have shown good results in controlling polyp location, size, and shape [22,24]. However, many existing approaches still generate samples with limited diversity, as they rely on deterministic mask-to-image mappings, rather than an explicit modeling of the underlying data distribution via a probabilistic latent space.
Overview and contributions. For polyp synthesis, existing methods depend on deterministic mappings which, for a given segmentation mask, tend to produce nearly the same image each time, which inherently limits the generation of diverse samples and fails to capture the rich textural variability observed in real data. The key insight is simple: although polyp images appear highly heterogeneous, their appearances tend to group into a limited set of recurring texture patterns. These patterns can be learned without manual annotation and then used to guide the synthesis of more diverse and realistic images. To the best of our knowledge, this is the first approach that integrates fuzzy clustering into the latent space of a VAE-GAN architecture for polyp data augmentation. Our main contributions are as follows. (1) Texture-aware conditional synthesis. In this study we introduce FCVGAN, a generative approach that integrates Fuzzy C-Means clustering within a conditional VAE-GAN. FCVGAN is able to separate texture characteristics from spatial structure as it learns cluster assignments alongside image generation. This allows it to produce diverse polyp appearances while still following the shape defined by the input mask. (2) Hybrid generative architecture. Our architecture is hybrid as it integrates different components, including FCM clustering within a conditional VAE-GAN with SPADE normalization. This design helps to achieve two main things: first, it helps preserve fine-grained details that are considered challenging in polyp synthesis. Second, the probabilistic latent representation allows the model to introduce variations in the generated images while preserving the spatial structure defined by the input mask. (3) Comprehensive cross-dataset evaluation. We conducted experiments using four publicly available polyp datasets. We evaluate both per-dataset augmentation and a transfer setting in which the generative model trained on Kvasir-SEG is directly applied to target datasets without fine-tuning. This allows us to study whether a single trained generator can produce useful synthetic samples across different target distributions. The downstream segmentation experiments were further repeated across five independent seeds to evaluate the stability of the observed improvements. (4) Systematic ablation analysis. Through an ablation study, we evaluate five model configurations: the full FCVGAN, FCM+GAN without the CVAE branch, CVAE-GAN without FCM clustering, CVAE without adversarial training, and GAN-only without both CVAE and FCM. The results show that the contribution of the different components depends on the target dataset. In particular, removing the CVAE reduces the average segmentation improvement from +13.92% to +7.29% and causes a substantial performance degradation on ETIS-Larib. Overall, the full FCVGAN provides the most balanced performance across the evaluated datasets.

3. Methodology

We present FCVGAN, a hybrid architecture that integrates Fuzzy C-Means clustering into a conditional VAE-GAN framework for mask-conditional polyp image synthesis. Figure 1 illustrates the complete architecture comprising three main components: an encoder with integrated FCM clustering, a SPADE-based generator, and a multi-scale discriminator.
Figure 1. Overview of the proposed FCVGAN architecture. (A) The encoder extracts features from concatenated image–mask pairs and produces two latent components: z vae through variational sampling and z fcm through fuzzy clustering. (B) The generator progressively upsamples the combined latent code using SPADE residual blocks conditioned on the input mask. (C) The multi-scale discriminator evaluates realism at full and half resolutions using PatchGAN with feature matching.

3.1. Problem Formulation

Given a binary segmentation mask m ∈ 0 , 1 H × W , our objective is to generate a realistic colonoscopy image x ^ ∈ R 3 × H × W that follows the spatial structure defined by the input polyp mask. Rather than learning a single deterministic mapping from mask to image, the model learns the conditional distribution p ( x | m ) , which makes it possible to produce multiple realistic images from the same mask. The latent code of the proposed architecture has two main components: a variational component z vae capturing continuous appearance variations, and a clustering component z fcm encoding discrete texture patterns.

3.2. Encoder

The encoder E maps the image and the mask pair to a structured latent representation (Figure 1A). Then, the concatenated input [ x ; m ] ∈ R 4 × 256 × 256 is processed through six convolutional blocks with channel dimensions 64 → 128 → 256 → 512 → 512 → 512 , each using 4 × 4 kernels with stride 2, batch normalization, and LeakyReLU. The resulting features are flattened to f ∈ R 8192 .

3.2.1. Variational Latent Space

Two linear layers project f to the mean μ and log-variance log σ 2 of a Gaussian posterior. The latent code is sampled via reparameterization:
z vae = μ + σ ⊙ ϵ , ϵ ∼ N ( 0 , I )

3.2.2. Fuzzy C-Means Module

A third linear layer projects f to a clustering space f fcm ∈ R 128 . We use K = 12 clusters with a fuzzification parameter m = 2 . The cluster centers { c k } k = 1 K are initialized from a Gaussian distribution N ( 0 , 0 . 01 2 ) and are treated as learnable parameters.
For each feature representation f fcm , i , we first compute the Euclidean distance d i k between the feature and each cluster center c k . The fuzzy membership of sample i to cluster k is then computed as:
u i k = d i k − 2 m − 1 ∑ j = 1 K d i j − 2 m − 1 ,
where m > 1 controls the degree of fuzziness and is fixed to m = 2 in our experiments. In the implementation, this expression is computed as a softmax over − 2 m − 1 log ( d i k ) for numerical stability. Therefore, the softmax operation is only used to implement the FCM membership equation and does not introduce an additional temperature parameter τ .
The FCM latent representation is then obtained as the weighted combination of the cluster centers:
z fcm , i = ∑ k = 1 K u i k c k .
To encourage compact fuzzy clusters, we use the following FCM compactness loss:
L FCM = 1 N K ∑ i = 1 N ∑ k = 1 K u i k m d i k .
We additionally use an entropy regularization term to avoid the concentration of all samples into a limited number of clusters. Let u ¯ k = 1 N ∑ i = 1 N u i k denote the average membership of cluster k. The entropy regularization is defined as:
L ent = log K + ∑ k = 1 K u ¯ k log ( u ¯ k ) .
Unlike the conventional iterative FCM algorithm, the cluster centers in our architecture are not updated using alternating closed-form updates. Instead, they are optimized jointly with the encoder through backpropagation using the Adam optimizer. Therefore, the proposed module can be considered as a differentiable adaptation of FCM that allows fuzzy clustering to be integrated into the end-to-end generative framework.

3.3. Generator

The generator G synthesizes images from z combined conditioned on mask m (Figure 1B). During training, z combined is obtained from the encoder as described in Section 3.2. During transfer inference, the encoder is bypassed and z combined is sampled directly from a standard Gaussian distribution, as detailed in Section 4.2. A fully connected layer projects the latent code to a 1024 × 4 × 4 feature map, which is progressively upsampled to 256 × 256 through seven SPADE residual blocks. Each SPADE block applies spatially adaptive normalization [40] where modulation parameters are derived from the input mask:
SPADE ( h , m ) = γ ( m ) ⊙ h − μ ( h ) σ ( h ) + β ( m )
γ ( m ) and β ( m ) are produced by small convolutional networks taking the resized mask as input. This mechanism ensures polyp boundaries are preserved throughout generation. The final layer applies a 3 × 3 convolution with Tanh activation.

3.4. Multi-Scale Discriminator

We employ a multi-scale PatchGAN discriminator [39] operating at two scales (Figure 1C). Both discriminators receive the concatenation of image and mask as input. D 0 operates at full resolution ( 256 × 256 ), while D 1 operates at half resolution via average pooling. Each discriminator has convolutional layers with spectral normalization that produce patch-level predictions ( 30 × 30 and 14 × 14 respectively). This multi-scale evaluation provides complementary feedback on both fine textures and global structure.

3.5. Training Objectives

The training objective combines several loss terms to jointly optimize the generator, encoder, FCM module, and discriminator. We use the hinge loss for adversarial training. Since the discriminator operates at two different scales, the discriminator loss is computed as:
L D = ∑ s = 0 1 E max ( 0 , 1 − D s ( x , m ) ) + E max ( 0 , 1 + D s ( x ^ , m ) ) ,
where x is the real image, x ^ is the generated image, m is the segmentation mask, and D s represents the discriminator at scale s. The corresponding adversarial loss for the generator is defined as:
L adv = − ∑ s = 0 1 E D s ( x ^ , m ) .
We also use a feature matching loss to encourage the generated images to produce intermediate discriminator representations similar to those of real images. This loss is defined as:
L FM = 1 S ∑ s = 0 1 ∑ l D s ( l ) ( x , m ) − D s ( l ) ( x ^ , m ) 1 ,
where D s ( l ) represents the feature representation from layer l of the discriminator at scale s, and S = 2 is the number of discriminator scales.
To preserve the overall appearance of the original image, we include an L 1 reconstruction loss
L recon = x − x ^ 1 .
We also use a VGG19-based perceptual loss to preserve higher-level visual features between the real and generated images:
L perc = ∑ l w l ϕ l ( x ) − ϕ l ( x ^ ) 1 ,
where ϕ l represents the features extracted from layer l of the pre-trained VGG19 network and w l represents the corresponding layer weight.
For the variational component, the KL divergence is used to regularize the latent distribution toward a standard Gaussian distribution:
L KL = − 1 2 E 1 + log σ 2 − μ 2 − exp ( log σ 2 ) .
In addition, the FCM compactness loss L FCM and the entropy regularization loss L ent , defined in Section 3.2.2, are included during training. The FCM loss encourages the latent features to remain close to their cluster centers according to their fuzzy memberships, while the entropy loss helps avoid the concentration of the samples into only a limited number of clusters.
The complete objective used to train the generator, encoder, and FCM module is therefore:
L G = λ adv L adv + λ FM L FM + λ recon L recon + λ perc L perc + λ KL L KL + λ FCM L FCM + λ ent L ent .

4. Experiments

4.1. Datasets

To evaluate the effectiveness of our proposed framework, we conducted experiments on four datasets for polyp segmentation. Figure 2 presents representative image–mask pairs from the four datasets:
Figure 2. Representative samples from the four polyp segmentation datasets. Each panel shows two colonoscopy images (top) with corresponding ground truth masks (bottom). (a) Kvasir-SEG: large, prominent polyps with varied morphology. (b) CVC-ClinicDB: polyps captured under controlled imaging conditions. (c) Kvasir-Sessile: flat sessile polyps with subtle boundaries. (d) ETIS-Larib: small polyps with challenging imaging conditions including specular reflections.
  • Kvasir-SEG [12] is a large-scale polyp segmentation dataset containing 1000 colono scopy images with corresponding pixel-wise ground truth masks. This dataset has considerable variation in polyp appearance, including differences in size, shape, color, and texture. The images have variable resolutions ranging from 332 × 487 to 1920 × 1072 pixels. In our experiments, the Kvasir-SEG dataset is the source dataset for training the generative model.
  • CVC-ClinicDB [13] contains 612 frames extracted from 29 colonoscopy video sequences. All images have a resolution of 384 × 288 pixels and were acquired at the Hospital Clinic of Barcelona. This dataset is widely used as a benchmark for polyp segmentation algorithms and contains polyps with varying sizes and morphological characteristics.
  • Kvasir-Sessile [12] is a specialized dataset containing 196 images of sessile (flat) polyps. These polyps are particularly challenging to detect and segment due to their subtle elevation and similarity in appearance to the surrounding mucosa. Sessile polyps are clinically significant as they have a higher risk of malignancy compared to pedunculated polyps. This dataset provides an important test case for evaluating segmentation methods on difficult polyp morphologies.
  • ETIS-Larib [14] consists of 196 high-resolution colonoscopy images (1225 × 966 pixels). This dataset is considered one of the most challenging benchmarks due to significant variations in polyp size, shape, texture, and the presence of artifacts such as specular reflections and motion blur.
Table 1 summarizes the key characteristics of each dataset. For all experiments, we applied stratified sampling based on polyp size categories to ensure balanced representation in the training, validation, and test splits. Polyp size was computed as the ratio of polyp pixels to total image pixels, with samples that were grouped into small (<5%), medium (5–15%), or large (>15%). Each dataset was then divided into training (70%), validation (15%), and test (15%) sets, while preserving the same polyp-size distribution across splits to avoid the bias toward specific morphologies.
Table 1. Summary of polyp segmentation datasets used in our experiments.

4.2. Implementation Details

All experiments were implemented in PyTorch 2.0 and conducted on a single NVIDIA GPU with CUDA acceleration. Mixed-precision training using automatic mixed precision (AMP) was employed to reduce the memory consumption and accelerate computation without compromising model performance.
The generator was trained using the Adam optimizer with a learning rate 1 × 10 − 4 , while the discriminator used learning rate 4 × 10 − 4 , both with β 1 = 0.0 and β 2 = 0.999 . Loss weights were set as λ a d v = 1.0 , λ F M = 10.0 , λ r e c o n = 10.0 , λ p e r c = 10.0 , λ K L = 0.05 , λ F C M = 0.5 , and λ e n t = 0.01 . The number of epochs was set to 150 epochs with early stopping based on validation reconstruction loss, using a patience of 20 epochs. The batch size was set to 16.
For the downstream segmentation task, we employed UNet++ with an encoder featuring five levels of 32, 64, 128, 256, and 512 channels respectively. The segmentation model was trained using the Adam optimizer with learning rate 1 × 10 − 4 and a combined loss function of Dice loss and binary cross-entropy with equal weights. Also, for the segmentation, the number of epochs was set to 150 epochs with early stopping based on validation Dice score (patience of 20 epochs). A learning rate scheduler reduced the learning rate by a factor 0.5 when validation performance plateaued for 5 consecutive epochs.
All images were resized to 256 × 256 pixels and normalized to the range [ − 1 , 1 ] for generative model training. For segmentation, images were normalized using ImageNet statistics (mean = [0.485, 0.456, 0.406], std = [0.229, 0.224, 0.225]). For data augmentation during training, we included random horizontal flipping, random vertical flipping, and random rotation by multiples of 90 degrees, applied consistently to both images and masks.
For synthetic data generation, we generated one synthetic image per training mask, so we doubled the training set size. In the transfer learning protocol, the generative model trained on Kvasir-SEG was directly applied to the target datasets without any fine-tuning. In this setting, the target image is not required and the encoder is bypassed. Instead, a 256-dimensional latent vector is sampled directly from a standard Gaussian distribution, z combined ∼ N ( 0 , I ) , and then provided to the generator together with the target segmentation mask, such that x ^ = G ( z combined , m ) . Therefore, the segmentation mask is the only input required from the target dataset, while the latent vector is sampled internally. The mask controls the spatial structure of the generated image through the SPADE layers, while the sampled latent vector introduces appearance variation. In this transfer setting, the encoder and the FCM module are not explicitly used during inference.

4.3. Computational Cost and Efficiency

We performed an analysis of the computational requirements to assess the practical efficiency and deployability of our proposed architecture FCVGAN. Table 2 presents the parameter distribution across the main architectural components. The encoder E has 14.25 M parameters: 11.1 M for convolutional feature extraction, 1.05 M for the FCM module (clustering layer and projection), and 2.1 M for the VAE latent projections ( μ and log σ 2 heads). The generator G (SPADE decoder) accounts for 96.2 M parameters and represents the largest component of the architecture due to the repeated high-dimensional convolutional operations and learnable modulation networks required for spatially adaptive normalization. The multi-scale discriminator D has 14.0 M parameters. Hence the total of parameters is 124.45 M trainable parameters during adversarial training.
Table 2. Parameter distribution and active parameter counts of FCVGAN.
The FCM module has only 1.05 M parameters, corresponding to approximately 0.84% of the total model. In contrast, the SPADE-based generator accounts for approximately 77% of the total parameter count, reflecting the computational capacity required for mask-conditioned image synthesis.
Table 3 shows the computational cost of each FCVGAN component in terms of floating-point operations (FLOPs) per forward pass. As presented in the table, the SPADE-based generator requires most of the computation, with 130.18 GFLOPs per forward pass, and the encoder requires 2.02 GFLOPs, while the discriminator requires 3.25 GFLOPs. During transfer inference and as only the generator is used, the cost is 130.18 GFLOPs per generated image. In the per-dataset setting, the encoder is also used, which increases the cost slightly to 132.20 GFLOPs per image.
Table 3. Computational cost (GFLOPs) per forward pass of FCVGAN components.
The training was performed on a single NVIDIA RTX 4090 GPU. On the Kvasir-SEG dataset, each epoch took about 20 s, and training stopped after 104 epochs using early stopping, for a total training time of about 35 min. Compared to other diffusion-based methods such as Polyp-DDPM and ArSDM, FCVGAN is more efficient during image generation because it produces a synthetic image in a single forward pass instead of using several denoising steps.
At inference, the required components depend on which generation strategy is used. In the per-dataset setting, the encoder is used to obtain z FCM and z VAE , so both the encoder and generator are active, corresponding to 110.45M parameters. In the transfer setting, only the target segmentation mask is available. In this case, z combined ∼ N ( 0 , I ) is sampled directly and passed to the generator together with the mask. The encoder is not used, reducing the active parameter count to 96.2M. In this setting, the FCM module is used during training to organize the latent representation and guide the learning of the generator, but it is not used directly during inference. In the transfer setting, FCVGAN generates one 256 × 256 synthetic image in about 18 ms, which corresponds to roughly 55 images per second. This makes it suitable for generating large numbers of synthetic samples efficiently. In contrast, diffusion-based models require several denoising steps for each generated image.
Although FCVGAN generates images quickly, the model remains relatively large, mainly because of the SPADE decoder. This may make it harder to use on devices with limited memory or computing resources. Future work could explore techniques such as knowledge distillation, channel pruning, or lighter SPADE designs to reduce the model size while maintaining image quality.

4.4. Evaluation Metrics

We evaluate our approach using metrics for both generation quality and downstream segmentation performance.

4.4.1. Segmentation Metrics

For evaluation of segmentation performance we used two complementary metrics that capture different aspects of spatial overlap between predicted and ground truth masks. The Dice Similarity Coefficient (DSC), also known as the F1 score, measures the harmonic mean of precision and recall and is defined as:
DSC = 2 | P ∩ G | | P | + | G |
where P denotes the set of pixels predicted as polyp and G denotes the ground truth polyp pixels. The Dice coefficient ranges from 0 (no overlap) to 1 (perfect agreement) and is widely adopted in medical image segmentation due to its robustness to class imbalance.
We used also the Intersection over Union (IoU), also known as the Jaccard index, and it provides a stricter measure of overlap by computing the ratio of intersection to union:
IoU = | P ∩ G | | P ∪ G |
The IoU metric is informative in particular for segmentation quality in clinical applications as it penalizes both false positives and false negatives more heavily than Dice.

4.4.2. Generation Quality Metrics

To assess the quality of synthetic images, we use the following metrics:
(a)
Fréchet Inception Distance (FID): measures the distributional similarity between real and generated images in the feature space of a pre-trained InceptionV3 network [58]. Lower FID indicates higher visual quality and realism:
FID = | | μ r − μ g | | 2 + Tr ( Σ r + Σ g − 2 ( Σ r Σ g ) 1 / 2 )
where ( μ r , Σ r ) and ( μ g , Σ g ) are the mean and covariance of real and generated feature distributions, respectively.
(b)
Structural Similarity Index (SSIM): [59] measures perceptual similarity between image pairs based on luminance, contrast, and structure, ranging from −1 to 1, with higher values indicating greater similarity.
(c)
Peak Signal-to-Noise Ratio (PSNR): quantifies reconstruction quality in decibels (dB), with higher values indicating lower pixel-wise distortion.

4.4.3. Memorization Metric

Following [60], we evaluate potential memorization by computing the maximum Pearson correlation between each synthetic image and all real training images. For descriptive analysis, correlations above 0.9 are treated as a warning of strong similarity and potential memorization, while values below 0.8 indicate lower similarity to the training data. Pearson correlation is used here as a practical similarity indicator rather than as a complete test of memorization.

4.4.4. Statistical Analysis

To evaluate the consistency of the segmentation improvements, all segmentation experiments were repeated using five independent random seeds. For each seed, we use the mean test Dice score as one independent observation, resulting in five observations for each experimental setting. The same seeds were used across the baseline, transfer augmentation, and per-dataset augmentation settings to enable paired comparisons.
Since the number of the independent runs is limited, we used the two-sided paired Wilcoxon signed-rank test instead of assuming a normal distribution. In addition to the p-values, we report the rank-biserial correlation as an effect size and 95% bootstrap confidence intervals for the Dice differences between the compared methods. A Bonferroni correction is also applied to account for multiple pairwise comparisons across the four datasets.

5. Results

We evaluate FCVGAN using two training strategies: (1) transfer learning, where the model is trained on Kvasir-SEG and applied to generate synthetic images for all datasets, and (2) per-dataset training, where separate models are trained for each dataset. Our evaluation includes qualitative visual assessment, downstream segmentation performance, generation quality metrics, memorization analysis, and also an ablation study, to provide a thorough characterization of the proposed approach.
Figure 3 presents synthetic polyp images generated by both approaches across all four datasets. As visually observed in the figure, there are distinct characteristics between the two strategies. The per-dataset approach (right panel) produces images that are similar to real ones with more accurate color reproduction and texture consistency. In contrast, in the transfer learning approach (left panel) we see noticeable color shifts and artifacts (yellowish or greenish), especially on target datasets such as CVC-ClinicDB and Kvasir-Sessile.
Figure 3. Qualitative comparison of synthetic polyp images generated by FCVGAN. Left (transfer learning): The model is trained on Kvasir-SEG (source) and applied to all datasets. Right (per-dataset training): Separate models are trained for each dataset. Rows: (a) Kvasir-SEG, (b) CVC-ClinicDB, (c) Kvasir-Sessile, and (d) ETIS-Larib. The per-dataset approach produces more visually consistent images, while transfer learning shows color artifacts on target datasets but generates more diverse samples.
Despite these visual differences, both approaches have successfully preserved the structural characteristics that are defined by the input masks, generating anatomically plausible polyp images with appropriate boundaries and shapes. This observation gives an important question: does higher visual fidelity translate to better downstream task performance? The following subsections will try to give an answer based on segmentation experiments, generation quality metrics, and diversity analysis.

5.1. Segmentation Performance

We evaluate the impact of synthetic data augmentation on polyp segmentation using UNet++ as the backbone network. For each dataset, we compare baseline training (using only real images) against augmented training (combining real images with synthetic images generated by FCVGAN) with two augmentation strategies: transfer and per-dataset. All configurations were evaluated using identical train, validation, and test splits. The experiments were conducted across five independent random seeds and the results are reported as the mean Dice coefficient ± standard deviation across the five runs. Table 4 summarizes the segmentation performance obtained with the three training configurations across the four datasets.
Table 4. Comparison of baseline, transfer, and per-dataset FCVGAN augmentation strategies. Results are reported as mean ± standard deviation of the Dice coefficient across five independent training runs.
Looking at the results presented on the table, we notice that both augmentation strategies have improved the mean Dice coefficient over the baseline across all of the four benchmark datasets with variation on the magnitude between datasets.
On the larger datasets, which are Kvasir-SEG and CVC-ClinicDB, the transfer strategy has achieved the highest segmentation performance. For Kvasir-SEG, transfer augmentation increased the mean Dice coefficient from 0.8318 to 0.8518 , which corresponds to an improvement of 2.4 % , compared with 1.7 % for the per-dataset strategy. A similar trend was observed on CVC-ClinicDB, where transfer augmentation reached a Dice coefficient of 0.8887 compared with 0.8830 for per-dataset augmentation and 0.8812 for the baseline, and despite the fact that these improvements on these two datasets were relatively modest, both augmentation strategies maintained performance above the corresponding baseline models.
Larger improvements were observed on the smaller and more challenging Kvasir-Sessile and ETIS-Larib datasets. On the Kvasir-Sessile dataset, the per-dataset augmentation has achieved the largest improvement, increasing the mean Dice coefficient from 0.4015 to 0.4832 , corresponding to a gain of 20.3 % , and for the transfer augmentation, it also produced a substantial improvement on this dataset, reaching a Dice coefficient of 0.4659 , or 16.0 % above the baseline. On ETIS-Larib, transfer augmentation achieved the highest Dice coefficient of 0.7443 , compared with 0.7388 for per-dataset augmentation and 0.7007 for the baseline, corresponding to improvements of 6.2 % and 5.4 % , respectively. These results show basically that synthetic augmentation provides its largest relative benefit when the baseline segmentation performance is lower and also the amount of available training data is limited.
The variability across the five independent training runs also differed between datasets. On the Kvasir-SEG and ETIS-Larib datasets, the augmented models showed lower standard deviations compared with the baseline, especially for the transfer strategy. However, this behavior was not observed consistently on CVC-ClinicDB and Kvasir-Sessile, where the augmented models have exhibited similar or greater variability. Therefore, the main benefit of FCVGAN augmentation is actually reflected in the increase in mean segmentation performance rather than in a uniform reduction in training variability.
When the two augmentation strategies are compared directly, their average relative improvements across the four datasets are similar, reaching approximately 6.4 % for transfer augmentation and 6.9 % for per-dataset augmentation. However, these averages hide an important difference in their behavior. The transfer augmentation has achieved the highest mean Dice coefficient on three of the four datasets, which are Kvasir-SEG, CVC-ClinicDB, and ETIS-Larib, whereas the per-dataset augmentation has achieved the best result in only one dataset which is the Kvasir-Sessile dataset. For that, and from this observation we can say the results support the use of transfer augmentation as the preferred strategy when the same generative framework is intended to support several target datasets as it achieved the best segmentation performance on three out of four datasets while avoiding the additional cost of training a separate generator for each domain. The per-dataset training can also be considered as a useful alternative when the target data present distinct characteristics and the additional training cost is acceptable.

5.2. Statistical Comparison of Augmentation Strategies

A statistical analysis was also performed across the five independent training runs to further examine the differences between the baseline, transfer, and per-dataset configurations, as summarized in Table 5. For each random seed, we obtained a single mean Dice coefficient on the fixed test set, and the comparisons were therefore performed at the training-run level, with results obtained using the same random seed treated as paired observations. As the number of runs is limited, a two-sided Wilcoxon signed-rank test was used without assuming normality of the paired differences. Three comparisons were considered for each dataset: baseline versus transfer, baseline versus per-dataset, and transfer versus per-dataset. The Wilcoxon statistic (W) and the corresponding raw p-values were reported, and a Bonferroni correction was applied across the 12 pairwise comparisons. To complement the hypothesis tests, the matched-pairs rank-biserial correlation ( r r b ) was used to quantify the magnitude and direction of the paired effect. The mean paired Dice difference ( Δ Dice) and its 95% bootstrap confidence interval were also calculated. Finally, the confidence intervals were obtained by resampling the five seed-level paired differences and therefore they describe uncertainty across independent training runs rather than patient-level uncertainty.
Table 5. Statistical comparison of baseline, transfer, and per-dataset augmentation across five independent training runs. Δ Dice represents the mean paired difference between the second and first strategy in each comparison. CI denotes the 95% bootstrap confidence interval of the paired seed-level difference, W is the Wilcoxon signed-rank statistic, p adj is the Bonferroni-adjusted p-value, and r r b is the matched-pairs rank-biserial correlation.
None of the pairwise comparisons has reached statistical significance after correction for multiple comparisons. However, this result should be interpreted in relation to the limited number of independent training runs. With five observations, the smallest possible two-sided Wilcoxon p-value is 0.0625, which means that even if all five runs favor the same strategy, the test cannot reach p < 0.05 .
This was the case for Kvasir-SEG. Transfer augmentation improved the mean Dice coefficient by 0.0200 over the baseline, and all five paired runs favored the transfer strategy. This produced the maximum rank-biserial effect size of r r b = 1.000 , while the Wilcoxon test gave p = 0.0625 . The confidence interval was also fully positive, [0.0105, 0.0285], supporting the consistency of the improvement.
A similar trend was observed on the other datasets. Transfer augmentation showed large positive effects over the baseline on CVC-ClinicDB ( r r b = 0.733 ), Kvasir-Sessile ( r r b = 0.867 ), and ETIS-Larib ( r r b = 0.733 ). The largest mean improvements were observed on Kvasir-Sessile ( Δ Dice = 0.0644) and ETIS-Larib ( Δ Dice = 0.0437).
The per-dataset strategy also showed large positive effects over the baseline on the Kvasir-SEG, Kvasir-Sessile, and ETIS-Larib datasets. On the CVC-ClinicDB dataset, however, we see that the effect was negligible ( r r b = − 0.067 ), with a very small mean difference of 0.0018 and a confidence interval that included zero.
The direct comparison between transfer and per-dataset augmentation showed smaller differences. Transfer was moderately favored on Kvasir-SEG and CVC-ClinicDB ( r r b = − 0.467 for both), while per-dataset augmentation was moderately favored on Kvasir-Sessile ( r r b = 0.333 ), and for the ETIS-Larib dataset, the effect was negligible ( r r b = 0.067 ), which indicates no clear advantage for either strategy across the five runs.
Overall, the statistical results agree with the segmentation results reported in Table 4. Transfer augmentation showed a large positive effect over the baseline on all four datasets and achieved the highest mean Dice score on three of them.
Although no direct comparison between the transfer and per-dataset augmentation has reached statistical significance, the overall results support that the transfer augmentation is to be considered as the preferred practical strategy when one generator is intended to support multiple target datasets. Per-dataset training remains a useful alternative when a target dataset benefits from a dedicated generator, as observed on Kvasir-Sessile.

5.3. Robustness Under Degraded Image Conditions

To further evaluate our approach and to examine the behavior of the segmentation models under changes in image quality, we conducted a robustness experiment on the ETIS-Larib dataset. The baseline, transfer-augmented, and per-dataset-augmented models obtained from the five independent training runs were evaluated without any additional training. The ETIS-Larib was the dataset selected as it is considered as a challenging evaluation case and because of its relatively small training set and lower baseline segmentation performance.
We considered for this experiment, four types of image degradation: Gaussian blur, Gaussian noise, reduced brightness, and reduced contrast. Gaussian blur was applied with a radius of 1.5 and Gaussian noise with a standard deviation of 0.03. Brightness and contrast were reduced using factors of 0.75 and 0.70, respectively. These transformations were applied only to the test images, while the corresponding ground-truth masks remained unchanged. Figure 4 illustrates the different image conditions used in this experiment.
Figure 4. Examples of the image conditions used in the robustness evaluation on ETIS-Larib. From left to right: clean image, Gaussian blur, Gaussian noise, reduced brightness, and reduced contrast.
Robustness was evaluated using the Dice coefficient under both clean and degraded conditions. In addition to the absolute Dice score, the performance drop was calculated as
Δ drop = D i c e clean − D i c e degraded ,
where a smaller value indicates better preservation of the segmentation performance after degradation. Results are reported as the mean ± standard deviation across the five independent training runs and are summarized in Table 6.
Table 6. Robustness evaluation on ETIS-Larib under different image degradation conditions. Dice values are reported as mean ± standard deviation across five independent training runs. Δ represents the performance drop relative to the clean condition.
As expected, all degradation conditions reduced segmentation performance compared with the clean images. However, the effect of the degradations varied across the three training configurations. Under the Gaussian blur degradation, the transfer strategy has achieved the highest Dice coefficient of 0.4957 , compared with 0.4845 for the per-dataset strategy and 0.4206 for the baseline. Transfer also showed the smallest performance drop under this condition ( Δ drop = 0.2486 ), which indicates better preservation of its clean performance.
Under Gaussian noise, the baseline model achieved the highest Dice coefficient ( 0.2113 ), followed by transfer ( 0.1899 ) and per-dataset augmentation ( 0.1739 ). Therefore, neither augmentation strategy improved robustness over the baseline under this type of degradation. However, when the two augmented models are compared directly, transfer augmentation remained more robust than the per-dataset strategy.
A different behavior was observed when the brightness and the contrast of the images were reduced. Under reduced brightness, the per-dataset model achieved the highest Dice coefficient ( 0.2482 ), followed by transfer ( 0.2008 ), while the baseline dropped to 0.0947 . A similar pattern was observed under reduced contrast, where per-dataset augmentation reached 0.2851 , transfer reached 0.2578 , and the baseline decreased to 0.0567 . The smaller performance drops obtained by the augmented models under these two conditions indicate a robustness benefit against intensity-related image degradation.
The comparison also between the two augmentation strategies also shows different robustness characteristics. Transfer augmentation achieved higher Dice coefficients than the per-dataset strategy under the clean, blurred, and noisy conditions. In contrast, the per-dataset strategy performed better under reduced brightness and reduced contrast, and this suggests that transfer augmentation provides a stronger general performance across several conditions, while the per-dataset training can offer an additional advantage under changes related to illumination and contrast.
Overall, the robustness experiment shows that the benefit of FCVGAN augmentation depends on the type of image degradation. Transfer augmentation provided the best overall performance on clean and blurred images and outperformed the per-dataset strategy under three of the five evaluated image conditions. The per-dataset augmentation was more effective under reduced brightness and the contrast degradations, while the Gaussian noise degradation remained the most challenging case for both of the augmented models. These results show that the proposed augmentation framework is robust under several image-quality variations, while also showing that no single augmentation strategy is uniformly superior under all types of degradation.

5.4. Generation Quality

To understand the factors contributing to segmentation improvements, we analyze the visual quality of generated images using three metrics: Fréchet Inception Distance (FID), Structural Similarity Index (SSIM), and Peak Signal-to-Noise Ratio (PSNR). Table 7 presents these metrics for both training strategies across all datasets.
Table 7. Generation quality metrics comparing per-dataset and transfer learning approaches. Lower FID indicates better quality. Best results per dataset in bold.
Per-dataset training achieves substantially lower FID scores across most datasets, with an average of 135.48 compared to 173.43 for transfer learning. Lower FID indicates that the distribution of the generated images more closely matches the distribution of real images, reflecting higher visual fidelity. The difference is particularly noticeable on CVC-ClinicDB (107.27 vs. 189.37, a difference of 82.10) and Kvasir-SEG (68.91 vs. 126.58, a difference of 57.67). These results quantitatively support the visual observations presented in Figure 3, where per-dataset training produces images with more realistic colors and textures. The FID scores also reveal dataset-specific differences. Kvasir-SEG achieves the lowest FID score (68.91 for per-dataset training), indicating a closer distribution between the generated and real images. In contrast, ETIS-Larib shows the highest FID scores for both approaches (195.11 and 191.02), indicating greater difficulty in reproducing its image characteristics. Interestingly, transfer learning slightly outperforms per-dataset training on ETIS-Larib (191.02 vs. 195.11). This suggests that knowledge transferred from the larger Kvasir-SEG dataset can provide some benefit for generation on this challenging target dataset.
For SSIM and PSNR, moderate values were obtained for both approaches, ranging from 0.29 to 0.54 for SSIM and from 11.25 to 14.49 dB for PSNR. These metrics measure structural and pixel-level similarity between generated and real images. However, these values should be interpreted in the context of FCVGAN, which performs mask-conditional generation rather than reconstruction of a specific real image. Therefore, high pixel-wise similarity to an individual reference image is not necessarily expected. These metrics should consequently be considered together with FID, qualitative evaluation, and downstream segmentation performance when assessing generation quality.
A critical observation is obtained when comparing the generation quality with the segmentation performance. Despite achieving lower FID scores, the per-dataset training does not consistently lead to better downstream segmentation. Transfer augmentation achieved the highest mean Dice coefficient on three of the four datasets, while per-dataset training achieved the highest result on Kvasir-Sessile. This finding suggests that visual quality, as measured by FID, does not directly predict augmentation effectiveness. Other characteristics of the generated images, including their similarity to the original training samples and their diversity, may also influence their usefulness for segmentation augmentation. We investigate this relationship further through the memorization analysis presented in the following subsection.

5.5. Memorization Analysis

The observation that better visual quality does not necessarily translate into better segmentation performance motivates a deeper investigation of the diversity of the generated samples. Generative models, particularly when trained on limited datasets, may produce samples that closely resemble their training examples rather than introducing sufficiently new variations. In a data augmentation setting, such behavior can reduce the additional information provided to the downstream segmentation model.
Following the correlation-based analysis used by Akbar et al. [60], we quantified similarity by computing the maximum Pearson correlation between each synthetic image and all real training images. For each generated image, the highest correlation with any training image was retained. Therefore, lower values indicate that the generated image is less similar to its nearest training example, while high values indicate stronger similarity and a greater possibility of memorization. For descriptive interpretation, correlations below 0.8 were considered to indicate high diversity, values between 0.8 and 0.9 moderate similarity, and values above 0.9 were considered a warning of strong similarity to the training data.
Table 8 and Figure 5 present the results for the two training strategies. A clear difference can be observed between transfer and per-dataset training.
Table 8. Memorization analysis based on the maximum Pearson correlation between each synthetic image and the real training images. Lower values indicate more diverse generation.
Figure 5. Memorization analysis comparing per-dataset and transfer learning approaches. Lower maximum correlation indicates greater diversity with respect to the training images.
Transfer learning achieves lower mean maximum correlations across all four datasets, with an average of 0.767 compared with 0.926 for per-dataset training. The difference is particularly large on CVC-ClinicDB, where the mean correlation decreases from 0.948 for per-dataset training to 0.683 for transfer learning. Similar differences are observed on ETIS-Larib (0.931 vs. 0.726) and Kvasir-Sessile (0.899 vs. 0.771). On Kvasir-SEG, the difference is smaller, with correlations of 0.923 and 0.886 for per-dataset and transfer training, respectively.
The percentage of generated samples with a maximum correlation above 0.9 further illustrates this difference. For per-dataset training, 84.4% of Kvasir-SEG samples, 99.1% of CVC-ClinicDB samples, 49.3% of Kvasir-Sessile samples, and 86.7% of ETIS-Larib samples exceed this threshold. In contrast, the corresponding value for transfer learning is 35.9% on Kvasir-SEG and 0% on the three target datasets. These results indicate that the per-dataset models generally generate samples that are more similar to their training images, while transfer learning introduces greater image-level variation.
This observation provides an additional explanation for the relationship between generation quality and downstream segmentation. Although per-dataset training achieves better FID scores, its generated images also show stronger similarity to the corresponding training data. Transfer learning, in contrast, produces more diverse samples and achieved the highest mean segmentation Dice score on three of the four datasets. This suggests that visual fidelity alone is not sufficient to determine the usefulness of synthetic images for data augmentation. Diversity with respect to the original training data can also play an important role by introducing additional variations during segmentation training.
This behavior is particularly relevant for the smaller target datasets, where training a generative model using only the available target images provides a more limited amount of information from which to learn the image distribution. Transfer learning can reduce this limitation by first learning broader generative characteristics from the larger Kvasir-SEG dataset and then applying this learned representation to the target datasets.
It should be noted that the present analysis uses maximum pixel-level Pearson correlation as a practical indicator of similarity rather than as a complete test of memorization. Pearson correlation does not fully capture perceptual or feature-level similarity and does not separately analyze the polyp and background regions. A more extensive privacy or memorization assessment could include perceptual distances, deep feature-space similarity, nearest-neighbor visualization, and region-specific analysis. However, such analyses are beyond the scope of the present work. Here, the correlation analysis is used to provide a consistent comparison of the relative diversity produced by the two training strategies.
Finally, since FCVGAN performs mask-conditional generation, some structural similarity is naturally expected because the input mask constrains the spatial location and shape of the generated polyp. The relevant difference is therefore not the complete absence of similarity, but whether the model can preserve the required spatial structure while introducing variations in appearance, including texture, color, and local image details.

5.6. Ablation Study

To validate the contribution of each component in FCVGAN, we conduct an ablation experiment by systematically removing different architectural components. We have evaluated five configurations: (1) the full FCVGAN model, (2) FCM+GAN without the CVAE branch, (3) CVAE-GAN without FCM clustering, (4) CVAE without the GAN discriminator, and (5) GAN only without both CVAE and FCM. Table 9 presents the segmentation improvement obtained by the different configurations using the transfer learning approach.
Table 9. Ablation study comparing segmentation improvement (%) across different model configurations using transfer learning. Results correspond to a single training run for each configuration.
To keep the computational cost of the ablation study manageable, each ablation configuration was evaluated using a single training run. Therefore, the results reported in Table 9 should be interpreted as a component-wise comparison under the same experimental setting rather than as the five-run averages reported in the main segmentation experiments.
The full FCVGAN model achieves an average improvement of +13.92% across the four datasets. Positive improvements are obtained on all datasets, including +2.98% on Kvasir-SEG, +1.24% on CVC-ClinicDB, +37.43% on Kvasir-Sessile, and +14.02% on ETIS-Larib. In particular, the improvement obtained on ETIS-Larib shows that the complete architecture can provide useful augmentation on this challenging target dataset.
To better isolate the contribution of the CVAE beyond the FCM component, we also evaluate an FCM+GAN configuration in which the CVAE branch is removed while keeping the FCM and adversarial components. This configuration achieves improvements of +0.67% on Kvasir-SEG, +1.75% on CVC-ClinicDB, and +45.74% on Kvasir-Sessile. However, its performance decreases considerably on ETIS-Larib, with a degradation of −19.01%. The average improvement is +7.29%, compared with +13.92% for the full FCVGAN.
A direct comparison between FCM+GAN and the full FCVGAN provides a clearer view of the contribution of the CVAE. On Kvasir-SEG, the Dice score increases from 0.8397 with FCM+GAN to 0.8590 with the full model. The largest difference is observed on ETIS-Larib, where the Dice score increases from 0.5412 to 0.7619 when the CVAE is included. On the other hand, FCM+GAN performs slightly better on CVC-ClinicDB (0.8976 vs. 0.8932) and Kvasir-Sessile (0.5082 vs. 0.4792). Therefore, the contribution of the CVAE is not uniform across all datasets, but it provides an important benefit for the overall performance and particularly for the generalization to ETIS-Larib.
Removing FCM clustering from the full architecture (CVAE-GAN) results in an average improvement of +11.28%. This configuration performs well on Kvasir-Sessile (+44.87%) and also improves Kvasir-SEG (+2.58%) and CVC-ClinicDB (+1.65%). However, it degrades the performance on ETIS-Larib by −4.00%. This suggests that the FCM component can contribute to the generalization of the model when transferring to a more challenging target distribution.
We also experiment with removing the GAN discriminator and keeping the CVAE configuration. In this case, the performance degrades substantially on ETIS-Larib, reaching −21.53%, while Kvasir-Sessile shows an improvement of +18.27%. The average improvement becomes −0.50%. This indicates that adversarial training plays an important role in generating images that are useful for downstream segmentation.
The GAN-only configuration achieves an average improvement of +7.51%. It improves CVC-ClinicDB by +2.36% and Kvasir-Sessile by +38.78%, but its performance decreases on ETIS-Larib by −11.36%. Therefore, although the GAN alone can provide useful augmentation on some datasets, its performance is less consistent across the different target domains.
Overall, the ablation results show that the contribution of the different components depends on the target dataset. The GAN discriminator plays an important role in maintaining useful image generation, while the FCM component contributes to the organization of the latent representation. The new FCM+GAN experiment also shows more clearly the contribution of the CVAE. Although removing the CVAE can still produce strong results on some datasets, especially Kvasir-Sessile, it considerably reduces the performance on ETIS-Larib and lowers the average improvement from +13.92% to +7.29%. These results suggest that the combination of FCM clustering, CVAE encoding, and adversarial training provides a more balanced augmentation framework across the evaluated datasets.

5.7. Comparison with Related Work

In this section, we compare our proposed approach (FCVGAN) with recent polyp synthesis methods for segmentation augmentation in Table 10. The comparison across methods is a bit challenging due to variations in segmentation backbones, training protocols, datasets, and evaluation metrics. Nevertheless, we organize the results by generative approach and highlight the main methodological differences. All improvement values ( Δ ) are computed from Dice scores. For FCVGAN, the reported Dice scores correspond to the mean performance across five independent training runs.
Table 10. Comparison with recent polyp synthesis methods for segmentation augmentation. Improvements ( Δ ) are relative to the corresponding baseline using real data only and are computed from Dice scores. FCVGAN results are reported as mean ± standard deviation across five independent runs. Methods vary in segmentation backbones and evaluation protocols, limiting direct comparison. Only Polyp-DDPM uses the same backbone (UNet++) as our method.
Among the compared methods, Polyp-DDPM [55] provides one of the most direct comparisons as it also employs UNet++ for segmentation evaluation. On Kvasir-SEG, FCVGAN achieves a mean Dice improvement of +2.4%, compared with +0.7% reported for Polyp-DDPM. Although the two methods use the same segmentation backbone, the training and evaluation protocols are not identical, and therefore this comparison should be interpreted with caution. Polyp-DDPM also reports performance degradation in some cross-dataset evaluations, while FCVGAN shows positive mean improvements across all four datasets evaluated in our transfer setting.
GAN-based approaches also show varying results. Adjei et al. [23] report a +3.0% Dice improvement on Kvasir-SEG using a modified Pix2Pix architecture with a U-Net backbone, which is comparable to the +2.4% mean improvement obtained by FCVGAN. However, their evaluation is limited to a single dataset without cross-dataset validation. MinimalGAN [42] reports improvements of +0.4% on CVC-ClinicDB and +2.2% on ETIS-Larib using PraNet as the segmentation backbone. FCVGAN achieves improvements of +0.8% and +6.2% on the same datasets, respectively. SemanticPolypGAN [52] also reports a +2.8% improvement on Kvasir-SEG using UACANet. These results show that GAN-based augmentation can provide useful improvements, although the magnitude of the benefit depends strongly on the dataset, segmentation backbone, and training protocol.
Among diffusion-based methods, ArSDM [54] achieves the highest absolute performance on CVC-ClinicDB with a Dice score of 0.933 and reports a substantial improvement of +20.2% on ETIS-Larib. However, its improvement on Kvasir-SEG is considerably smaller (+0.1%), which shows that the benefit of synthetic augmentation can vary substantially between datasets. ControlPolypNet [56] achieves an improvement of +2.95% on CVC-ClinicDB through ControlNet-based synthesis with a quality filtering mechanism, while its performance decreases by −0.66% on Kvasir-SEG. These results further show that even advanced generative approaches do not provide uniform improvements across all evaluation settings.
From this comparison, we gain several observations. First, methods using different segmentation backbones can achieve substantially different absolute Dice scores, which makes it difficult to separate the contribution of the generative model from the capability of the downstream segmentation network. Second, the effectiveness of synthetic augmentation is generally dataset-dependent, with several methods showing strong improvements on some datasets and limited or negative effects on others. Third, the experimental protocols differ considerably between studies, including the datasets, synthetic-to-real data ratios, segmentation architectures, and evaluation procedures, which limits direct ranking between the methods.
FCVGAN provides positive mean improvements across all four evaluated datasets under the transfer setting, with gains of +2.4%, +0.8%, +16.0%, and +6.2% on Kvasir-SEG, CVC-ClinicDB, Kvasir-Sessile, and ETIS-Larib, respectively. The largest relative improvement is observed on Kvasir-Sessile, which has the lowest baseline segmentation performance. Moreover, FCVGAN uses a single generator trained on Kvasir-SEG to augment several target datasets without additional generative model training. Overall, these results indicate that FCVGAN can provide useful synthetic augmentation across different target datasets, while the comparison with existing methods should be considered in light of the differences in backbones and experimental protocols.

6. Discussion and Conclusions

This study introduced FCVGAN, a generative architecture that combines Fuzzy C-Means clustering with a Conditional VAE-GAN for synthetic polyp image generation. Our evaluation across the four benchmark datasets provides several observations regarding the use of generative models for polyp segmentation augmentation.
The proposed architecture combines three main components. The SPADE-based generator enables control over the location, shape, and boundaries of the synthesized polyps by preserving the spatial information provided by the input segmentation masks. The multi-scale discriminator helps the generator reproduce realistic appearance and texture at different resolutions, while the FCM module organizes the latent representation through soft clustering. Together, these components allow FCVGAN to preserve the structure defined by the input masks while introducing variations in the generated images.
We evaluated two augmentation strategies: per-dataset training, where a separate generative model is trained for each dataset, and transfer learning, where a single FCVGAN model trained on Kvasir-SEG is used to generate synthetic images for all target datasets. Across five independent segmentation runs, transfer augmentation achieved the highest mean Dice score on three of the four datasets, with improvements of +2.4%, +0.8%, +16.0%, and +6.2% on Kvasir-SEG, CVC-ClinicDB, Kvasir-Sessile, and ETIS-Larib, respectively. Per-dataset augmentation performed best on Kvasir-Sessile, where an improvement of +20.3% was obtained. Although none of the pairwise comparisons remained statistically significant after correction for multiple comparisons, transfer augmentation showed large effect sizes compared with the baseline across all four datasets. Considering also that only five independent runs were available, these results are interpreted mainly based on the direction and magnitude of the effects rather than statistical significance alone.
An important observation is obtained when comparing generation quality with downstream segmentation performance. Per-dataset training achieved a better average FID score than transfer learning (135.48 compared with 173.43), indicating that its generated images were generally closer to the distribution of the corresponding real datasets. However, this did not consistently result in better segmentation performance, as transfer augmentation achieved the highest mean Dice score on three of the four datasets. The memorization analysis provides additional information about this behavior. Per-dataset generated images showed a higher average maximum correlation with the training data (0.926) compared with transfer-generated images (0.767), indicating stronger similarity to the training samples. These observations suggest that visual fidelity alone does not completely determine the usefulness of synthetic images for data augmentation and that diversity with respect to the original training data may also play an important role.
The robustness experiment further showed that the effect of synthetic augmentation depends on the type of image degradation. On ETIS-Larib, the transfer model achieved the highest Dice score on clean and blurred images, while the per-dataset model performed best under reduced brightness and contrast. Under Gaussian noise, however, the baseline model obtained the highest performance. Therefore, synthetic augmentation does not provide uniform robustness against all image-quality degradations, but it can improve performance under several acquisition-related variations.
The ablation study also showed that the contribution of the different FCVGAN components is dataset-dependent. The complete FCVGAN achieved an average improvement of +13.92% across the four datasets. Removing the CVAE branch while keeping FCM and adversarial training reduced the average improvement to +7.29%, with a particularly large degradation on ETIS-Larib. Removing FCM resulted in an average improvement of +11.28%, while the CVAE-only configuration produced an overall degradation of −0.50%. The GAN-only configuration achieved an average improvement of +7.51%. These results indicate that individual components can contribute differently depending on the target dataset. In particular, the complete FCVGAN provides the most balanced performance across the evaluated datasets rather than each component being uniformly beneficial in every case.
The comparison with related work further shows that the benefit of synthetic augmentation depends strongly on the generative method, segmentation backbone, dataset, and evaluation protocol. FCVGAN provides positive mean improvements across all four datasets in the transfer setting while using a single generator trained on Kvasir-SEG. However, direct ranking against existing methods should be interpreted carefully because the compared studies use different segmentation architectures and experimental settings.
From a sensing perspective, FCVGAN can be directly connected to the data-processing stage of endoscopic visual sensing. Colonoscopy images are acquired through endoscopic imaging systems, and the performance of automatic segmentation methods depends strongly on the quality and diversity of the acquired images. Variations in illumination, blur, noise, contrast, and other acquisition conditions can affect the visual information provided to the segmentation model. In this context, FCVGAN aims to complement the endoscopic sensing pipeline by generating additional mask-conditioned images that increase the diversity of the training data. The robustness experiments under different image degradations further show how the trained segmentation models behave when the quality of the acquired images is altered. Future work could extend this evaluation using images acquired from different endoscopic devices and acquisition settings to better study the generalization of the proposed approach across sensing systems.

7. Limitations and Future Work

Several limitations should be acknowledged. First, our evaluation relies only on publicly available benchmark datasets and uses a single segmentation backbone (UNet++). Therefore, the generalization of the observed improvements to other segmentation architectures and independent clinical datasets remains to be evaluated. In addition, patient- or video-level identifiers were not available in the dataset files used in our experiments, which prevented us from performing patient- or video-grouped resampling during the statistical analysis.
The transfer learning strategy also assumes the availability of a suitable source dataset. When the source and target domains differ substantially, the benefit of transfer augmentation may become more limited. Furthermore, the robustness analysis was performed using controlled image degradations rather than images acquired directly under different endoscopic devices or clinical acquisition conditions. Future work should therefore include external datasets collected from different endoscopic systems and imaging settings.
Another limitation concerns the memorization analysis. Maximum Pearson correlation was used as a practical indicator of image-level similarity, but it does not fully capture perceptual or feature-level similarity between generated and real images. Future work could extend this analysis using perceptual distances, deep feature representations, nearest-neighbor visualization, and region-specific similarity analysis.
FCVGAN also remains a relatively large generative model, mainly due to the SPADE-based generator. Although image generation is relatively fast during inference, future work could investigate lighter architectures, pruning, or knowledge distillation to reduce the computational requirements.
Finally, improved segmentation performance on benchmark datasets does not necessarily translate into improved clinical outcomes. This study did not evaluate adenoma detection rates, histological characterization, or endoscopist performance. Future work should include blinded evaluation of the generated images by multiple endoscopists, assessment of anatomical and pathological plausibility, and prospective validation on independent clinical data. Extending FCVGAN to other medical imaging modalities and incorporating domain adaptation strategies could also improve its applicability across different source and target domains.

Author Contributions

Conceptualization, N.M. and M.A.A.; methodology, N.M. and M.A.A.; software, N.M. and M.A.A.; validation, N.M. and M.A.A.; formal analysis, N.M. and M.A.A.; writing—original draft preparation, N.M.; writing—review and editing, N.M. and M.A.A.; visualization, N.M. and M.A.A.; supervision, M.A.A.; funding acquisition, M.A.A. All authors have read and agreed to the published version of the manuscript.

Funding

This research was enabled in part by support provided by the Natural Sciences and Engineering Research Council of Canada (NSERC), funding reference number RGPIN-2024-05287, and by the AI in Health Research Chair at the Université de Moncton.

Institutional Review Board Statement

Not applicable.

Data Availability Statement

The datasets used in this study are publicly available from the original sources cited in the manuscript. The synthetic data generated during this study have not been deposited in a public repository. Further inquiries can be directed to the corresponding author.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
CVAEConditional Variational Autoencoder
FCMFuzzy C-Means
FCVGANFuzzy C-Means Conditional Variational Autoencoder GAN
GANGenerative Adversarial Network
IoUIntersection over Union
KLKullback–Leibler
SPADESpatially Adaptive Denormalization
VAEVariational Autoencoder

References

  1. Morgan, E.; Arnold, M.; Gini, A.; Lorenzoni, V.; Cabasag, C.; Laversanne, M.; Vignat, J.; Ferlay, J.; Murphy, N.; Bray, F. Global burden of colorectal cancer in 2020 and 2040: Incidence and mortality estimates from GLOBOCAN. Gut 2023, 72, 338–344. [Google Scholar] [CrossRef] [Scilit]
  2. Simon, K. Colorectal cancer development and advances in screening. Clin. Interv. Aging 2016, 11, 967–976. [Google Scholar] [CrossRef] [Scilit]
  3. Siegel, R.L.; Miller, K.D.; Wagle, N.S.; Jemal, A. Cancer statistics, 2023. CA Cancer J. Clin. 2023, 73, 17. [Google Scholar] [CrossRef] [Scilit]
  4. Siegel, R.L.; Wagle, N.S.; Cercek, A.; Smith, R.A.; Jemal, A. Colorectal cancer statistics, 2023. CA Cancer J. Clin. 2023, 73, 233–254. [Google Scholar] [CrossRef] [Scilit]
  5. Kim, N.H.; Jung, Y.S.; Jeong, W.S.; Yang, H.J.; Park, S.K.; Choi, K.; Park, D.I. Miss rate of colorectal neoplastic polyps and risk factors for missed polyps in consecutive colonoscopies. Intest. Res. 2017, 15, 411. [Google Scholar] [CrossRef] [Scilit]
  6. Fan, D.P.; Ji, G.P.; Zhou, T.; Chen, G.; Fu, H.; Shen, J.; Shao, L. Pranet: Parallel reverse attention network for polyp segmentation. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Lima, Peru, 4–8 October 2020; pp. 263–273. [Google Scholar] [CrossRef] [Scilit]
  7. Jha, D.; Smedsrud, P.H.; Johansen, D.; De Lange, T.; Johansen, H.D.; Halvorsen, P.; Riegler, M.A. A comprehensive study on colorectal polyp segmentation with ResUNet++, conditional random field and test-time augmentation. IEEE J. Biomed. Health Inform. 2021, 25, 2029–2040. [Google Scholar] [CrossRef] [Scilit]
  8. Peng, C.; Qian, Z.; Wang, K.; Zhang, L.; Luo, Q.; Bi, Z.; Zhang, W. MugenNet: A novel combined convolution neural network and transformer network with application in colonic polyp image segmentation. Sensors 2024, 24, 7473. [Google Scholar] [CrossRef] [Scilit]
  9. Qiu, W.; Yang, X.; Liu, Z.; Qiu, C. Advancing real-time polyp detection in colonoscopy imaging: An anchor-free deep learning framework with adaptive multi-scale perception. Sensors 2025, 25, 7524. [Google Scholar] [CrossRef] [Scilit]
  10. Ali, S.; Ghatwary, N.; Jha, D.; Isik-Polat, E.; Polat, G.; Yang, C.; Li, W.; Galdran, A.; Ballester, M.Á.G.; Thambawita, V.; et al. Assessing generalisability of deep learning-based polyp detection and segmentation methods through a computer vision challenge. Sci. Rep. 2024, 14, 2032. [Google Scholar] [CrossRef] [Scilit]
  11. Litjens, G.; Kooi, T.; Bejnordi, B.E.; Setio, A.A.A.; Ciompi, F.; Ghafoorian, M.; Van Der Laak, J.A.; Van Ginneken, B.; Sánchez, C.I. A survey on deep learning in medical image analysis. Med. Image Anal. 2017, 42, 60–88. [Google Scholar] [CrossRef] [Scilit]
  12. Jha, D.; Smedsrud, P.H.; Riegler, M.A.; Halvorsen, P.; de Lange, T.; Johansen, D.; Johansen, H.D. Kvasir-seg: A segmented polyp dataset. In Proceedings of the International Conference on Multimedia Modeling; Springer: Berlin/Heidelberg, Germany, 2020; pp. 451–462. [Google Scholar]
  13. Jha, D.; Smedsrud, P.H.; Riegler, M.A.; Johansen, D.; De Lange, T.; Halvorsen, P.; Johansen, H.D. Resunet++: An advanced architecture for medical image segmentation. In Proceedings of the 2019 IEEE International Symposium on Multimedia (ISM); IEEE: New York, NY, USA, 2019; pp. 225–2255. [Google Scholar]
  14. Silva, J.; Histace, A.; Romain, O.; Dray, X.; Granado, B. Toward embedded detection of polyps in wce images for early diagnosis of colorectal cancer. Int. J. Comput. Assist. Radiol. Surg. 2014, 9, 283–293. [Google Scholar] [CrossRef] [Scilit]
  15. Jia, X.; Shen, Y.; Yang, J.; Song, R.; Zhang, W.; Meng, M.Q.H.; Liao, J.C.; Xing, L. PolypMixNet: Enhancing semi-supervised polyp segmentation with polyp-aware augmentation. Comput. Biol. Med. 2024, 170, 108006. [Google Scholar] [CrossRef] [Scilit]
  16. Shorten, C.; Khoshgoftaar, T.M. A survey on image data augmentation for deep learning. J. Big Data 2019, 6, 60. [Google Scholar] [CrossRef] [Scilit]
  17. Kingma, D.P.; Welling, M. Auto-encoding variational bayes. arXiv 2013, arXiv:1312.6114. [Google Scholar]
  18. Goodfellow, I.J.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; Bengio, Y. Generative adversarial nets. Adv. Neural Inf. Process. Syst. 2014, 27, 2672–2680. [Google Scholar]
  19. Larsen, A.B.L.; Sønderby, S.K.; Larochelle, H.; Winther, O. Autoencoding beyond pixels using a learned similarity metric. In Proceedings of the International Conference on Machine Learning, New York, NY, USA, 19–24 June 2016; pp. 1558–1566. [Google Scholar]
  20. Frid-Adar, M.; Diamant, I.; Klang, E.; Amitai, M.; Goldberger, J.; Greenspan, H. GAN-based synthetic medical image augmentation for increased CNN performance in liver lesion classification. Neurocomputing 2018, 321, 321–331. [Google Scholar] [CrossRef] [Scilit]
  21. Han, C.; Rundo, L.; Araki, R.; Nagano, Y.; Furukawa, Y.; Mauri, G.; Nakayama, H.; Hayashi, H. Combining noise-to-image and image-to-image GANs: Brain MR image augmentation for tumor detection. IEEE Access 2019, 7, 156966–156977. [Google Scholar] [CrossRef] [Scilit]
  22. Shin, H.C.; Tenenholtz, N.A.; Rogers, J.K.; Schwarz, C.G.; Senjem, M.L.; Gunter, J.L.; Andriole, K.P.; Michalski, M. Medical image synthesis for data augmentation and anonymization using generative adversarial networks. In Proceedings of the International Workshop on Simulation and Synthesis in Medical Imaging, Granada, Spain, 16 September 2018; pp. 1–11. [Google Scholar] [CrossRef] [Scilit]
  23. Adjei, P.E.; Lonseko, Z.M.; Du, W.; Zhang, H.; Rao, N. Examining the effect of synthetic data augmentation in polyp detection and segmentation. Int. J. Comput. Assist. Radiol. Surg. 2022, 17, 1289–1302. [Google Scholar] [CrossRef] [Scilit]
  24. Qadir, H.A.; Balasingham, I.; Shin, Y. Simple U-net based synthetic polyp image generation: Polyp to negative and negative to polyp. Biomed. Signal Process. Control 2022, 74, 103491. [Google Scholar] [CrossRef] [Scilit]
  25. Garcea, F.; Serra, A.; Lamberti, F.; Morra, L. Data augmentation for medical imaging: A systematic literature review. Comput. Biol. Med. 2023, 152, 106391. [Google Scholar] [CrossRef] [Scilit]
  26. Chen, Y.; Yang, X.H.; Wei, Z.; Heidari, A.A.; Zheng, N.; Li, Z.; Chen, H.; Hu, H.; Zhou, Q.; Guan, Q. Generative adversarial networks in medical image augmentation: A review. Comput. Biol. Med. 2022, 144, 105382. [Google Scholar] [CrossRef] [Scilit]
  27. Cacciola, M.; Angiulli, G.; Burrascano, P.; Laganà, F.; Versaci, M. A Prototypical Fuzzy Similarity-Based Classification Framework for Ultrasonic Defect Detection in Concrete. Eng 2026, 7, 88. [Google Scholar] [CrossRef] [Scilit]
  28. Rais, K.; Amroune, M.; Benmachiche, A.; Haouam, M.Y. Exploring variational autoencoders for medical image generation: A comprehensive study. arXiv 2024, arXiv:2411.07348. [Google Scholar]
  29. Yeoh, P.S.Q.; Hasikin, K.; Wu, X.; Goh, S.L.; Lai, K.W. Trends and applications of variational autoencoders in medical imaging analysis. Comput. Med. Imaging Graph. 2025, 126, 102647. [Google Scholar] [CrossRef] [Scilit]
  30. Liu, D.; Zhou, X. Abdominal CT Image Synthesis with Generative Autoencoders. Blog Post. 2021. Available online: https://www.researchgate.net/publication/364607293_Abdominal_CT_Image_Synthesis_with_Generative_Autoencoders (accessed on 12 August 2026).
  31. Cetin, I.; Stephens, M.; Camara, O.; Ballester, M.A.G. Attri-VAE: Attribute-based interpretable representations of medical images with variational autoencoders. Comput. Med. Imaging Graph. 2023, 104, 102158. [Google Scholar] [CrossRef] [Scilit]
  32. Chadebec, C.; Thibeau-Sutre, E.; Burgos, N.; Allassonnière, S. Data augmentation in high dimensional low sample size setting using a geometry-based variational autoencoder. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 45, 2879–2896. [Google Scholar] [CrossRef] [Scilit]
  33. Rguibi, Z.; Hajami, A.; Zitouni, D.; Maleh, Y.; Elqaraoui, A. Medical variational autoencoder and generative adversarial network for medical imaging. Indones. J. Electr. Eng. Comput. Sci. 2023, 32, 494–505. [Google Scholar] [CrossRef] [Scilit]
  34. Bao, J.; Chen, D.; Wen, F.; Li, H.; Hua, G. CVAE-GAN: Fine-grained image generation through asymmetric training. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; pp. 2745–2754. [Google Scholar]
  35. Diamantis, D.E.; Gatoula, P.; Iakovidis, D.K. Endovae: Generating endoscopic images with a variational autoencoder. In Proceedings of the 2022 IEEE 14th Image, Video, and Multidimensional Signal Processing Workshop (IVMSP), Nafplio, Greece, 26–29 June 2022; pp. 1–5. [Google Scholar] [CrossRef] [Scilit]
  36. Yi, X.; Walia, E.; Babyn, P. Generative adversarial network in medical imaging: A review. Med. Image Anal. 2019, 58, 101552. [Google Scholar] [CrossRef] [Scilit]
  37. Kazeminia, S.; Baur, C.; Kuijper, A.; Van Ginneken, B.; Navab, N.; Albarqouni, S.; Mukhopadhyay, A. GANs for medical image analysis. Artif. Intell. Med. 2020, 109, 101938. [Google Scholar] [CrossRef] [Scilit]
  38. Isola, P.; Zhu, J.Y.; Zhou, T.; Efros, A.A. Image-to-image translation with conditional adversarial networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 1125–1134. [Google Scholar]
  39. Wang, T.C.; Liu, M.Y.; Zhu, J.Y.; Tao, A.; Kautz, J.; Catanzaro, B. High-resolution image synthesis and semantic manipulation with conditional gans. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 8798–8807. [Google Scholar]
  40. Park, T.; Liu, M.Y.; Wang, T.C.; Zhu, J.Y. Semantic image synthesis with spatially-adaptive normalization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019; pp. 2337–2346. [Google Scholar]
  41. Skandarani, Y.; Jodoin, P.M.; Lalande, A. Gans for medical image synthesis: An empirical study. J. Imaging 2023, 9, 69. [Google Scholar] [CrossRef] [Scilit]
  42. Zhang, Y.; Wang, Q.; Hu, B. MinimalGAN: Diverse medical image synthesis for data augmentation using minimal training data. Appl. Intell. 2023, 53, 3899–3916. [Google Scholar] [CrossRef] [Scilit]
  43. Karras, T.; Aittala, M.; Hellsten, J.; Laine, S.; Lehtinen, J.; Aila, T. Training generative adversarial networks with limited data. Adv. Neural Inf. Process. Syst. 2020, 33, 12104–12114. [Google Scholar]
  44. Sohl-Dickstein, J.; Weiss, E.; Maheswaranathan, N.; Ganguli, S. Deep unsupervised learning using nonequilibrium thermodynamics. In Proceedings of the International Conference on Machine Learning, Lille, France, 6–11 July 2015; pp. 2256–2265. [Google Scholar]
  45. Kazerouni, A.; Aghdam, E.K.; Heidari, M.; Azad, R.; Fayyaz, M.; Hacihaliloglu, I.; Merhof, D. Diffusion models in medical imaging: A comprehensive survey. Med. Image Anal. 2023, 88, 102846. [Google Scholar] [CrossRef] [Scilit]
  46. Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 19–24 June 2022; pp. 10684–10695. [Google Scholar]
  47. Pinaya, W.H.; Tudosiu, P.D.; Dafflon, J.; Da Costa, P.F.; Fernandez, V.; Nachev, P.; Ourselin, S.; Cardoso, M.J. Brain imaging generation with latent diffusion models. In Proceedings of the MICCAI Workshop on Deep Generative Models, Singapore, 22 September 2022; pp. 117–126. [Google Scholar] [CrossRef] [Scilit]
  48. Konz, N.; Chen, Y.; Dong, H.; Mazurowski, M.A. Anatomically-controllable medical image generation with segmentation-guided diffusion models. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Marrakesh, Morocco, 6–10 October 2024; pp. 88–98. [Google Scholar] [CrossRef] [Scilit]
  49. Song, J.; Meng, C.; Ermon, S. Denoising diffusion implicit models. arXiv 2020, arXiv:2010.02502. [Google Scholar] [CrossRef] [Scilit]
  50. Song, Y.; Dhariwal, P.; Chen, M.; Sutskever, I. Consistency models. In Proceedings of the 40th International Conference on Machine Learning; Proceedings of Machine Learning Research, 202; JMLR.org: Norfolk, MA, USA, 2023; pp. 32211–32252. [Google Scholar]
  51. Shin, Y.; Qadir, H.A.; Balasingham, I. Abnormal colon polyp image synthesis using conditional adversarial networks for improved detection performance. IEEE Access 2018, 6, 56007–56017. [Google Scholar] [CrossRef] [Scilit]
  52. Song, H.; Shin, Y. Semantic polyp generation for improving polyp segmentation performance. J. Med. Biol. Eng. 2024, 44, 280–292. [Google Scholar] [CrossRef] [Scilit]
  53. Yoon, D.; Kong, H.J.; Kim, B.S.; Cho, W.S.; Lee, J.C.; Cho, M.; Lim, M.H.; Yang, S.Y.; Lim, S.H.; Lee, J.; et al. Colonoscopic image synthesis with generative adversarial network for enhanced detection of sessile serrated lesions using convolutional neural network. Sci. Rep. 2022, 12, 261. [Google Scholar] [CrossRef] [Scilit]
  54. Du, Y.; Jiang, Y.; Tan, S.; Wu, X.; Dou, Q.; Li, Z.; Li, G.; Wan, X. Arsdm: Colonoscopy images synthesis with adaptive refinement semantic diffusion models. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Vancouver, Canada, 8–12 October 2023; pp. 339–349. [Google Scholar] [CrossRef] [Scilit]
  55. Dorjsembe, Z.; Pao, H.K.; Xiao, F. Polyp-ddpm: Diffusion-based semantic polyp synthesis for enhanced segmentation. In Proceedings of the 2024 46th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), Orlando, FL, USA, 15–19 July 2024; pp. 1–7. [Google Scholar] [CrossRef] [Scilit]
  56. Sharma, V.; Kumar, A.; Jha, D.; Bhuyan, M.K.; Das, P.K.; Bagci, U. Controlpolypnet: Towards controlled colon polyp synthesis for improved polyp segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 2325–2334. [Google Scholar]
  57. Hu, K.; Zhao, L.; Feng, S.; Zhang, S.; Zhou, Q.; Gao, X.; Guo, Y. Colorectal polyp region extraction using saliency detection network with neutrosophic enhancement. Comput. Biol. Med. 2022, 147, 105760. [Google Scholar] [CrossRef] [Scilit]
  58. Heusel, M.; Ramsauer, H.; Unterthiner, T.; Nessler, B.; Hochreiter, S. GANs trained by a two time-scale update rule converge to a local Nash equilibrium. In Proceedings of the Advances in Neural Information Processing Systems 30, Long Beach, CA, USA, 4–9 December 2017; pp. 6626–6637. [Google Scholar]
  59. Wang, Z.; Bovik, A.C.; Sheikh, H.R.; Simoncelli, E.P. Image quality assessment: From error visibility to structural similarity. IEEE Trans. Image Process. 2004, 13, 600–612. [Google Scholar] [CrossRef] [Scilit]
  60. Usman Akbar, M.; Wang, W.; Eklund, A. Beware of diffusion models for synthesizing medical images—A comparison with GANs in terms of memorizing brain MRI and chest x-ray images. Mach. Learn. Sci. Technol. 2025, 6, 015022. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.