A Taxonomy of Generative Models with a Focus on Diffusion Models and Denoising Techniques
Abstract
1. Introduction
- Position diffusion models within a broader taxonomy of generative models, including explicit likelihood-based models, variational or score-based models, and implicit models.
- Provide a structured taxonomy of diffusion models based on key design dimensions, including representation space, forward process formulation, reverse parameterization, sampling strategy, and conditioning mechanisms.
- Analyze architectural variations in diffusion models, including pixel-space, latent-space, and transformer-based approaches.
- Discuss noise characteristics in real-world imaging domains, and their impact on diffusion-based generative modeling.
- Summarize commonly used datasets and evaluation metrics for diffusion models, covering pixel-level, perceptual, and distribution-based measures.
- Highlight practical trade-offs across models, including computational cost, memory requirements, sampling efficiency, and output fidelity across different resolution regimes.
2. Taxonomy by Architecture Generative Models for Image and Video Generation
- Explicit Density Models;
- Variational/Score-Based Density Models;
- Implicit Models.
2.1. Explicit Density Models
2.1.1. Autoregressive Models
2.1.2. Flow-Based Models
2.2. Variational and Score-Based Density Models
2.2.1. Variational Autoencoders (VAEs)
2.2.2. Diffusion and Score-Based Models
2.3. Implicit Models
- A generator , which maps a latent variable to the data space.
- A discriminator , which attempts to distinguish real samples from generated ones.
2.4. Comparative Summary of Generative Model Paradigms
3. Diffusion Model: Foundations and Taxonomy
3.1. Diffusion Model
3.1.1. Forward Diffusion Process
3.1.2. Reverse Diffusion Process
3.1.3. Training Objective
3.2. Diffusion Model Variants
3.2.1. Denoising Diffusion Probabilistic Models (DDPM)
3.2.2. Latent Diffusion Models
3.2.3. Text-Guided Diffusion
3.2.4. Score-Based Generative Models
3.2.5. Sampling Acceleration Techniques
3.2.6. Large Language Diffusion Models
3.2.7. Transformer-Based Models
Standard DiT Block
DiT Block with Cross-Attention
DiT Block with In-Context Conditioning
3.2.8. Consistency and Distillation-Based Models
3.2.9. Flow Matching and Rectified Flow Models
3.2.10. Hybrid Diffusion Models
3.3. Taxonomy of Diffusion Models
- Representation Space;
- Forward Process Formulation;
- Reverse Parameterization;
- Sampling Strategy;
- Conditioning Mechanism.
3.3.1. Axis 1: Representation Space
- Pixel-Space: Diffusion is applied directly to raw data (e.g., images). This provides full fidelity but is computationally expensive.
- Latent-Space: Data are first encoded into a lower-dimensional representation, and diffusion is performed in this latent space. This reduces computational cost while preserving semantic structure.
- Patch/Token-Space: Data are represented as discrete or continuous tokens (e.g., patches or embeddings), enabling transformer-based diffusion models.
3.3.2. Axis 2: Forward Process Formulation
- Variance Preserving (VP): Maintains a bounded variance during the diffusion process, commonly used in discrete-time diffusion models such as DDPM.
- Variance Exploding (VE): Increases variance over time, often used in score-based generative models formulated via stochastic differential equations.
- EDM Parameterization: A continuous formulation that generalizes noise scaling and improves stability and sampling efficiency.
3.3.3. Axis 3: Reverse Parameterization
- Noise Prediction (ε-prediction): The model predicts the noise added at each step. This is the standard formulation in DDPM.
- Data Prediction (x0-prediction): The model directly predicts the original clean data.
- Velocity Prediction (v-prediction): The model predicts a combination of data and noise, improving stability in some settings.
- Score Function Parameterization: The model predicts the gradient of the log-density (score), connecting diffusion models to score-based generative modeling.
3.3.4. Axis 4: Sampling Strategy
- Stochastic Sampling (DDPM): Uses a Markov chain with injected noise at each step, leading to diverse samples but higher computational cost.
- Deterministic Sampling (DDIM): Uses a non-Markovian deterministic mapping, enabling faster generation with fewer steps.
- ODE-Based Solvers (e.g., DPM-Solver, Heun): Treat sampling as solving an ordinary differential equation, allowing efficient high-quality generation.
- Acceleration/Distillation: Reduces the number of sampling steps through learned or approximated processes.
3.3.5. Axis 5: Conditioning Mechanism
- Unconditional: Generation is based only on the learned data distribution.
- Class-Conditional: Conditioning on discrete labels.
- Text-Conditional (Cross-Attention): Uses text embeddings to guide generation.
- Classifier-Free Guidance (CFG) [52]: Combines conditional and unconditional predictions to control generation strength without requiring a separate classifier.
- Structural Control (e.g., depth, edges, pose): Incorporates spatial constraints or control signals.
3.4. Mapping Diffusion Models to Taxonomy
4. Noise in Imaging and Video Domains
4.1. Domain
4.1.1. Low Light Images
4.1.2. Medical Imaging
4.1.3. Remote Sensing
4.1.4. Infrared Imaging
4.2. Sources of Noise
4.2.1. Sensor Noise
4.2.2. Compression Artifacts
4.2.3. Quantization Noise
4.2.4. Low Light/High ISO Noise
4.2.5. Environmental Effects (Atmospheric Distortion)
4.2.6. Implications for Diffusion Models
4.3. Noise Distributions
4.4. Practical Guidance for Noise Aware Diffusion Model Design
5. Denoising Strategies in Diffusion Models
5.1. In-Model Denoising Techniques in Diffusion Models
- ε: the noise added at each step.
- x0: the original uncorrupted data.
- v: a combination of x0 and ε (velocity formulation), which stabilizes training and improves conditioning.
- DDPM (stochastic sampling);
- DDIM (deterministic, non-Markovian);
- DPM-Solver, EDM (few-step ODE-based solvers).
5.2. External Denoising Techniques
- Pre-process training data (e.g., removing compression artifacts);
- Post-process outputs (e.g., cleaning up residual noise from low-step samplers).
5.2.1. Mean Filter
5.2.2. Median Filter
5.2.3. Non-Local Filtering
5.2.4. Transform Domain Denoising
5.2.5. Bilateral Filtering
5.2.6. Wavelet Transform Denoising
5.2.7. Denoising Autoencoders (DAE)
5.3. Architectural and Denoising Design Choices Across Applications
6. Datasets
- Image generation datasets, such as ImageNet, LSUN, and FFHQ, which are commonly used for unconditional and class-conditional generation;
- Text-to-image datasets, including LAION-5B, CC3M, and CC12M, which enable multimodal conditioning;
- Video datasets, such as Kinetics-700 and Vimeo90K, used for temporal generation and motion modeling;
- 3D and scene datasets, including ShapeNet, SUNCG, and Matterport, supporting spatial and geometric modeling;
- Autonomous driving datasets, such as KITTI and Cityscapes, which provide structured real-world environments;
- Domain-specific datasets, including medical, remote sensing, and infrared imaging datasets, which introduce unique noise characteristics.
Implications of Dataset for Diffusion Models
7. Evaluation Metrics
7.1. Peak Signal-to-Noise Ratio (PSNR)
7.2. Structural Similarity Index (SSIM) and Multi-Scale SSIM (MS-SSIM)
7.3. Perceptual Similarity
7.4. Mean Opinion Score (MOS)
7.5. Fréchet Inception Distance (FID)
7.6. Fréchet Video Distance (FVD)
7.7. Inception Score (IS)
7.8. Kernel Inception Distance (KID)
7.9. Learned Perceptual Image Patch Similarity (LPIPS)
8. Experimental Results and Model Comparison
8.1. Comparative Results
- DDPM (Pixel-space models): DDPMs demonstrate strong performance in low-resolution settings, where direct pixel-space modeling enables accurate reconstruction and high structural fidelity. However, as resolution increases, the computational cost and memory requirements grow significantly, which can limit scalability.
- Stable Diffusion (Latent-space models with conditioning): Stable Diffusion models exhibit strong performance in high-resolution and perceptual tasks, benefiting from latent-space compression and conditioning mechanisms such as text guidance. These models tend to produce visually realistic outputs, although pixel-level fidelity metrics may not fully capture perceptual quality.
- Latent Diffusion Models (LDM): Latent Diffusion models provide a balanced trade-off between computational efficiency and output quality. By operating in a compressed latent space, they reduce resource requirements while maintaining competitive perceptual performance, making them suitable for scalable generation tasks.
8.2. Model Performance and Practical Trade-Offs
- Pixel-space diffusion models tend to achieve higher scores on reconstruction-based metrics (e.g., PSNR, SSIM) in low-resolution settings. However, their computational cost increases significantly with resolution.
- Latent diffusion models perform better on perceptual and distribution-based metrics (e.g., FID, LPIPS), especially at higher resolutions.
- Conditioning mechanisms (e.g., text guidance) improves semantic alignment but may not always improve pixel-level fidelity.
- No single model consistently outperforms others across all metrics, emphasizing the importance of task-specific model selection.
8.3. Implementation Challenges and Practical Considerations
- Latent-space models may require input sizes compatible with the encoder–decoder architecture, particularly for low-resolution datasets.
- High-resolution datasets often require preprocessing steps such as resizing or padding to maintain consistency during training.
- Pixel-space models can be computationally intensive at higher resolutions due to the iterative sampling process, necessitating memory-efficient training strategies.
8.4. Scope and Limitations
9. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- van den Oord, A.; Kalchbrenner, N.; Kavukcuoglu, K. Pixel recurrent neural networks. arXiv 2016, arXiv:1601.06759. [Google Scholar]
- Salimans, T.; Karpathy, A.; Chen, X.; Kingma, D.P. PixelCNN++: Improving the PixelCNN with discretized logistic mixture likelihood and other modifications. arXiv 2017, arXiv:1701.05517. [Google Scholar]
- Dinh, L.; Krueger, D.; Bengio, Y. NICE: Non-linear independent components estimation. arXiv 2014, arXiv:1410.8516. [Google Scholar]
- Dinh, L.; Sohl-Dickstein, J.; Bengio, S. Density estimation using Real NVP. arXiv 2016, arXiv:1605.08803. [Google Scholar]
- Kingma, D.P.; Dhariwal, P. Glow: Generative flow with invertible 1x1 convolutions. Adv. Neural Inf. Process. Syst. 2018, 31, 10215–10224. [Google Scholar]
- Kingma, D.P.; Welling, M. Auto-encoding variational Bayes. In Proceedings of the 2nd International Conference on Learning Representations (ICLR), Banff, AB, Canada, 14–16 April 2014. [Google Scholar]
- Doersch, C. Tutorial on variational autoencoders. arXiv 2016, arXiv:1606.05908. [Google Scholar]
- Sohn, K.; Lee, H.; Yan, X. Learning structured output representation using deep conditional generative models. Adv. Neural Inf. Process. Syst. 2015, 28, 3483–3491. [Google Scholar]
- Kipf, T.N.; Welling, M. Variational graph autoencoders. arXiv 2016, arXiv:1611.07308. [Google Scholar]
- Zhang, C.; Barbano, R.; Jin, B. Conditional variational autoencoder for learned image reconstruction. Computation 2021, 9, 114. [Google Scholar] [CrossRef]
- Oord, A.V.; Vinyals, O.; Kavukcuoglu, K. Neural Discrete Representation Learning. Neural Inf. Process. Syst. 2017, 30, 6309–6318. [Google Scholar]
- Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; Bengio, Y. Generative adversarial nets. Adv. Neural Inf. Process. Syst. 2014, 27, 2672–2680. [Google Scholar]
- Karras, T.; Laine, S.; Aila, T. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019; pp. 4401–4410. [Google Scholar]
- Zhu, J.Y.; Park, T.; Isola, P.; Efros, A.A. Unpaired image-to-image translation using cycle-consistent adversarial networks. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 27–29 October 2017; pp. 2223–2232. [Google Scholar]
- Zhu, J.; Park, T.; Isola, P.; Efros, A.A. BicycleGAN: Toward Realistic Image Decomposition and Translation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Venice, Italy, 24–27 October 2017. [Google Scholar]
- Brock, A.; Donahue, J.; Simonyan, K. Large Scale GAN Training for High Fidelity Natural Image Synthesis. In Proceedings of the International Conference on Learning Representations (ICLR), Vancouver, BC, Canada, 30 April–3 May 2018. [Google Scholar]
- Radford, A.; Metz, L.; Chintala, S. Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks. In Proceedings of the International Conference on Learning Representations (ICLR), San Juan, Puerto Rico, 2–4 May 2016. [Google Scholar]
- Park, T.; Liu, M.Y.; Wang, T.C.; Zhu, J.Y. GauGAN: Semantic Image Synthesis with Spatially-Adaptive Normalization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019; pp. 2337–2346. [Google Scholar]
- Chen, X.; Duan, Y.; Houthooft, R.; Schulman, J.; Sutskever, I.; Abbeel, P. InfoGAN: Interpretable Representation Learning by Information Maximizing Generative Adversarial Nets. In Proceedings of the 30th International Conference on Neural Information Processing Systems (NIPS), Barcelona, Spain, 5–10 December 2016. [Google Scholar]
- Li, Y.; Huang, X.; Hu, Z.; Zhang, L.; Liu, L. LayoutGAN: Generating Graphic Layouts with Layout Variational Autoencoders. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019. [Google Scholar]
- Zhang, L.; Zhang, X.; Wang, H.; Sun, Y. MaskGAN: Towards High-Resolution Image Manipulation with Deep Generative Models. IEEE Trans. Image Process. 2020, 29, 4387–4400. [Google Scholar]
- Liu, X.; Xu, L.; Zhang, H.; Jin, X. Object-Centric Generative Adversarial Networks. In Proceedings of the 33rd AAAI Conference on Artificial Intelligence (AAAI), Honolulu, HI, USA, 27 January–1 February 2019. [Google Scholar]
- Zhang, H.; Goodfellow, I.; Metaxas, D.; Odena, A. Self-Attention Generative Adversarial Networks. In Proceedings of the 36th International Conference on Machine Learning (ICML), Long Beach, CA, USA, 9–15 June 2019. [Google Scholar]
- Miyato, T.; Kataoka, T.; Koyama, M.; Yoshida, Y. Spectral Normalization for Generative Adversarial Networks. In Proceedings of the International Conference on Learning Representations (ICLR), Vancouver, BC, Canada, 30 April–3 May 2018. [Google Scholar]
- Ledig, C.; Theis, L.; Huszár, F.; Caballero, J.; Cunningham, A.; Acosta, A.; Aitken, A.; Tejani, A.; Totz, J.; Wang, Z.; et al. Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 4681–4690. [Google Scholar]
- Esser, P.; Rombach, R.; Ommer, B. Taming Transformers for High-Resolution Image Synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 12873–12882. [Google Scholar]
- Xiao, Z.; Kreis, K.; Vahdat, A. Tackling the Generative Learning Trilemma with Denoising Diffusion GANs. In Proceedings of the International Conference on Learning Representations, Virtual, 25–29 April 2022. [Google Scholar]
- Ho, J.; Jain, A.; Abbeel, P. Denoising diffusion probabilistic models. arXiv 2020, arXiv:2006.11239. [Google Scholar]
- Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; Ommer, B. High-Resolution Image Synthesis with Latent Diffusion Models. arXiv 2022, arXiv:2112.10752. [Google Scholar]
- Sohl-Dickstein, J.; Weiss, E.; Maheswaranathan, N.; Ganguli, S. Deep unsupervised learning using nonequilibrium thermodynamics. In Proceedings of the 32nd International Conference on Machine Learning (ICML), Lille, France, 6–11 July 2015; pp. 2256–2265. [Google Scholar]
- Nichol, A.; Dhariwal, P.; Ramesh, A.; Shyam, P.; Mishkin, P.; McGrew, B.; Sutskever, I.; Chen, M. GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models. arXiv 2022, arXiv:2112.10741. [Google Scholar]
- Kim, G.; Kwon, T.; Ye, J.C. DiffusionCLIP: Text-Guided Diffusion Models for Robust Image Manipulation. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 2416–2425. [Google Scholar] [CrossRef]
- Yang, S.; Hwang, H.; Ye, J.C. Zero-Shot Contrastive Loss for Text-Guided Diffusion Image Style Transfer. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 1–6 October 2023; pp. 22816–22825. [Google Scholar] [CrossRef]
- Lim, S.; Yoon, E.; Byun, T.; Kang, T.; Kim, S.; Lee, K.; Choi, S. Score-based generative modeling through stochastic evolution equations in Hilbert spaces. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NeurIPS ’23); Curran Associates Inc.: Red Hook, NY, USA, 2023; Volume 1645, pp. 37799–37812. [Google Scholar]
- Song, Y.; Sohl Dickstein, J.; Kingma, D.P.; Kumar, A.; Ermon, S.; Poole, B. Score Based Generative Modeling through Stochastic Differential Equations. In Proceedings of the International Conference on Learning Representations (ICLR), Virtual, 3–7 May 2021. [Google Scholar]
- Song, J.; Meng, C.; Ermon, S. Denoising Diffusion Implicit Models. arXiv 2022, arXiv:2010.02502. [Google Scholar]
- Lu, C.; Zhou, Y.; Bao, F.; Chen, J.; Li, C.; Zhu, J. DPM-Solver: A fast ODE solver for diffusion probabilistic model sampling in around 10 steps. In Proceedings of the 36th International Conference on Neural Information Processing Systems (NeurIPS ’22); Curran Associates Inc.: Red Hook, NY, USA, 2022; pp. 5775–5787. [Google Scholar]
- Nie, S.; Zhu, F.; You, Z.; Zhang, X.; Ou, J.; Hu, J.; Zhou, J.; Lin, Y.; Wen, J.-R.; Li, C. Large Language Diffusion Models. arXiv 2025, arXiv:2502.09992. [Google Scholar]
- Ramesh, A.; Pavlov, M.; Goh, G.; Gray, S.; Voss, C.; Radford, A.; Chen, M.; Sutskever, I. Zero-shot text-to-image generation. In Proceedings of the 38th International Conference on Machine Learning (ICML), Virtual, 18–24 July 2021; pp. 8821–8831. [Google Scholar]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, L.; Polosukhin, I. Attention is all you need. Adv. Neural Inf. Process. Syst. 2017, 30, 5998–6008. [Google Scholar]
- Brown, T.B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. Language models are few-shot learners. Adv. Neural Inf. Process. Syst. 2020, 33, 1877–1901. [Google Scholar]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An Image is Worth 16 × 16 Words: Transformers for Image Recognition at Scale. arXiv 2020, arXiv:2010.11929. [Google Scholar]
- Peebles, W.; Xie, S. Scalable Diffusion Models with Transformers. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 1–6 October 2023; pp. 4195–4205. [Google Scholar]
- Luo, S.; Tan, Y.; Huang, L.; Li, J.; Zhao, H. Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference. arXiv 2023, arXiv:2310.04378. [Google Scholar]
- Kitaev, N.; Kaiser, L.; Levskaya, A. Reformer: The Efficient Transformer. In Proceedings of the International Conference on Learning Representations (ICLR), Virtual, 26 April–1 May 2020. [Google Scholar]
- Dao, T.; Fu, D.; Ermon, S.; Rudra, A.; Re, C. FlashAttention: Fast and Memory Efficient Exact Attention with IO Awareness. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), New Orleans, LA, USA, 28 November–9 December 2022; pp. 16344–16359. [Google Scholar]
- Song, Y.; Dhariwal, P.; Chen, M.; Sutskever, I. Consistency Models. In Proceedings of the 40th International Conference on Machine Learning (ICML), Honolulu, HI, USA, 23–29 July 2023; Volume 202, pp. 32211–32252. [Google Scholar]
- Luo, S.; Tan, Y.; Patil, S.; Gu, D.; von Platen, P.; Passos, A.; Huang, L.; Li, J.; Zhao, H. LCM-LoRA: A Universal Stable-Diffusion Acceleration Module. arXiv 2023, arXiv:2311.05556. [Google Scholar]
- Salimans, T.; Ho, J. Progressive Distillation for Fast Sampling of Diffusion Models. arXiv 2022, arXiv:2202.00512. [Google Scholar]
- Lipman, Y.; Chen, R.T.; Ben-Hamu, H.; Nickel, M.; Le, M. Flow Matching for Generative Modeling. In Proceedings of the 11th International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, 1–5 May 2023. [Google Scholar]
- Liu, Q. Rectified flow: A marginal preserving approach to optimal transport. arXiv 2022, arXiv:2209.14577. [Google Scholar]
- Ho, J.; Salimans, T. Classifier Free Diffusion Guidance. In Proceedings of the NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications, Virtual, 14 December 2021. [Google Scholar]
- Jo, Y.; Chun, S.Y.; Choi, J. Rethinking Deep Image Prior for Denoising. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 10–17 October 2021; pp. 5067–5076. [Google Scholar]
- Su, H.; Yu, L.; Jung, C. Joint Contrast Enhancement and Noise Reduction of Low Light Images via JND Transform. IEEE Trans. Multimed. 2022, 24, 17–32. [Google Scholar] [CrossRef]
- Yu, H.; Wang, G. Compressed Sensing Based Interior Tomography. Phys. Med. Biol. 2009, 54, 2791–2805. [Google Scholar] [CrossRef]
- Jensen, J.A.; Svendsen, N.B. Calculation of Pressure Fields from Arbitrarily Shaped, Apodized, and Excited Ultrasound Transducers. IEEE Trans. Ultrason. Ferroelectr. Freq. Control. 1992, 39, 262–267. [Google Scholar] [CrossRef] [PubMed]
- Yan, F.; Wu, S.; Zhang, Q.; Liu, Y.; Sun, H. Destriping of Remote Sensing Images by an Optimized Variational Model. Sensors 2023, 23, 7529. [Google Scholar] [CrossRef]
- Geng, J.; Jiang, W.; Deng, X. Multi-Scale Deep Feature Learning Network with Bilateral Filtering for SAR Image Classification. ISPRS J. Photogramm. Remote Sens. 2020, 167, 201–213. [Google Scholar] [CrossRef]
- Cui, G.M.; Feng, H.J.; Xu, Z.H.; Li, Q.; Chen, Y.T. Multi-Scale Detail-Preserving Denoising Method of Infrared Image via Relative Total Variation. In International Symposium on Photoelectronic Detection and Imaging 2013: Infrared Imaging and Applications; SPIE: Bellingham, WA, USA, 2013; Volume 8907, pp. 268–274. [Google Scholar]
- Wang, E.; Jiang, P.; Li, X.; Cao, H. Infrared Stripe Correction Algorithm Based on Wavelet Decomposition and Total Variation-Guided Filtering. J. Eur. Opt. Soc.-Rapid Publ. 2020, 16, 1. [Google Scholar] [CrossRef]
- Ni, C.; Li, Q.; Xia, L.Z. A Novel Method of Infrared Image Denoising and Edge Enhancement. Signal Process. 2008, 88, 1606–1614. [Google Scholar] [CrossRef]
- Dabov, K.; Foi, A.; Katkovnik, V.; Egiazarian, K. Image Denoising by Sparse 3D Transform Domain Collaborative Filtering. IEEE Trans. Image Process. 2007, 16, 2080–2095. [Google Scholar] [CrossRef]
- Meng, Q.; Li, L.; Nießner, M.; Dai, A. LT3SD: Latent Trees for 3D Scene Diffusion. arXiv 2024, arXiv:2409.08215. [Google Scholar]
- Yang, B.; Luo, Y.; Chen, Z.; Wang, G.; Liang, X.; Lin, L. LAW-Diffusion: Complex Scene Generation by Diffusion with Layouts. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 1–6 October 2023; pp. 22612–22622. [Google Scholar]
- Zhang, R. Perceptual Similarity Guidance and Text Guidance Optimization for Editing Real Images using Guided Diffusion Models. arXiv 2023, arXiv:2312.06680. [Google Scholar]
- Sui, J.; Ma, X.; Zhang, X.; Pun, M.-O.; Wu, H. Adaptive Semantic-Enhanced Denoising Diffusion Probabilistic Model for Remote Sensing Image Super-Resolution. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2025, 18, 892–906. [Google Scholar] [CrossRef]
- Sui, J.; Wu, Q.; Pun, M.-O. Denoising Diffusion Probabilistic Model with Adversarial Learning for Remote Sensing Super-Resolution. Remote Sens. 2024, 16, 1219. [Google Scholar] [CrossRef]
- Tan, W.; Liu, B.; Zhang, J.; Song, R.; Fu, J. RoLD: Robot Latent Diffusion for Multi-task Policy Modeling. arXiv 2024, arXiv:2403.07312. [Google Scholar]
- Shaoul, Y.; Mishani, I.; Vats, S.; Li, J.; Likhachev, M. Multi-Robot Motion Planning with Diffusion Models. arXiv 2024, arXiv:2410.03072. [Google Scholar]
- Fung, A.; Benhabib, B.; Nejat, G. LDTrack: Dynamic People Tracking by Service Robots using Diffusion Models. arXiv 2024, arXiv:2402.08774. [Google Scholar]
- Wang, Z.; Hao, Z.; Lin, J.; Feng, Y.; Guo, Y. UP-Diff: Latent Diffusion Model for Remote Sensing Urban Prediction. arXiv 2024, arXiv:2407.11578. [Google Scholar]
- Tang, K.; Chen, J. ChangeAnywhere: Sample Generation for Remote Sensing Change Detection via Semantic Latent Diffusion Model. arXiv 2024, arXiv:2404.08892. [Google Scholar]
- Zhou, J.-Y.; Fu, T.-H. Word card generation for language education using latent diffusion model. IET Conf. Proc. 2024, 2023, 146–147. [Google Scholar] [CrossRef]
- Arbelaez, P.; Maire, M.; Fowlkes, C.; Malik, J. Contour Detection and Hierarchical Image Segmentation. IEEE Trans. Pattern Anal. Mach. Intell. 2011, 33, 898–916. [Google Scholar]
- Krizhevsky, A.; Hinton, G. Learning Multiple Layers of Features from Tiny Images; Technical Report; University of Toronto: Toronto, ON, Canada, 2009. [Google Scholar]
- Deng, J.; Dong, W.; Socher, R.; Li, L.J.; Li, K.; Fei-Fei, L. ImageNet: A Large-Scale Hierarchical Image Database. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Miami, FL, USA, 20–25 June 2009. [Google Scholar]
- Yu, F.; Zhang, Y.; Song, S.; Funkhouser, T.; Xiao, J. LSUN: Construction of a Large-scale Image Dataset using Deep Learning with Humans in the Loop. arXiv 2015, arXiv:1506.03365. [Google Scholar]
- Zhou, B.; Lapedriza, A.; Khosla, A.; Oliva, A.; Torralba, A. Places: A 10 Million Image Database for Scene Recognition. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 40, 1452–1464. [Google Scholar] [CrossRef] [PubMed]
- Lee, C.-H.; Liu, Z.; Wu, L.; Luo, P. MaskGAN: Towards Diverse and Interactive Facial Image Manipulation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020. [Google Scholar]
- Agustsson, E.; Timofte, R. NTIRE 2017 Challenge on Single Image Super-Resolution: Dataset and Study. In CVPR Workshops; IEEE: New York, NY, USA, 2017; pp. 1122–1131. [Google Scholar]
- Zhou, B.; Zhao, H.; Puig, X.; Fidler, S.; Barriuso, A.; Torralba, A. Scene Parsing through ADE20K Dataset. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017. [Google Scholar]
- Caesar, H.; Uijlings, J.; Ferrari, V. COCO-Stuff: Thing and Stuff Classes in Context. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–23 June 2018. [Google Scholar]
- Chao, Y.-W.; Liu, Y.; Liu, X.; Zeng, H.; Deng, J. Learning to Detect Human-Object Interactions. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision (WACV), Lake Tahoe, NV, USA, 12–15 March 2018. [Google Scholar]
- Lin, T.-Y.; Maire, M.; Belongie, S.; Bourdev, L.; Girshick, R.; Hays, J.; Perona, P.; Ramanan, D.; Zitnick, C.L.; Dollár, P. Microsoft COCO: Common Objects in Context. In Proceedings of the European Conference on Computer Vision (ECCV), Zurich, Switzerland, 6–12 September 2014. [Google Scholar]
- Kuznetsova, A.; Rom, H.; Alldrin, N.; Uijlings, J.; Krasin, I.; Pont-Tuset, J.; Kamali, S.; Popov, S.; Malloci, M.; Kolesnikov, A.; et al. The Open Images Dataset V4. Int. J. Comput. Vis. 2020, 128, 1956–1981. [Google Scholar] [CrossRef]
- Everingham, M.; Van Gool, L.; Williams, C.K.I.; Winn, J.; Zisserman, A. The PASCAL Visual Object Classes (VOC) Challenge. Int. J. Comput. Vis. 2010, 88, 303–338. [Google Scholar] [CrossRef]
- Sharma, P.; Ding, N.; Goodman, S.; Soricut, R. Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset for Automatic Image Captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, Melbourne, Australia, 15–20 July 2018. [Google Scholar]
- Changpinyo, S.; Sharma, P.; Ding, N.; Soricut, R. Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training to Recognize Long-Tail Visual Concepts. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021. [Google Scholar]
- Young, P.; Lai, A.; Hodosh, M.; Hockenmaier, J. From Image Descriptions to Visual Denotations: New Similarity Metrics for Semantic Inference over Event Descriptions. Trans. Assoc. Comput. Linguist. 2017, 2, 67–78. [Google Scholar] [CrossRef]
- Schuhmann, C.; Beaumont, R.; Vencu, R.; Gordon, C.; Wightman, R.; Cherti, M.; Coombes, T.; Katta, A.; Mullis, C.; Wortsman, M.; et al. LAION-5B: An Open Large-Scale Dataset for Training Next Generation Image-Text Models. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), New Orleans, LA, USA, 28 November–9 December 2022; Volume 35, pp. 25278–25294. [Google Scholar]
- Chua, T.-S.; Tang, J.; Hong, R.; Li, H.; Luo, Z.; Zheng, Y. NUS-WIDE: A Real-World Web Image Database from National University of Singapore. In Proceedings of the ACM International Conference on Image and Video Retrieval, Santorini Island, Greece, 8–10 July 2009. [Google Scholar]
- Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.J.; Shamma, D.A.; et al. Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations. Int. J. Comput. Vis. 2017, 123, 32–73. [Google Scholar] [CrossRef]
- Sigurdsson, G.A.; Varol, G.; Wang, X.; Farhadi, A.; Laptev, I.; Gupta, A. Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding. In Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands, 11–14 October 2016; pp. 510–526. [Google Scholar]
- Pont-Tuset, J.; Perazzi, F.; Caelles, S.; Arbeláez, P.; Sorkine-Hornung, A.; Van Gool, L. The 2017 DAVIS Challenge on Video Object Segmentation. arXiv 2017, arXiv:1704.00675. [Google Scholar]
- Kuehne, H.; Jhuang, H.; Garrote, E.; Poggio, T.; Serre, T. HMDB: A Large Video Database for Human Motion Recognition. In Proceedings of the International Conference on Computer Vision (ICCV), Barcelona, Spain, 6–13 November 2011. [Google Scholar]
- Carreira, J.; Zisserman, A. Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017. [Google Scholar]
- Karpathy, A.; Toderici, G.; Shetty, S.; Leung, T.; Sukthankar, R.; Fei-Fei, L. Large-Scale Video Classification with Convolutional Neural Networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Columbus, OH, USA, 23–28 June 2014. [Google Scholar]
- Wang, X.; Wu, J.; Chen, J.; Li, L.; Wang, Y.-F.; Wang, W.Y. VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 4581–4591. [Google Scholar]
- Xue, T.; Chen, B.; Wu, J.; Wei, D.; Freeman, W.T. Video Enhancement with Task-Oriented Flow. Int. J. Comput. Vis. 2019, 127, 1106–1125. [Google Scholar] [CrossRef]
- Abu-El-Haija, S.; Kothari, N.; Lee, J.; Natsev, P.; Toderici, G.; Varadarajan, B.; Vijayanarasimhan, S. YouTube-8M: A Large-Scale Video Classification Benchmark. arXiv 2016, arXiv:1609.08675. [Google Scholar]
- Fu, H.; Cai, B.; Gao, L.; Zhang, L.; Wang, J.; Li, C.; Xun, Z.; Sun, C.; Jia, R.; Zhao, B.; et al. 3D-FRONT: 3D Furnished Rooms with Layout and semaNTics. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 11–17 October 2021. [Google Scholar]
- Hu, Z.; Iscen, A.; Jain, A.; Kipf, T.; Yue, Y.; Ross, D.A.; Schmid, C.; Fathi, A. SceneCraft: An LLM Agent for Synthesizing 3D Scenes as Blender Code. In Proceedings of the 32nd ACM International Conference on Multimedia (MM), Melbourne, Australia, 28 October–1 November 2024. [Google Scholar]
- Chang, A.; Dai, A.; Funkhouser, T.; Halber, M.; Nießner, M.; Savva, M.; Song, S.; Zeng, A.; Zhang, Y. Matterport3D: Learning from RGB-D Data in Indoor Environments. In Proceedings of the International Conference on 3D Vision (3DV), Qingdao, China, 10–12 October 2017. [Google Scholar]
- Wu, Z.; Song, S.; Khosla, A.; Yu, F.; Zhang, L.; Tang, X.; Xiao, J. 3D ShapeNets: A Deep Representation for Volumetric Shapes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA, 7–12 June 2015. [Google Scholar]
- Chen, W.; Qian, S.; Fan, D.; Kojima, N.; Hamilton, M.; Deng, J. OASIS: A Large-Scale Dataset for Single Image 3D in the Wild. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 679–688. [Google Scholar]
- Wu, W.; Fu, X.M.; Tang, R.; Wang, Y.; Qi, Y.H.; Liu, L. Data-Driven Interior Plan Generation for Residential Buildings. ACM Trans. Graph. 2019, 38, 1–12. [Google Scholar] [CrossRef]
- Dai, A.; Chang, A.; Savva, M.; Halber, M.; Funkhouser, T.; Nießner, M. ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 June 2017. [Google Scholar]
- Chang, A.X.; Funkhouser, T.; Guibas, L.; Hanrahan, P.; Huang, Q.; Li, Z.; Savarese, S.; Savva, M.; Song, S.; Su, H.; et al. ShapeNet: An Information-Rich 3D Model Repository. arXiv 2015, arXiv:1512.03012. [Google Scholar]
- Song, S.; Yu, F.; Zeng, A.; Chang, A.X.; Savva, M.; Funkhouser, T. Semantic Scene Completion from a Single Depth Image. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 June 2017. [Google Scholar]
- Chen, W.; Qian, S.; Deng, J. Learning Single-Image Depth from Videos Using Quality Assessment Networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019; pp. 5604–5613. [Google Scholar]
- Cordts, M.; Omran, M.; Ramos, S.; Rehfeld, T.; Enzweiler, M.; Benenson, R.; Franke, U.; Roth, S.; Schiele, B. The Cityscapes Dataset for Semantic Urban Scene Understanding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016. [Google Scholar]
- Varma, G.; Subramanian, A.; Namboodiri, A.; Chandraker, M.; Jawahar, C.V. IDD: A Dataset for Exploring Problems of Autonomous Navigation in Unconstrained Environments. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019. [Google Scholar]
- Geiger, A.; Lenz, P.; Urtasun, R. Are We Ready for Autonomous Driving? The KITTI Vision Benchmark Suite. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Providence, RI, USA, 16–21 June 2012. [Google Scholar]
- Wrenninge, M.; Unger, M. Synscapes: A Photorealistic Synthetic Dataset for Street Scene Parsing. arXiv 2018, arXiv:1810.08705. [Google Scholar]
- Calli, B.; Walsman, A.; Singh, A.; Srinivasa, S.; Abbeel, P.; Dollar, A.M. Benchmarking in Manipulation Research: The YCB Object and Model Set and Benchmarking Protocols. In Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Hamburg, Germany, 28 September–2 October 2015. [Google Scholar]
- Vasudevan, S.; Bleyer, M.; Geiger, A. DIODE: A Dense Indoor and Outdoor Depth Dataset. arXiv 2019, arXiv:1908.01940. [Google Scholar]
- Li, Z.; Snavely, N. MegaDepth: Learning Single-View Depth Prediction from Internet Photos. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–23 June 2018. [Google Scholar]
- Johnson, J.; Hariharan, B.; van der Maaten, L.; Fei-Fei, L. CLEVR-G: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017. [Google Scholar]
- LeCun, Y.; Cortes, C.; Burges, C. The MNIST Database of Handwritten Digits. 1998. Available online: http://yann.lecun.com/exdb/mnist (accessed on 15 March 2026).
- Liu, Z.; Luo, P.; Wang, X.; Tang, X. DeepFashion: Powering Robust Clothes Recognition and Retrieval. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016. [Google Scholar]
- Xiao, H.; Rasul, K.; Vollgraf, R. Fashion-MNIST: A Novel Image Dataset for Benchmarking Machine Learning Algorithms. arXiv 2017, arXiv:1708.07747. [Google Scholar]
- Wang, Y.; Ostermann, J.; Zhang, Y. Video Processing and Communications; Prentice Hall: Upper Saddle River, NJ, USA, 2001. [Google Scholar]
- Wang, Z.; Bovik, A.C.; Sheikh, H.R.; Simoncelli, E.P. Image quality assessment: From error visibility to structural similarity. IEEE Trans. Image Process. 2004, 13, 600–612. [Google Scholar] [CrossRef] [PubMed]
- Wang, Z.; Simoncelli, E.P.; Bovik, A.C. Multiscale structural similarity for image quality assessment. Proc. IEEE Int. Conf. Image Process. (ICIP) 2003, 2, 1398–1402. [Google Scholar]
- Zhang, R.; Isola, P.; Efros, A.A.; Shechtman, E.; Wang, O. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–23 June 2018; pp. 586–595. [Google Scholar]
- ITU-T. Recommendation P.910: Subjective Video Quality Assessment Methods for Multimedia Applications; ITU: Geneva, Switzerland, 2008. [Google Scholar]
- Heusel, M.; Ramsauer, H.; Unterthiner, T.; Nessler, B.; Hochreiter, S. GANs trained by a two time-scale update rule converge to a local Nash equilibrium. Adv. Neural Inf. Process. Syst. 2017, 30, 6626–6637. [Google Scholar]
- Unterthiner, T.; van Steenkiste, S.; Kurach, K.; Marinier, R.; Michalski, M.; Gelly, S. Towards accurate generative models of video. arXiv 2018, arXiv:1812.01717. [Google Scholar]
- Salimans, T.; Goodfellow, I.; Zaremba, W.; Cheung, V.; Radford, A.; Chen, X. Improved techniques for training GANs. Adv. Neural Inf. Process. Syst. 2016, 29, 2234–2242. [Google Scholar]
- Bińkowski, M.; Sutherland, D.J.; Arbel, M.; Gretton, A. Demystifying MMD GANs. arXiv 2021, arXiv:1801.01401. [Google Scholar]
- Dhariwal, P.; Nichol, A. Diffusion Models Beat GANs on Image Synthesis. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Virtual, 6 December–14 December 2021; pp. 8780–8794. [Google Scholar]













| Variant | Conditioning/Structure | Latent Representation | Architectural Modification | Representative Use Cases | Key Advantage |
|---|---|---|---|---|---|
| CVAE [10] | Uses auxiliary conditions (e.g., class labels, side information) in both encoder and decoder | Continuous latent distribution conditioned on inputs | Encoder and decoder take conditional inputs; optimize conditional ELBO | Class-conditional generation, controlled synthesis | Enables guided generation using auxiliary information |
| VGAE [9] | Designed for graph-structured data; node features are encoded | Continuous latent embeddings per graph node | GCN layers in encoder; inner-product decoder | Link prediction, graph embedding | Captures relational and topological structure |
| VQ-VAE [11] | Unconditional or conditional via codebook guidance | Discrete codebook of latent embeddings | Encoder outputs nearest codebook vector; commitment losses | High-fidelity image/audio generation, tokenized representations | Discrete latent codes improve representation and alleviate posterior collapse. |
| GAN Variant | Core Modification | Primary Application | Key Features |
|---|---|---|---|
| BicycleGAN | Adds a bijection between output and latent space to prevent models from generating few repetitive samples. | Diverse image-to-image translation | Maps output back to latent space; avoids repetitive sample generation [15] |
| BigGAN | Expands model size with increased batch size and layer width, adding skip connections and orthogonal regularization. | High-fidelity class-conditional synthesis | Large model size; improved performance with skip connection; orthogonal regularization [16] |
| CycleGAN | Works in a cyclic manner, mapping input to output and vice versa, supporting paired and unpaired images. | Unpaired image-to-image translation | Cycle consistency loss; mapping between paired/unpaired images; mean squared error calculation [14] |
| DCGANs | Uses convolutional layers to generate realistic images, effective in unsupervised learning. | General image synthesis | Convolutional layers in both generator and discriminator; hierarchical image representation learning [17] |
| FRVSR-GAN (Frame Recurrent VSR-GAN) | Designed for video super-resolution, generating high-quality frames sequentially using a recurrent architecture. | Video super-resolution | Frame-to-frame dependency; recurrent neural network architecture for video frames. |
| GauGAN | Developed by NVIDIA, generates images based on text, semantic segmentation, sketches, and style. | Semantic image synthesis | Multi-input capability; supports interactive sketch-based generation [18] |
| InfoGAN | Maximizes mutual information to learn interpretable features, separating style from content (e.g., writing style from digit shape). | Disentangled representation learning | Mutual information maximization; interpretable latent variables [19] |
| LayoutGAN | Generates object layouts with multiple self-attention layers on 2D planes to create semantic relationships. | Graphic and layout generation | Multi-layer attention for layout generation; rendering layer for layout [20] |
| MaskGAN | Maps semantic segmentation to target images, allowing interactive manipulation of segmentation maps to produce new images. | Semantic editing and manipulation | Dense Mapping Network and Editing Behavior Simulator; interactive manipulation of segmentation maps [21] |
| OC-GAN | Addresses complex scenes with multiple objects, avoiding spurious or overlapping object generation using bounding boxes. | Multi-object scene synthesis | Object-centric approach; bounding box generation and management [22] |
| SAGAN (Self-Attention GAN) | Uses self-attention mechanisms to capture long-range dependencies, allowing fine-grained details in high-resolution images. | High-resolution synthesis | Self-attention layers; long-range dependency modeling for detailed image consistency [23] |
| SN-GAN (Spectral Normalization GAN) | Resolves gradient vanishing/exploding and model collapse issues using spectral normalization. | Stable adversarial training | Spectral normalization for stability; improved training process [24] |
| SRGAN | Enhance low-resolution images to high-resolution ones by leveraging adversarial training and perceptual loss, producing photo-realistic details | Super-resolution | Perceptual loss, realistic outputs, texture preservation, adversarial training, and versatile applications [25] |
| StyleGAN | Employs an alternative generator architecture inspired by style transfer, with progressive training and mixing regularization to improve image generation quality. | High-fidelity image synthesis | Progressive growing, style vectors from latent variables, mixing regularization [13] |
| VQ-GAN (Vector Quantised GAN) | Combines Vector Quantized approach with GAN to produce higher-quality images with reduced noise. | High-resolution synthesis | Vector quantizer with discrete codebook vectors; noise reduction and object constraints [26] |
| Criterion | VAE | GAN | Flow | Diffusion |
|---|---|---|---|---|
| Likelihood | Variational lower bound (ELBO) | Not applicable (implicit) | Exact (change in variables) | Variational bound or score matching |
| Training Stability | Stable; straightforward optimization | Prone to mode collapse and instability | Stable; bijective constraint | Stable; simple denoising objective |
| Sample Quality | Moderate; often blurry | High; sharp and realistic | Moderate to high | High; comparable or superior to GANs |
| Sampling Speed | Fast (single forward pass) | Fast (single forward pass) | Fast (single forward pass) | Slow (iterative; 10 to 1000 steps) |
| Mode Coverage | Good but may underfit | Prone to mode dropping | Full (bijective) | Excellent coverage |
| Controllability | Latent interpolation | Limited without auxiliary losses | Invertible; supports editing | Strong (CFG, ControlNet) |
| Scalability | Moderate | High with progressive training | Limited by invertibility | High with latent space formulations |
| Typical Applications | Representation learning, anomaly detection | Image synthesis, style transfer | Density estimation, exact inference | Image/video synthesis, inpainting, conditional generation |
| Model | Representation Space | Forward | Reverse Parameterization | Sampling Strategy | Conditioning Mechanism |
|---|---|---|---|---|---|
| DDPM | Pixel | VP | Noise | Stochastic | None |
| DDIM | Pixel | VP | Noise | Deterministic | None |
| Score-SDE | Pixel | VE | Score | ODE/SDE | Optional |
| LDM | Latent | VP | Noise | DDIM/ODE | Text |
| Stable Diffusion | Latent | VP | Noise | DDIM + CFG | Text |
| DiT | Latent/Patch | VP | Noise/Velocity | ODE | Text |
| ControlNet | Latent | VP | Noise | DDIM | Structural |
| Consistency Model | Pixel/Latent | VP/VE | Score/Data | Single-step/Few-step ODE | Optional/CFG |
| Progressive Distillation | Pixel/Latent | VP | Noise | Few-step Distilled Sampler | Same as base model |
| Noise Type | Domain | Distribution | Recommended Forward Process | Preprocessing Strategy |
|---|---|---|---|---|
| Photon shot noise | Low light photography | Poisson | Variance Exploding (VE) SDE | Variance stabilizing transform (e.g., Anscombe) |
| Speckle noise | Medical ultrasound | Rayleigh/multiplicative | VE SDE with adapted schedule | Log transform; bilateral filtering for smoothing |
| Gaussian sensor noise | General photography | Gaussian | Variance Preserving (VP) SDE (standard DDPM) | Standard normalization |
| Compression artifacts | JPEG images, web images | Structured/blocky | VP SDE | DCT domain denoising as preprocessing |
| Thermal/read noise | High ISO, infrared | Gaussian mixture | VP or EDM parameterization | BM3D or wavelet denoising as preprocessing |
| Atmospheric distortion | Remote sensing, satellite | Non Gaussian (complex) | VE SDE or EDM | Wavefront correction; variational filtering |
| Technique | Key Idea | Advantages | Limitations | Typical Use Cases (with Diffusion Models) |
|---|---|---|---|---|
| Mean Filter | Replaces each pixel with the average value of its local neighborhood. |
|
| Rarely used; coarse smoothing in early preprocessing. |
| Median Filter | Replaces each pixel with the median value of its local neighborhood. |
|
| Preprocessing for corrupted training data with impulse noise. |
| Non-Local Filtering (BM3D) | Groups similar blocks (patches) into 3D arrays and applies collaborative filtering. |
|
| Optional post-processing for fine-detail preservation in high-fidelity outputs. |
| Transform Domain Denoising (e.g., DCT, Fourier) | Transforms the image to a frequency domain, then thresholds or smooths high-frequency coefficients. |
|
| Removing banding or ringing artifacts in post-processing. |
| Bilateral Filtering | Uses spatial and intensity differences to weight neighboring pixels, preserving edges while smoothing. |
|
| Post-processing when edge fidelity is visually important. |
| Wavelet Transform Denoising | Decomposes image into wavelet coefficients; thresholds high-frequency coefficients to reduce noise. |
|
| High-resolution domains (e.g., medical imaging, remote sensing) after generation. |
| Denoising Autoencoders (DAE) | Neural networks trained to map noisy inputs to clean outputs, learning complex noise patterns. |
|
| Post-processing in domain-specific pipelines (e.g., scientific visualization, medical content). |
| Use Case | Preferred Architecture (s) | Conditioning/Post-Processing | Key Metrics for Evaluation |
|---|---|---|---|
| Movie & Scene Generation | Stable Diffusion, Latent Diffusion [63,64] | BM3D, Non-local filtering | PSNR, SSIM |
| Recreational Content | Text-Guided Diffusion [31,32,65] | Mean/Median filtering | VGG-based perceptual similarity |
| Robotics, Remote Sensing | DDPM [66,67], Latent Diffusion [68,69,70,71,72] | Robust detail-preserving methods | SSIM, Temporal Consistency |
| Educational Content | DDPM, Latent Diffusion [73] | Transform domain denoising | SSIM, Perceptual Similarity |
| Category | Dataset | Images/Frame | Description | Use Case |
|---|---|---|---|---|
| General Image Generation Benchmarks | BSDS500 [74] | 500 images | Boundary detection dataset, beneficial for edge detection and segmentation applications. | Edge detection, boundary analysis |
| CIFAR-10 [75] | 60,000 images (50 k train, 10 k test) | Low-resolution (32 × 32) natural images across 10 object classes; widely used as a benchmark dataset for generative modeling and diffusion models | Benchmarking diffusion models, low-resolution image generation | |
| Imagenet [76] | 14 million | Widely used image classification dataset covering a broad spectrum of object categories. | Image classification, foundational model training | |
| LSUN [77] | 1 million+ | High-quality images for large-scale scene understanding, covering bedrooms, churches, outdoor scenes, etc. | Large-scale scene generation | |
| Places365 [78] | 1.8 million | Scene dataset covering various place types, such as landscapes, urban areas, and indoor scenes. | Scene generation, place recognition | |
| High-Fidelity Image Generation | BigGAN [16] | 14 million | Dataset specifically curated for GAN training, covering multiple high-quality categories. | High-fidelity image generation |
| CelebAMask-HQ [79] | ~30,000 | High-resolution face images with facial segmentation masks. | High-fidelity face generation | |
| DIV2K [80] | 1000 high-resolution images | High-quality 2K resolution images designed for super-resolution tasks, containing diverse natural scenes with fine details | Super-resolution, high-fidelity image generation | |
| FFHQ (Flickr-Faces-HQ) [13] | 70,000 | High-quality dataset of human faces for training and evaluating face generation models. | Face generation, high-resolution synthesis | |
| Scene Understanding | ADE20K [81] | 25,574 (train), 2000 (val) | Used for segmentation and object detection. | Scene segmentation, object detection |
| COCO-Stuff [82] | 118 k (train), 5 k (val) | Scene understanding tasks, including segmentation, object detection, and image captioning. | Scene generation, object detection | |
| HICO-DET [83] | 15,963 (train), 4034 (val) | Focuses on detecting Human–Object Interactions (HOI). | Human–object interaction modeling | |
| MS-COCO [84] | 328 k | Object Detection, Segmentation, Captioning, and Keypoint detection dataset. | Object detection, image captioning | |
| OpenImages [85] | ~9 million | Contains images annotated with object bounding boxes, segmentation masks, and visual relationships. | Object detection, segmentation, relationship modeling | |
| PASCAL VOC [86] | 11,530 images | Classic dataset for object detection and segmentation in general visual categories. | Object detection, segmentation | |
| Multimodal | CC3M [87] | 3.3 million | Image–caption pairs, designed for image captioning model training and evaluation. | Image captioning, multimodal training |
| CC12M [88] | 12 million | Image–text pairs for vision and language pre-training. | Vision–language pre-training | |
| Flickr [89] | 70,000 (train), 10,000 (val) | Benchmark for sentence-based image description. | Image captioning, sentence-based image tasks | |
| LAION-5B [90] | 5.85 billion | Large-scale image–text dataset, open source, for next-generation multimodal model training. | Text-to-image/video generation, multimodal training | |
| NUS-WIDE [91] | 269,648 images | Image–text pairs with diverse tags, ideal for text-to-image generation tasks with varied scene categories. | Text-to-image generation, image tagging | |
| Visual Genome [92] | ~108 k | Region descriptions and Q&A pairs associated with WordNet synsets for multimodal tasks. | Text-to-image generation, Q&A | |
| Video Generation | Charades [93] | 9848 videos | Real-world video dataset for action recognition and captioning, covering complex activities and interactions. | Human activity synthesis, video Q&A |
| DAVIS 2017 [94] | 90 video sequences | High-quality video dataset with segmentation masks, useful for video segmentation tasks. | Video segmentation, motion tracking | |
| HMDB51 [95] | ~7000 videos | Video dataset containing clips from various action categories, such as sports and human interactions. | Human action modeling, video generation | |
| Kinetics-700 [96] | ~650,000 videos | Video action recognition dataset with multiple classes, useful for generating human activity sequences. | Human action generation, video synthesis | |
| Sports1M [97] | ~1 million | Video dataset with sports-related content, tagged with metadata about activity types. | Sports action generation, video recognition | |
| VaTeX [98] | 41,250 videos, 825,000 captions | A large-scale multilingual video-and-language dataset, containing captions in both English and Chinese. | Video captioning, multilingual video research | |
| Vimeo90K [99] | 90,000 video clips | Video dataset for tasks like frame interpolation and video super-resolution. | Video frame interpolation, enhancement | |
| YouTube-8M [100] | 6.1 million video segments | Video dataset with audio–visual content across a wide range of categories. | Video generation, multimodal tasks | |
| 3D and Scene Generation | 3D Front [101] | - | Large-scale synthetic indoor scenes with professionally designed layouts. | Interior design, 3D scene generation |
| BlenderBot 3D [102] | 10 k scenes | 3D dataset for virtual environment generation, suitable for training 3D-aware generative models. | 3D scene generation, virtual simulations | |
| MatterPort [103] | 4939 (train), 456 (val) | RGB-D dataset for indoor scene understanding with 3D layouts. | Indoor scene understanding, VR/AR applications | |
| ModelNet [104] | 9843 (train), 2468 (val) | Contains synthetic object point clouds for 3D object recognition. | 3D object recognition | |
| OASIS [105] | 9664 (train), 1120 (val) | Large-scale dataset for single-image 3D reconstruction in wild settings. | 3D single-image reconstruction | |
| RPLAN [106] | 118 k | Densely annotated real residential floor plans. | Architectural layout and floor plan generation | |
| ScanNet [107] | 1201 (train), 312 (val) | Instance-level indoor RGB-D data with 2D and 3D annotations. | Indoor scene understanding, 3D object recognition | |
| ShapeNet [108] | 100,000+ | Large-scale dataset of 3D shapes, used for 3D object and scene understanding. | 3D shape generation and recognition | |
| SUNCG [109] | 53,860 | Large-scale synthetic 3D scenes with dense volumetric annotations. | 3D scene generation, virtual environment modeling | |
| YouTube 3D [110] | 47,125 (train), 1525 (val) | 3D vertex coordinates derived from YouTube videos. | 3D object tracking, video synthesis | |
| Autonomous Driving | Cityscapes [111] | 5000 | Focuses on semantic segmentation of urban street scenes. | Autonomous driving, urban scene generation |
| Indian Driving Dataset (IDD) [112] | ~10,000 | Road scene understanding dataset for unstructured environments. | Road scene generation, autonomous driving | |
| KITTI [113] | 22,000 | Street-level images with 3D bounding boxes for autonomous driving and object detection tasks. | Autonomous driving, street scene generation | |
| Synscapes [114] | - | Synthetic dataset for street scene parsing, created using photorealistic rendering techniques. | Street scene generation, autonomous driving | |
| YCB [115] | 80,000 (train), 5000 (val) | RGB-D scans for robotic manipulation, with high-resolution object meshes. | Robotics, manipulation training | |
| Depth and Geometry | DIODE [116] | 25,458 (train), 771 (val) | High-resolution color images for indoor and outdoor scenes. | Depth estimation, indoor–outdoor scene understanding |
| MegaDepth [117] | 130 k | Depth maps of complex scenes, valuable for depth-aware generation tasks. | Depth map generation, 3D modeling | |
| Specialized Datasets | CLEVR-G [118] | 10,000 (train), 10,000 (val) | Visual Question Answering dataset with 3D-rendered object images. | Visual Q&A, synthetic object generation |
| ColorMNIST [119] | 8000 (train), 8000 (val) | Synthetic binary classification task derived from MNIST. | Simple classification tasks | |
| DeepFashion [120] | ~800,000 | Large-scale clothing dataset for tasks such as garment segmentation and clothing category classification. | Clothing generation, virtual try-on | |
| Fashion-MNIST [121] | 70,000 images | Dataset for clothing item classification, useful for generative model testing in the fashion domain. | Fashion synthesis, classification |
| Metric | Key Idea | Advantages | Limitations | Typical Use Cases (in Diffusion Models) |
|---|---|---|---|---|
| PSNR (Peak Signal-to-Noise Ratio) | Measures pixel-level fidelity using logarithmic scale of inverse MSE between generated and ground truth images. | Simple to compute; widely used in denoising/restoration tasks. | Poor correlation with human perception; not meaningful for generative sampling. | Used in reconstruction tasks (e.g., super-resolution, inpainting) where ground truth is known. |
| SSIM/MS-SSIM (Structural Similarity Index) | Compares structure, contrast, and luminance between ground truth and generated images; MS-SSIM does so at multiple resolutions. | Better alignment with perceptual quality than PSNR; sensitive to structural distortions. | Still a handcrafted metric; may miss semantic mismatches or large shifts. | Used in pixel-level tasks with structural expectations (e.g., denoising, enhancement). |
| Perceptual Similarity | Evaluates feature-level differences using a pre-trained network (e.g., VGG), aiming to mimic human visual perception better than pixel-based comparisons. | - Focuses on semantic and structural aspects. - Generally aligns more closely with human perception than SSIM or PSNR. - Less sensitive to small pixel shifts. | - Depends on choice of pre-trained network (domain mismatch can affect results). - Computationally heavier (feature extraction). - Does not measure coverage of entire data distribution (pairwise only). | - Ideal for tasks requiring photo-realism or style consistency (e.g., artistic generation, super-resolution). - Useful for iterative refinement of diffusion outputs. |
| MOS (Mean Opinion Score) | Averages human ratings of visual quality across evaluators. | Captures real human perception; useful in creative domains. | Expensive, time-consuming, prone to bias. | Used in user-facing applications (e.g., art, UI, entertainment). |
| Fréchet Inception Distance (FID) | Distribution-based metric: compares means and covariances of real vs. generated images in a deep feature space (Inception v3). | Captures both quality and diversity; widely adopted benchmark. | Assumes Gaussian distribution; sensitive to batch size and extractor quality. | Standard metric for generative image models (unconditional, text-conditioned). |
| FVD (Fréchet Video Distance) | Extension of FID to video: compares temporal feature distributions using I3D network. | Captures temporal consistency; correlates well with human realism judgments. | Requires large sample size; assumes Gaussian statistics. | Standard for generative video evaluation (e.g., text-to-video diffusion). |
| IS (Inception Score) | Measures image recognizability and class diversity using pre-trained Inception classifier. | No need for real data; fast to compute. | Poor correlation with human judgment; insensitive to distribution mismatch. | Legacy metric; sometimes reported alongside FID for completeness. |
| Kernel Inception Distance (KID) | Uses a non-parametric MMD in the Inception feature space to compare real vs. generated distributions, avoiding Gaussian assumptions. | - More flexible than FID (no Gaussian assumption). - Can be more robust with certain data distributions. - Like FID, it looks at distribution coverage. | - Potentially more computational than FID (kernel computations). - Still depends on a feature extractor (Inception v3). - Requires real data for comparison. | - Alternative or complement to FID when data distribution does not fit Gaussian assumptions. - Used in high-quality generative tasks to confirm distribution alignment. |
| LPIPS (Learned Perceptual Image Patch Similarity) | Measures perceptual similarity in deep feature space using pretrained models like VGG or AlexNet. | Strong perceptual correlation; captures high-level textures and semantics. | Dependent on network architecture; does not assess distribution-wide realism. | Common for assessing perceptual quality in generative tasks (e.g., image synthesis, editing). |
| Model | Dataset | PSNR | SSIM | Perceptual Similarity | FID | IS | KID | LPIPS |
|---|---|---|---|---|---|---|---|---|
| DDPM | CIFAR-10 | 39.0722 | 0.9988 | 0.0234 | 2.5883 | 9.1185 ± 0.7170 | −0.00056 ± 0.000040 | 0.0006 |
| DDPM | DIV2K | 5.0102 | 0.0029 (MS: 0.0120) | 184.109 | 414.7513 | 1.0077 ± 0.0021 | 0.1066 ± 0.0008 (mean ± std) | 0.8485 |
| Stable Diff. | CIFAR-10 | 7.59 | 0.0099 | 0.1181 | 562.1695 | 1.2171 ± 0.0147 | 0.7698 ± 0.0000 | 0.9602 |
| Stable Diff. | DIV2K | 9.5 | 0.1401 | 20.0151 | 365.3383 | 2.3525 ± 0.2319 | 0.3239 ± 0.0000 | 0.789 |
| Latent Diff. | CIFAR-10 | 8.3245 | 0.0307 | 0.2108 | 147.9928 | 1.1469 ± 0.0045 | −0.0005 ± 0.0001 | 0.1875 |
| Latent Diff. | DIV2K | 6.69 | 0.0044 | 0.0662 | 426.47 | 1.0000 ± 0.0000 | 0.3720 ± 0.0011 | 0.8175 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Singh, A.; Chatta, N.K.; Vagula, Y.; Ehtesham, A.; Kumar, S.; Talaei Khoei, T. A Taxonomy of Generative Models with a Focus on Diffusion Models and Denoising Techniques. Electronics 2026, 15, 1293. https://doi.org/10.3390/electronics15061293
Singh A, Chatta NK, Vagula Y, Ehtesham A, Kumar S, Talaei Khoei T. A Taxonomy of Generative Models with a Focus on Diffusion Models and Denoising Techniques. Electronics. 2026; 15(6):1293. https://doi.org/10.3390/electronics15061293
Chicago/Turabian StyleSingh, Aditi, Nikhil Kumar Chatta, Yuvaraj Vagula, Abul Ehtesham, Saket Kumar, and Tala Talaei Khoei. 2026. "A Taxonomy of Generative Models with a Focus on Diffusion Models and Denoising Techniques" Electronics 15, no. 6: 1293. https://doi.org/10.3390/electronics15061293
APA StyleSingh, A., Chatta, N. K., Vagula, Y., Ehtesham, A., Kumar, S., & Talaei Khoei, T. (2026). A Taxonomy of Generative Models with a Focus on Diffusion Models and Denoising Techniques. Electronics, 15(6), 1293. https://doi.org/10.3390/electronics15061293

