From Pixels to Volumes: Generative AI in 3D Medical Imaging
Abstract
1. Introduction
Review Methodology and Search Strategy
2. Background and Preliminaries
2.1. Medical Imaging Modalities
2.2. Generative Model Fundamentals
2.3. Generative Model Families and Unified Notation
2.4. Key Mathematical Formulations
3. Taxonomy of 3D Generative Architectures
3.1. Voxel-Based Generative Models
3.2. Implicit Neural Representations
3.3. Latent Diffusion Models
4. Modality-Specific Advances
4.1. Magnetic Resonance Imaging (MRI)
4.2. Computed Tomography (CT)
4.3. Positron Emission Tomography (PET)
4.4. Foundational Cross-Modality Methods
4.5. Summary and Cross-Modality Observations
5. Cross-Modal and Multi-Modal Synthesis
5.1. MRI-to-CT and MRI-to-PET Synthesis
5.2. CT-to-MRI Synthesis
5.3. Universal and Contrast-Agnostic Synthesis
5.4. Summary and Emerging Directions
6. Applications in Clinical Practice
6.1. Data Augmentation for Fairness and Rare Disease
6.2. Radiology Education and AI-Assisted Training
6.3. Cardiac Digital Twins and Personalized Simulation
6.4. Federated Learning and Privacy-Preserving Synthesis
7. Evaluation Metrics and Benchmarks
7.1. Unified Evaluation Protocol
7.2. Limitations of Current Evaluation Metrics and Standardization Recommendations
8. Ethical Considerations and Limitations
8.1. Data Provenance and Privacy
8.2. Bias and Fairness
8.3. Anatomical Validity and Clinical Risk
8.4. Hallucination and Misrepresentation
8.5. Governance, Regulatory Classification, and Post-Deployment Oversight
8.6. An Operational Deployment Blueprint
8.7. Reproducibility and the Gap Between Research Performance and Clinical Deployment
- External validation. Only a few of the 16 landmark architectures studied have been reported with external validation; most reported external validation data originate from a single institution’s held-out test set and not from an independent test cohort. Any model that works well on a set of data acquired from the same acquisition pipeline as the training data will only have weak evidence of generalization.
- Cross-institutional generalization. External validation is related but different from the question of whether a model trained at one or a few sites will transfer to institutions with patient populations, referral patterns, and clinical protocols different from those used for training. Federated approaches (Section 6.4) partially overcome the data sharing barrier to multi-institutional evaluation, but federated training is in itself does not ensure the model’s equal performance on all of the institutions on which it is federated, nor for unseen institutions.
- Scanner and protocol variability. Report acquisition variability by modality (MRI sequence and field-strength dependence, CT reconstruction kernel and dose protocol, and PET tracer and scanner model variations); Section 8.2 also covers scanner-domain shift as another source of demographic and institutional bias. Do not assume that a model that is validated on one scanner manufacturer’s protocol will transfer to another without re-validation.
- Computational requirements. In Section 9.1, we will discuss the real-time inference bottleneck of iterative diffusion sampling, while the taxonomy discussion in Section 3.3 highlights the fact that latent diffusion models take a lot of data, denoising iterations, and compute to use. This is a different constraint apart from model performance: deployability on standard clinical hardware, not research grade GPU clusters.
- Data privacy. In Section 8.1, we cover the issue of memorization, membership-inference and training-data-reconstruction risk in depth and ultimately come to the conclusion that the synthetic origin of an image is not a guarantee of anonymity. These risks are applicable to any deployment pipeline that shares or releases model output or weights and should be explicitly tested for in the deployment pipeline.
- Expert assessment. Blinded expert (radiologist) evaluation is preferred over voxel-level metrics according to the three-tier evaluation protocol used in Section 7.1, and expert/human review is listed as a necessary step before operational deployment in the operational deployment blueprint in Section 8.6. Voxel-level metrics should only be considered as weak evidence for any translational claim without at least this level of evidence.
9. Open Challenges and Future Directions
9.1. Computational Efficiency and Real-Time Inference
9.2. Physics-Informed Generation
9.3. Rare Disease and Long-Tail Learning
9.4. Standardization of Benchmarks and Metrics
9.5. Foundation Models for Medical Imaging
9.6. Interpretability and Explainability
10. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; Bengio, Y. Generative adversarial nets. Adv. Neural Inf. Process. Syst. 2014, 27, 2672–2680. [Google Scholar]
- Kingma, D.P.; Welling, M. Auto-encoding variational Bayes. arXiv 2014, arXiv:1312.6114. [Google Scholar]
- Chung, H.; Ye, J.C. Score-based diffusion models for accelerated MRI. Med. Image Anal. 2022, 80, 102479. [Google Scholar] [CrossRef] [Scilit]
- Mildenhall, B.; Srinivasan, P.P.; Tancik, M.; Barron, J.T.; Ramamoorthi, R.; Ng, R. NeRF: Representing scenes as neural radiance fields for view synthesis. In Proceedings of the Computer Vision—ECCV 2020; Springer: Cham, Switzerland, 2020; pp. 405–421. [Google Scholar]
- Lustig, M.; Donoho, D.; Pauly, J.M. Sparse MRI: The application of compressed sensing for rapid MR imaging. Magn. Reson. Med. 2007, 58, 1182–1195. [Google Scholar] [CrossRef] [Scilit]
- Kalender, W.A. X-ray computed tomography. Phys. Med. Biol. 2006, 51, R29–R43. [Google Scholar] [CrossRef] [Scilit]
- Cherry, S.R.; Sorenson, J.A.; Phelps, M.E. Physics in Nuclear Medicine, 4th ed.; Elsevier: Amsterdam, The Netherlands, 2012. [Google Scholar]
- Vosoughi, S.; Roy, D.; Aral, S. The spread of true and false news online. Science 2018, 359, 1146–1151. [Google Scholar] [CrossRef] [Scilit]
- Abdali, S.; Shaham, S.; Krishnamachari, B. Multi-modal misinformation detection: Approaches, challenges and opportunities. ACM Comput. Surv. 2024, 57, 76. [Google Scholar] [CrossRef] [Scilit]
- Bickley, S.J.; Torgler, B. Cognitive architectures for artificial intelligence ethics. AI Soc. 2023, 38, 501–519. [Google Scholar] [CrossRef] [Scilit]
- Rezende, D.; Mohamed, S. Variational inference with normalizing flows. In Proceedings of the 32nd International Conference on Machine Learning, Lille, France, 6–11 July 2015; ACM: New York, NY, USA, 2015; Volume 37, pp. 1530–1538. [Google Scholar]
- Zhu, J.-Y.; Park, T.; Isola, P.; Efros, A.A. Unpaired image-to-image translation using cycle-consistent adversarial networks. In Proceedings of the International Conference on Computer Vision; IEEE: New York, NY, USA, 2017; pp. 2223–2232. [Google Scholar]
- Wu, J.; Zhang, C.; Xue, T.; Freeman, B.; Tenenbaum, J. Learning a probabilistic latent space of object shapes via 3D generative-adversarial modeling. In Proceedings of the NeurIPS Annual Conference on Neural Information Processing Systems; Neural Information Processing Systems Foundation, Inc.: San Diego, CA, USA, 2016; Volume 29, pp. 82–90. [Google Scholar]
- Zhao, A.; Balakrishnan, G.; Durand, F.; Guttag, J.V.; Dalca, A.V. Data Augmentation Using Learned Transformations for One-shot Medical Image Segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019; IEEE: New York, NY, USA, 2019; pp. 8543–8553. [Google Scholar]
- Dalca, A.V.; Guttag, J.; Sabuncu, M.R. Anatomical Priors in Convolutional Networks for Unsupervised Biomedical Segmentation. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; IEEE: New York, NY, USA, 2018; pp. 9290–9299. [Google Scholar]
- Xie, J.; Zheng, Z.; Gao, R.; Wang, W.; Zhu, S.-C.; Wu, Y.N. Learning Descriptor Networks for 3D Shape Synthesis and Analysis. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; IEEE: New York, NY, USA, 2018; pp. 8629–8638. [Google Scholar]
- Baur, C.; Denner, S.; Wiestler, B.; Navab, N.; Albarqouni, S. Autoencoders for unsupervised anomaly segmentation in brain MR images: A comparative study. Med. Image Anal. 2021, 69, 101952. [Google Scholar] [CrossRef] [Scilit]
- Luo, G.; Blumenthal, M.; Heide, M.; Uecker, M. Bayesian MRI reconstruction with joint uncertainty estimation using diffusion models. Magn. Reason. Med. 2023, 90, 295–311. [Google Scholar] [CrossRef] [Scilit]
- Khader, F.; Müller-Franzes, G.; Arasteh, S.T.; Han, T.; Haarburger, C.; Schulze-Hagen, M.; Schad, P.; Engelhardt, S.; Baeßler, B.; Foersch, S.; et al. Denoising diffusion probabilistic models for 3D medical image generation. Sci. Rep. 2023, 13, 7303. [Google Scholar] [CrossRef] [Scilit]
- Chen, H.; Zhang, Y.; Kalra, M.K.; Lin, F.; Chen, Y.; Liao, P.; Zhou, J.; Wang, G. Low-dose CT with a residual encoder-decoder convolutional neural network. IEEE Trans. Med. Imaging 2017, 36, 2524–2535. [Google Scholar] [CrossRef] [Scilit]
- Maspero, M.; Savenije, M.H.F.; Dinkla, A.M.; Seevinck, P.R.; Intven, M.P.W.; Juergenliemk-Schulz, I.M.; Kerkmeijer, L.G.W.; Berg, C.A.T.V.D. Dose evaluation of fast synthetic-CT generation using a generative adversarial network for general pelvis MR-only radiotherapy. Phys. Med. Biol. 2018, 63, 185001. [Google Scholar] [CrossRef] [Scilit]
- Reed, A.W.; Kim, H.; Anirudh, R.; Mohan, K.A.; Champley, K.; Kang, J.; Jayasuriya, S. Dynamic CT Reconstruction from Limited Views with Implicit Neural Representations and Parametric Motion Fields. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 10–17 October 2021; IEEE: New York, NY, USA, 2021; pp. 2258–2268. [Google Scholar]
- Kunz, J.F.; Ruschke, S.; Heckel, R. Implicit neural networks with Fourier-feature inputs for free-breathing cardiac MRI reconstruction. IEEE Trans. Comput. Imaging 2024, 10, 1280–1289. [Google Scholar] [CrossRef] [Scilit]
- Song, Y.; Sohl-Dickstein, J.; Kingma, D.P.; Kumar, A.; Ermon, S.; Poole, B. Score-based generative modeling through stochastic differential equations. arXiv 2021, arXiv:2011.13456. [Google Scholar]
- Pinaya, W.H.L.; Tudosiu, P.-D.; Dafflon, J.; Da Costa, P.F.; Fernandez, V.; Nachev, P.; Ourselin, S.; Cardoso, M.J. Brain imaging generation with latent diffusion models. In Deep Generative Models—DGM4MICCAI 2022 Workshop; Springer: Berlin/Heidelberg, Germany, 2022; Volume 13609, pp. 117–126. [Google Scholar] [CrossRef] [Scilit]
- Wolleb, J.; Sandkühler, R.; Bieder, F.; Valmaggia, P.; Cattin, P.C. Diffusion models for implicit image segmentation ensembles. In Proceedings of the 5th International Conference on Medical Imaging with Deep Learning; PMLR: New York, NY, USA, 2022. [Google Scholar]
- Gong, K.; Johnson, K.; El Fakhri, G.; Li, Q.; Pan, T. PET image denoising based on denoising diffusion probabilistic model. Eur. J. Nucl. Med. Mol. Imaging 2024, 51, 358–368. [Google Scholar] [CrossRef] [Scilit]
- Chung, H.; Kim, J.; McCann, M.T.; Klasky, M.L.; Ye, J.C. Diffusion posterior sampling for general noisy inverse problems. In Proceedings of the International Conference on Learning Representations, Kigali, Rwanda, 1–5 May 2023. [Google Scholar]
- Dewey, B.E.; Zhao, C.; Reinhold, J.C.; Carass, A.; Fitzgerald, K.C.; Sotirchos, E.S.; Saidha, S.; Oh, J.; Pham, D.L.; Calabresi, P.A.; et al. DeepHarmony: A deep learning approach to contrast harmonization across scanner changes. NeuroImage Clin. 2019, 24, 101945. [Google Scholar]
- Zha, R.; Zhang, Y.; Li, H. NAF: Neural Attenuation Fields for Sparse-View CBCT Reconstruction. In Proceedings of the 25th International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI 2022), Singapore, 18–22 September 2022; pp. 442–452. [Google Scholar]
- Elbakri, I.A.; Fessler, J.A. Statistical image reconstruction for polyenergetic X-ray computed tomography. IEEE Trans. Med. Imaging 2002, 21, 89–99. [Google Scholar] [CrossRef] [Scilit]
- Ozbey, M.; Dalmaz, O.; Dar, S.U.H.; Bedel, H.A.; Özturk, Ş.; Güngör, A.; Çukur, T. Unsupervised medical image translation with adversarial diffusion models. IEEE Trans. Med. Imaging 2023, 42, 3524–3539. [Google Scholar] [CrossRef] [Scilit]
- Luo, X.; Chen, J.; Song, T.; Wang, G.; Zhang, S. Semi-supervised medical image segmentation through dual-task consistency. In Proceedings of the Thirty-Fifth AAAI Conference on Artificial Intelligence; AAAI Press: Palo Alto, CA, USA, 2021; Volume 35, pp. 8801–8809. [Google Scholar]
- Wang, Y.; Zhou, L.; Yu, B.; Wang, L.; Zu, C.; Lalush, D.S.; Lin, W.; Wu, X.; Zhou, J.; Shen, D. 3D auto-context-based locality adaptive multi-modality GANs for PET synthesis. IEEE Trans. Med. Imaging 2019, 38, 1328–1339. [Google Scholar] [CrossRef] [Scilit]
- Gong, C.; Huang, Y.; Luo, M.; Cao, S.; Gong, X.; Ding, S.; Yuan, X.; Zheng, W.; Zhang, Y. Channel-wise attention enhanced and structural similarity constrained cycleGAN for effective synthetic CT generation from head and neck MRI images. Radiat. Oncol. 2024, 19, 37. [Google Scholar] [CrossRef] [Scilit]
- Pan, S.; Abouei, E.; Wynne, J.; Chang, C.W.; Wang, T.; Qiu, R.L.; Li, Y.; Peng, J.; Roper, J.; Patel, P.; et al. Synthetic CT generation from MRI using 3D transformer-based denoising diffusion model. Med. Phys. 2024, 51, 2538–2548. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.; Yakushev, I.; Hedderich, D.M.; Wachinger, C. Translating MRI to PET through conditional diffusion models with enhanced pathology awareness. Med. Image Anal. 2026, 111, 104035. [Google Scholar] [CrossRef] [Scilit]
- Ni, Y.; Ma, J.; Chen, J. Medical volume CT-to-MRI translation with multi-dimensional diffusion architecture. Biomed. Signal Process. Control 2026, 112, 108627. [Google Scholar] [CrossRef] [Scilit]
- Billot, B.; Greve, D.N.; Puonti, O.; Thielscher, A.; Van Leemput, K.; Fischl, B.; Dalca, A.V.; Iglesias, J.E. SynthSeg: Segmentation of brain MRI scans of any contrast and resolution without retraining. Med. Image Anal. 2023, 86, 102789. [Google Scholar] [CrossRef] [Scilit]
- Pan, S.; Abouei, E.; Peng, J.; Qian, J.; Wynne, J.F.; Wang, T.; Chang, C.W.; Roper, J.; Nye, J.A.; Mao, H.; et al. Full-dose whole-body PET synthesis from low-dose PET using high-efficiency denoising diffusion probabilistic model: PET consistency model. Med. Phys. 2024, 51, 5468–5478. [Google Scholar] [CrossRef] [Scilit]
- Ktena, I.; Wiles, O.; Albuquerque, I.; Rebuffi, S.A.; Tanno, R.; Roy, A.G.; Azizi, S.; Belgrave, D.; Kohli, P.; Cemgil, T.; et al. Generative models improve fairness of medical classifiers under distribution shifts. Nat. Med. 2024, 30, 1166–1173. [Google Scholar] [CrossRef] [Scilit]
- Sambhu, P.; Guin, O.; Sambhu, M.; Cha, J. Curriculum learning with synthetic data for enhanced pulmonary nodule detection in chest radiographs. arXiv 2025, arXiv:2510.07681. [Google Scholar]
- Lyu, X.; Dong, L.; Fan, Z.; Sun, Y.; Zhang, X.; Liu, N.; Wang, D. Artificial intelligence-based graded training of pulmonary nodules for junior radiology residents and medical imaging students. BMC Med. Educ. 2024, 24, 740. [Google Scholar] [CrossRef] [Scilit]
- Qian, S.; Ugurlu, D.; Fairweather, E.; Toso, L.D.; Deng, Y.; Strocchi, M.; Cicci, L.; Jones, R.E.; Zaidi, H.; Prasad, S.; et al. Developing cardiac digital twin populations powered by machine learning provides electrophysiological insights in conduction and repolarization. Nat. Cardiovasc. Res. 2025, 4, 624–636. [Google Scholar] [CrossRef] [Scilit]
- Koopsen, T.; Gerrits, W.; van Osta, N.; van Loon, T.; Wouters, P.; Prinzen, F.W.; Cicci, L.; Jones, R.E.; Zaidi, H.; Prasad, S.; et al. Virtual pacing of a patient’s digital twin to predict left ventricular reverse remodelling after cardiac resynchronization therapy. Europace 2024, 26, euae009. [Google Scholar] [CrossRef] [Scilit]
- Thangaraj, P.M.; Benson, S.H.; Oikonomou, E.K.; Asselbergs, F.W.; Khera, R. Cardiovascular care with digital twin technology in the era of generative artificial intelligence. Eur. Heart J. 2024, 45, 4808–4821. [Google Scholar] [CrossRef] [Scilit]
- Kulkarni, P.; Kanhere, A.; Kukreja, H.; Zhang, V.; Yi, P.H.; Parekh, V.S. Improving multi-center generalizability of GAN-based fat suppression using federated learning. arXiv 2024, arXiv:2404.07374. [Google Scholar]
- Perumal, M.; Srinivas, M. FMed-diffusion federated learning on medical image diffusion. bioRxiv 2025. [Google Scholar] [CrossRef] [Scilit]
- Baid, U.; Ghodasara, S.; Mohan, S.; Bilello, M.; Calabrese, E.; Colak, E.; Farahani, K.; Kalpathy-Cramer, J.; Kitamura, F.C.; Pati, S.; et al. The RSNA-ASNR-MICCAI BraTS 2021 benchmark on brain tumor segmentation and radiogenomic classification. arXiv 2021, arXiv:2107.02314. [Google Scholar]
- Li, H.; Conte, G.M.; Anwar, S.M.; Kofler, F.; Ezhov, I.; van Leemput, K.; Piraud, M.; Diaz, M.; Cole, B.; Calabrese, E.; et al. Brain Tumor Segmentation (BraTS) challenge 2023: Brain MR image synthesis for tumor segmentation (BraSyn). arXiv 2023, arXiv:2305.09011. [Google Scholar]
- Armato, S.G., III; McLennan, G.; Bidaut, L.; McNitt-Gray, M.F.; Meyer, C.R.; Reeves, A.P.; Zhao, B.; Aberle, D.R.; Henschke, C.I.; Hoffman, E.A.; et al. The lung image database consortium (LIDC) and image database resource initiative (IDRI): A completed reference database of lung nodules on CT scans. Med. Phys. 2011, 38, 915–931. [Google Scholar] [CrossRef] [Scilit]
- Zbontar, J.; Knoll, F.; Sriram, A.; Murrell, T.; Huang, Z.; Muckley, M.J.; Defazio, A.; Stern, R.; Johnson, P.; Bruno, M.; et al. FastMRI: An open dataset and benchmarks for accelerated MRI. arXiv 2018, arXiv:1811.08839. [Google Scholar]
- Kavur, A.E.; Gezer, N.S.; Barış, M.; Aslan, S.; Conze, P.-H.; Groza, V.; Pham, D.D.; Chatterjee, S.; Ernst, P.; Özkan, S.; et al. CHAOS challenge—Combined (CT-MR) healthy abdominal organ segmentation. Med. Image Anal. 2021, 69, 101950. [Google Scholar] [CrossRef] [Scilit]
- Bernard, O.; Lalande, A.; Zotti, C.; Cervenansky, F.; Yang, X.; Heng, P.A.; Cetin, I.; Lekadir, K.; Camara, O.; Gonzalez Ballester, M.A.; et al. Deep learning for automatic MRI cardiac multi-structures segmentation and diagnosis. IEEE Trans. Med. Imaging 2018, 37, 2514–2525. [Google Scholar] [CrossRef] [Scilit]
- Jack, C.R., Jr.; Bernstein, M.A.; Fox, N.C.; Thompson, P.; Alexander, G.; Harvey, D.; Borowski, B.; Britson, P.J.; L Whitwell, J.; Ward, C.; et al. The Alzheimer’s Disease Neuroimaging Initiative (ADNI): MRI methods. J. Magn. Reson. Imaging 2008, 27, 685–691. [Google Scholar] [CrossRef] [Scilit]
- McCollough, C.; Chen, B.; Holmes, D.; Duan, X.; Yu, Z.; Xu, L.; Leng, S.; Fletcher, J. Low dose CT image and projection data [data set]. Cancer Imaging Arch. 2020, 10, 174. [Google Scholar]
- Andrearczyk, V.; Oreiller, V.; Abobakr, M.; Akhavanallaf, A.; Balermpas, P.; Boughdad, S.; Capriotti, L.; Castelli, J.; Le Rest, C.C.; Decazes, P.; et al. Overview of the HECKTOR challenge at MICCAI 2022: Automatic head and neck tumor segmentation and outcome prediction in PET/CT. In Head and Neck Tumor Segmentation and Outcome Prediction; Springer: Cham, Switzerland, 2023. [Google Scholar] [CrossRef] [Scilit]
- Wijethilake, N.; Dorent, R.; Ivory, M.; Kujawa, A.; Cornelissen, S.; Langenhuizen, P.; Okasha, M.; Oviedova, A.; Dong, H.; Kang, B.; et al. CrossMoDA Challenge: Evolution of cross-modality domain adaptation techniques for VS segmentation 2021–2023. arXiv 2025, arXiv:2506.12006. [Google Scholar]
- Farhoudian, A.; Heidari, A.; Shahhosseini, R. A new era in colorectal cancer: Artificial Intelligence at the forefront. Comput. Biol. Med. 2025, 196, 110926. [Google Scholar] [CrossRef] [Scilit]
- Toumaj, S.; Heidari, A. Explainable artificial intelligence in neurology: A holistic exploration of models for diagnosis, progression tracking, and treatment planning. Artif. Intell. Rev. 2026, 59, 217. [Google Scholar] [CrossRef] [Scilit]






| Symbol | Definition | Used in |
|---|---|---|
| Original (clean) image or volume, | Section 2.4 and Section 3.3 | |
| Noised/latent version of at diffusion timestep t | Section 2.4, Section 3.3 and Section 9.2 | |
| t, T | Diffusion timestep (t ∈ {1, …, T} discrete, or t ∈ [0, 1] continuous SDE); T = total number of denoising steps | Section 2.4, Section 3.3 and Section 9.1 |
| ε, εθ | Ground-truth Gaussian noise; noise predicted by the network with parameters θ | Section 2.4 and Section 3.3 |
| z | Latent variable in a compressed (VAE) representation | Section 2.4, Section 3.1 and Section 3.3 |
| θ | Learnable parameters of the generative network (generator, denoiser, or decoder) | Throughout |
| sθ(x,t) | Score function (gradient of the log-density) estimated by a neural network | Section 2.4 and Section 3.1 |
| G, D | Generator and discriminator networks in adversarial (GAN) training | Section 2.4 and Section 3.1 |
| q(xt|xt−1) | Forward (noising) transition distribution | Section 2.4 and Section 3.3 |
| pθ(xt−1|xt) | Reverse (denoising) transition distribution parameterized by θ | Section 2.4 and Section 3.3 |
| p(x), pdata(x) | Model distribution/true data distribution over images or volumes | Section 2.4 and Section 7 |
| Cumulative noise schedule (product of per-step retained-signal ratios) used in the closed-form forward diffusion process | Section 2.4 and Section 3.3 | |
| σ(t) | Noise scale/schedule as a function of diffusion time (SDE formulation) | Section 2.4 and Section 3.3 |
| f(x,t), g(t) | Drift and diffusion coefficients of the forward SDE, dx = f(x,t)dt + g(t)dw | Section 2.4 |
| Ref. No. | Model Type | Key Techniques Used | Strengths | Limitations | Target Task | Code | Dataset |
|---|---|---|---|---|---|---|---|
| [13] | Voxel-Based GAN | 3D convolutional GAN; volumetric occupancy grids; probabilistic latent space; ShapeNet training | First 3D-GAN for shape synthesis; learns shape distributions; enables sampling and interpolation | Memory-intensive; mode collapse risk; cubic scaling of resolution cost | Shape/volume synthesis | https://github.com/zck119/3dgan-release (accessed on 1 August 2026) | Public (ShapeNet) |
| [14] | Voxel-Based GAN | GAN-based augmentation; learned spatial transforms; one-shot registration; synthesized training pairs | Significantly improves segmentation with minimal labels; adaptable to rare anatomy | Relies on atlas quality; limited to deformable structures; requires aligned input pairs | Data augmentation (one-shot seg.) | Not reported | Not reported |
| [15] | Voxel-Based VAE | Probabilistic generative model; anatomical shape priors; unsupervised latent segmentation | No manual labels needed; incorporates domain knowledge; robust across scanners | Limited baseline comparison; requires pre-registration steps; slow convergence | Unsupervised segmentation | Not reported | Public (ADNI/brain MRI) |
| [16] | Voxel-Based EBM | Deep 3D energy-based model; MCMC sampling; analysis by synthesis via maximum likelihood | No auxiliary networks (unlike GANs/VAEs); avoids mode collapse; realistic 3D shape generation | MCMC sampling computationally slow; voxel representation memory-heavy at high resolution | Shape synthesis | Not reported | Not reported |
| [17] | Voxel-Based VAE (Anomaly) | Comparative VAE study; reconstruction-error anomaly scoring; unsupervised lesion detection | No labeled lesion data required; surveys multiple VAE variants; strong generalization | Performance depends on healthy data quality; sensitive to noise in voxel-wise detection | Unsupervised anomaly detection | Not publicly reported | Public (brain MRI) |
| [18] | Voxel-Based Score Model | Score-based diffusion (denoising score matching) + Markov chain Monte Carlo (MCMC/Langevin) posterior sampling; k-space data-consistency term incorporated into the reverse diffusion at every step | Works from highly undersampled k-space (tested at 4×–10× and up to 8× on FLAIR); no hand-crafted sparsity prior needed (learned generative prior replaces wavelet/ℓ1 regularization); also yields pixel-wise uncertainty (variance) maps as a by-product | Computationally intensive (~10 min per reconstruction vs. ~5 s for ℓ1-ESPIRiT with BART); hallucinations/artifacts can appear at high acceleration | MRI reconstruction (undersampled k-space), brain imaging (T1w, T2w, FLAIR, T2*w) | https://github.com/mrirecon/spreco (accessed on 1 August 2026) | Public (fastMRI) |
| [19] | Voxel-Based Diffusion | DDPM; 3D volume generation; noise scheduling adapted for clinical MRI/CT volumes | Realistic anatomically correct 3D images; addresses data privacy and scarcity | Computationally intensive training and inference; complex implementation for clinical deployment | 3D volumetric synthesis (MRI/CT) | https://github.com/FirasGit/medicaldiffusion (accessed on 1 August 2026) | Public (BraTS, LIDC-IDRI) |
| [20] | Voxel-Based CNN | Residual encoder–decoder; perceptual loss; 3D volumetric convolutions; paired low/full-dose CT | Clinically indistinguishable in 91% of cases; real-time inference | Requires paired training data; may over-smooth fine textures; site-specific training needed | Low-dose CT denoising | Not reported | Public (AAPM Low-Dose CT) |
| [21] | Voxel-Based GAN (Cross-modal) | CycleGAN (3D); unpaired MRI-to-CT synthesis; Hounsfield Unit prediction; dose-volume evaluation | Eliminates separate CT scan; low MAE; DSC 0.91 on bone structures; enables MRI-only workflows | High memory footprint; requires pre-registration; limited generalizability across scanners | Cross-modal translation (MRI→CT) | Not reported | Not reported (institutional) |
| [22] | Implicit (NeRF) | NeRF coordinate-based rendering; multi-view X-ray/CT synthesis; volume rendering integral; parametric motion | Memory-efficient 3D representation; continuous resolution; strong few-shot view synthesis | Long inference time per point evaluation; difficulty on non-rigid anatomy; noisy clinical data challenges | Sparse-view/dynamic CT reconstruction | https://github.com/awreed/DynamicCTReconstruction (accessed on 1 August 2026) | Public (LIDC-IDRI) |
| [23] | Untrained implicit neural representation | Fourier-feature MLP with separate spatial/temporal embeddings; per-scan fitting via k-space data-consistency loss only; no explicit regularizer | No training data/ground truth or ECG needed; on par with/slightly beats SOTA untrained CNN (t-DIP); clearly beats other implicit-representation baselines (NIK, KFMLP); easily extends to 3D | High computational cost; tested on only one healthy volunteer (limited generalizability); no real-data ground truth (relies on SER + visual comparison, with a phantom used for true SSIM/VIF) | Free-breathing, ungated, real-time cardiac cine MRI reconstruction | Public—https://github.com/MLI-lab/cinemri (accessed on 1 August 2026) | Public in-house free-breathing cardiac MRI dataset (DOI 10.21227/f057-dw29) + MRXCAT synthetic phantom |
| [24] | Latent Diffusion (LDM) | Score-based SDE framework; DDPM generalization; classifier-free guidance; stochastic generation | High sample quality; stable training dynamics; unifies DDPM and NCSN frameworks | Slow sequential sampling; computationally expensive at inference; large model size | Generative framework (theory/benchmark) | https://github.com/yang-song/score_sde (accessed on 1 August 2026) | Public (CIFAR-10, CelebA) |
| [25] | Latent Diffusion (3D) | VAE latent space diffusion; 3D U-Net backbone; multi-modal conditioning; brain MRI synthesis | Anatomically plausible MRI/CT synthesis; multi-modal capability; scalable compression | Requires large datasets; complex latent space tuning; resolution limited by VAE bottleneck | 3D brain MRI synthesis | https://github.com/Project-MONAI/GenerativeModels (accessed on 1 August 2026) | Public (BraTS-derived) |
| [26] | Latent Diffusion (Medical) | Task-specific conditioning; anatomy-guided diffusion; segmentation ensemble generation | Generates rare pathologies; preserves structural integrity; interpretable conditioning | Limited to isotropic resolutions; needs paired data for conditioning; slow inference | Segmentation-ensemble/rare-pathology synthesis | Not reported | Public (segmentation-conditioned) |
| [27] | Latent Diffusion (PET) | DDPM-based PET denoising; MRI-guided prior integration; Poisson noise-aware iterative refinement | Outperforms non-local mean and U-Net baselines; flexible use of MRI prior; handles multiple tracers | Slow inference time; relies on MRI prior availability; limited to 2D slice-wise processing | PET denoising | Not reported | Not reported (clinical PET) |
| [28] | Latent Diffusion (MRI) | Forward model integration; physics-guided DDPM sampling; k-space consistency enforcement | Accelerated MRI reconstruction; uncertainty quantification; physical measurement consistency | Sensitive to forward model errors; computationally intensive; limited multi-coil support | General inverse-problem solving (MRI/CT recon.) | https://github.com/DPS2022/diffusion-posterior-sampling (accessed on 1 August 2026) | Not reported |
| Modality | Key Data Characteristics | Modality-Specific Challenges | Important Preprocessing | Generative Applications | Recommended Evaluation Considerations |
|---|---|---|---|---|---|
| MRI | Multi-sequence (T1, T2, FLAIR, DWI, etc.); scanner-, vendor-, and site-dependent intensity distributions; variable contrast | Scanner-, vendor-, field-strength-, and sequence-dependent intensity variation; motion and noise sensitivity; intensities not directly quantitative | Registration, resampling, bias-field correction, and intensity normalization | Reconstruction, synthesis, cross-sequence translation, and augmentation | PSNR, SSIM, MAE, and NMSE; downstream segmentation/task performance |
| CT | Quantitative attenuation expressed in Hounsfield Units (HU); relatively standardized intensity scale across scanners | Preservation of HU values; dose–noise trade-off (low dose vs. full dose); anatomical fidelity at reduced dose | HU clipping/windowing, normalization, and resampling | Low-dose reconstruction, denoising, and modality synthesis (e.g., MR-to-CT) | HU-MAE/HU-RMSE, PSNR, SSIM, noise, and structural fidelity |
| PET | Tracer-dependent uptake distribution (e.g., FDG and PSMA); quantitative Standardized Uptake Value (SUV); dose- and count-dependent statistics | Tracer distribution and kinetics; low-count/low-dose statistical noise; limited spatial resolution; SUV preservation | SUV normalization, registration, and resampling | Low-dose PET reconstruction, PET synthesis, and denoising | SUV error/bias, MAE/RMSE, PSNR, SSIM, and lesion-level detection metrics |
| Key Model | Modality | Anatomical Region/Population | Conditioning Input | Evaluation Protocol | Main Challenge | Approach | Key Metric/Result | Dataset/Code Link If Available | Comparability Flag and Rationale | Target Task | Ref. |
|---|---|---|---|---|---|---|---|---|---|---|---|
| MRI—Magnetic Resonance Imaging | |||||||||||
| Bayesian Diffusion-MCMC Reconstruction (score-based generative model) | MRI | Brain; healthy volunteers (in-house, n = 13) + fastMRI brain subset (T1w post-contrast, T2w, FLAIR)—not knee | Undersampled multi-coil k-space; coil sensitivity maps from ESPIRiT; no paired ground-truth needed at inference | Retrospective, simulated undersampling (single-coil unfolding, multicoil, 4×/8×/10×); compared against ℓ1-wavelet-regularized reconstruction | Ill-posedness of reconstruction from undersampled k-space; quantifying reconstruction uncertainty; heavy computational cost of MCMC sampling | Score-based (denoising score matching) generative diffusion prior + Bayesian posterior sampling via MCMC/Langevin dynamics | PSNR ≈ 34–37 dB, SSIM ≈ 0.90–0.94 depending on setting; recovers finer detail than ℓ1-wavelet regularization | In-house brain dataset (1300 images, 13 healthy volunteers) + public fastMRI (brain subset). Code: https://github.com/mrirecon/spreco (accessed on 1 August 2026) | ✗ Not directly comparable to other rows—uses a private in-house dataset in addition to fastMRI, and MCMC-based posterior sampling | MRI reconstruction (undersampled k-space) with joint uncertainty estimation | [18] |
| Medical Diffusion (DDPM) | MRI/CT | Brain tumor (BraTS); lung nodule (LIDC-IDRI) | Unconditional generation | FID vs. real-data distribution | Data scarcity; privacy concerns | Denoising diffusion probabilistic model (DDPM) for 3D volumetric generation | Realistic 3D brain MRI and lung CT synthesis; FID scores comparable to real data | BraTS, LIDC-IDRI Link—https://github.com/FirasGit/medicaldiffusion (accessed on 1 August 2026) | △ Shares BraTS w/DeepHarmony, LIDC-IDRI w/NeRF-CT—different task/metric each time | 3D volumetric synthesis (MRI/CT) | [19] |
| Fourier-feature MLP implicit neural representation | MRI (cardiac cine, real-time) | • Heart • 1 healthy volunteer (30 y, male), 3T Philips Elition X—not fastMRI, not a population study | • Continuously acquired, undersampled multi-coil k-space • Free-breathing, ungated, partial-Fourier Cartesian sampling | • Retrospective; untrained/self-supervised (no training data) • Baselines: t-DIP (CNN), NIK, KFMLP | Continuous cardiac + respiratory motion during acquisition • No ground-truth data available for real-time free-breathing MRI | Untrained implicit neural network (MLP) with separate spatial and temporal Fourier-feature embeddings | SER ≈ 9–17 dB across datasets—on par with/slightly better than t-DIP, clearly better than NIK/KFMLP | • Public dataset: DOI 10.21227/f057-dw29 • Public code: https://github.com/MLI-lab/cinemri (accessed on 1 August 2026) | ✗ Not directly comparable to fastMRI-benchmarked rows—own single-subject dataset, no PSNR/SSIM ground truth (uses SER instead) | • Free-breathing, ungated real-time cardiac cine MRI reconstruction (untrained method) | [23] |
| DeepHarmony | MRI | Multi-site; mixed healthy + BraTS tumor cohorts | Paired T1→T2 style transfer | Downstream cross-scanner segmentation consistency (task-based) | Domain shift across scanner manufacturers/vendors | GAN-based contrast harmonization (T1→T2 style transfer) | Improved cross-scanner segmentation consistency; preserves anatomical fidelity | IXI, ABIDE, and BraTS | △ Shares BraTS with Medical Diffusion below—different task/metric, not comparable | Scanner/contrast harmonization (style transfer) | [29] |
| CT—Computed Tomography | |||||||||||
| RED-CNN | CT | Abdomen (AAPM cohort) | Paired low-dose/full-dose CT | Reader study (clinical indistinguishability) + PSNR | Ionizing radiation dose; low-dose noise | Residual encoder–decoder CNN with perceptual loss and adversarial training | PSNR improvement at 25% dose; preserves high-frequency details; clinically indistinguishable in 91% of cases | AAPM Low-Dose CT Challenge | ✗ Not comparable to NeRF-CT below—reader-study result vs. pure SSIM value | Low-dose CT denoising | [20] |
| NeRF for Sparse-View CT | CT | Lung nodule cohort (public LIDC-IDRI) | Sparse projections (≤10 views) vs. FBP | SSIM vs. filtered backprojection baseline only | Radiation exposure; limited projections | Neural radiance fields (NeRF) with volume rendering from sparse views (≤10 projections) | SSIM 0.91 from 10 views vs. 0.78 for FBP; reduced streak artifacts | LIDC-IDRI | △ Shares LIDC-IDRI with Medical Diffusion—different task, not comparable | Sparse-view CT reconstruction | [30] |
| Penalized-likelihood statistical image reconstruction with ordered-subsets algorithm for polyenergetic X-ray CT | CT | Simulated numerical phantom containing bone and soft tissue | Known energy-dependent mass attenuation coefficients per material; assumed polyenergetic X-ray source spectrum | Simulated X-ray CT transmission measurements of a two-material (bone/soft-tissue) phantom; | Conventional (monoenergetic) reconstruction methods ignore the polyenergetic nature of the X-ray source, causing severe beam-hardening artifacts. | Physical model accounting for polyenergetic spectrum and energy-dependent attenuation; penalized-likelihood cost function optimized via an ordered-subsets | Substantially reduced beam-hardening artifacts compared to conventional/monoenergetic reconstruction; improved voxel-density accuracy for bone/soft-tissue mixtures | No public dataset or code release; simulation-only data generated by the authors | ✗ Not comparable | Beam-hardening artifact reduction and material-density estimation in CT image reconstruction | [31] |
| PET—Positron Emission Tomography | |||||||||||
| Denoising Diffusion Probabilistic Model | PET | Human brain imaging; 120 18F-FDG datasets and 140 18F-MK-6240 (tau) datasets | Low-count/noisy PET image and/or MR prior image, supplied either as direct network input | Regional and surface-based quantification on FDG and MK-6240 datasets; | PET image quality is degraded by physical degradation factors and limited photon counts, especially in low-dose/low-count acquisitions s | DDPM iteratively transforms a normal distribution into the target PET data distribution; evaluates variants where PET/MR prior images are supplied as network input versus as a refinement-step data-consistency constraint | DDPM-based methods incorporating PET information outperform nonlocal-mean and U-Net-based denoising; best performance achieved using MR prior as network input combined with PET as a data-consistency constraint during inference | Clinical datasets not publicly released; no code | ✗ Comparable within the PET deep-generative denoising literature and to other DDPM-based PET works using similar tracers | PET image denoising/noise reduction leveraging prior anatomical or PET information | [27] |
| Attention-Based cGAN for PET Synthesis | PET | Alzheimer’s cohort (ADNI) | Paired MRI→PET (conditional) | MAE/PSNR/SSIM vs. CycleGAN baseline; SUV-preservation check | Missing PET data in multi-modal studies | Conditional GAN (cGAN) with attention mechanism for MRI→PET synthesis | Superior MAE/PSNR/SSIM compared to CycleGAN; SUV preservation | ADNI | ✗ Not comparable to Low-Dose PET Diffusion row—different dataset/task | Cross-modal synthesis (MRI→PET) | [30] |
| General/Multi-Modality (Foundational Methods) | |||||||||||
| Score-Based Generative Modeling (SDE) | General | N/A—natural images, not patients | Unconditional | Standard generative-modeling benchmark (non-clinical) | Theoretical foundation for diffusion models | Stochastic differential equations (SDE) unifying DDPM and NCSN frameworks | Unified framework; high-quality generation; stable training | CIFAR-10, CelebA Link—https://github.com/yang-song/score_sde (accessed on 1 August 2026) | ✗ NOT COMPARABLE to any medical imaging row—non-medical benchmark (CIFAR-10/CelebA) | General generative framework (non-medical benchmark) | [24] |
| Diffusion Posterior Sampling (DPS) | General | Not specified | Physics-guided forward-model consistency | Aggregated across multiple inverse problems; no single dataset | Noisy inverse problems (MRI, CT, etc.) | Physics-guided DDPM sampling with forward model consistency | Uncertainty quantification; outperforms supervised methods on several inverse problems | Various (MRI, CT, and deblurring) Link—https://github.com/DPS2022/diffusion-posterior-sampling (accessed on 1 August 2026) | ✗ Not comparable—dataset unspecified/aggregated across tasks | General inverse-problem solving (MRI/CT/deblurring) | [28] |
| Adversarial Diffusion Model for Unsupervised Translation | General | Not specified | Unpaired image to image | Qualitative structural-integrity preservation | Unpaired image-to-image translation (any modality) | Adversarial diffusion model combining DDPM with adversarial loss | Unsupervised translation without paired data; preserves structural integrity | MRI, CT, and fundus images | ✗ Not comparable—multi-modality aggregate, unpaired, no quantitative value | Unpaired cross-modality image translation | [32] |
| Dual-Task Consistency for Semi-Supervised Segmentation | General | Not specified | Shared representation (segmentation + reconstruction) | Segmentation accuracy with few labels | Limited labeled data for segmentation | Dual-task consistency (segmentation + reconstruction) with shared representation | Improved segmentation with few labels; not a generative model per se | Various medical images | ✗ Not comparable AND a scope concern—table notes this is “not a generative model per se” (see Comment 4) | Semi-supervised segmentation (reconstruction-assisted) | [33] |
| Ref | Dataset | Model/Method | Source Modality | Target Modality | Architecture | Key Metric/Result |
|---|---|---|---|---|---|---|
| [35] | Head & Neck (Jiangxi Cancer Hospital) | cycleSimulationGAN | MRI (T1/T2) | CT (HU map) | CycleGAN + channel-wise attention + structural similarity loss | MAE 52.3 HU; DSC 0.91 on bone; outperforms standard CycleGAN |
| [36] | Brain & Prostate (Emory University, institutional) | MC-IDDPM (3D Transformer Diffusion) | MRI (T1) | CT (HU map) | Swin-Vnet denoising diffusion probabilistic model (3D) | Brain: MAE 48.8 HU, SSIM 0.947, and NCC 0.976; Prostate: MAE 55.1 HU and SSIM 0.878 |
| [32] | IXI, BRATS (multi-contrast MRI); in-house MRI–CT | SynDiff | MRI (multi-contrast) | CT/PET (target) | Adversarial diffusion model with cycle-consistent architecture (unpaired) | Superior PSNR/SSIM/FID vs. CycleGAN and DDPM baselines on multi-contrast MRI and MRI–CT translation |
| [37] | ADNI (Alzheimer’s Disease Neuroimaging Initiative) | MRI-to-PET Diffusion with Pathology Awareness | MRI (T1) | FDG-PET | Conditional diffusion model with pathology-aware attention | Improved SUV correlation; better lesion-region synthesis vs. attention cGAN; evaluated for AD staging |
| [38] | Pelvic CT–MRI paired (institutional) | MD-DGA (Multi-Dimensional Diffusion Generation Architecture) | CT | MRI (T2) | 2D scalable diffusion model (2D-SDM) + 3D scalable latent diffusion model (3D-SLDM) with 3D-VQVAE | High-fidelity 3D MRI synthesis from paired CT volumes; preserves volumetric consistency |
| [39] | 5000 scans across 6 modalities (MRI T1/T2/FLAIR/PD, CT); multiple public datasets | SynthSeg | MRI (any contrast/resolution) | Synthetic segmentation label map | CNN trained on domain-randomized synthetic data from generative model conditioned on segmentations | DSC comparable to supervised CNNs across 5000 scans; 6 modalities; 10 resolutions; no retraining needed |
| [40] | Clinical PET (brain/whole-body) | PET Consistency Model (PET-CM) | Low-dose PET (1/8 or 1/4 dose) | Full-dose PET | Denoising diffusion with PET Shifted-window Vision Transformer (PET-VIT); consistency model | 1/8-dose: PSNR 33.9 dB, SSIM 0.964, and NCC 0.968; 12× faster inference than DDPM |
| Ref | Application Domain | Target Modality | Model Type | Key Benefit | Observed Outcome/Result |
|---|---|---|---|---|---|
| [41] | Data Augmentation—Fairness and Distribution Shift | Dermatology, chest X-ray, and histopathology | Latent diffusion model (LDM)—conditional generation | Improves classifier fairness for under-represented groups under distribution shift | Dermatology: 63.5% improvement in high-risk sensitivity; 7.5× reduction in fairness gap Chest radiology: 5.2% accuracy gain, 44.6% lower fairness gap |
| [42] | Data Augmentation—Pulmonary Nodule Detection | CT (chest radiograph) | DDPM (curriculum learning pipeline with synthetic nodule generation) | Improves detection sensitivity for small/low-contrast nodules; addresses class imbalance | AUC 0.95 vs. 0.89 baseline (p < 0.001); sensitivity 70% vs. 48% baseline; accuracy 82% vs. 70% |
| [43] | Radiology Education—AI-Assisted Graded Training | CT (pulmonary) | AI-assisted diagnosis system (detection + grading pipeline) | Structured graded training of residents using AI feedback on real clinical cases | AI-assisted groups (Groups 2 and 3) outperformed traditional teaching group in nodule sensitivity across all densities and sizes; confirmed over 7 rounds of testing on 1057 nodules |
| [44] | Cardiac Digital Twins—Electrophysiology at Scale | Cardiac MRI + ECG | Automated mesh generation pipeline + electromechanical simulation | Patient-specific ventricular models enabling electrophysiological insight at population scale | 3461 cardiac digital twins from UK Biobank; 359 from ischemic heart disease cohort; sex-specific QRS differences explained by anatomy; conduction velocity changes with age and obesity confirmed |
| [45] | Cardiac Digital Twins—CRT Treatment Planning | Cardiac MRI (LV/RV) | CircAdapt biomechanical model + imaging-based personalization | Virtual pacing prediction of left ventricular reverse remodeling after cardiac resynchronization therapy | 45 heart failure patients; direct correlation between virtual pacing response and actual LV reverse remodeling at follow-up; validated as patient selection tool for CRT |
| [46] | Cardiovascular Digital Twins—Generative AI Integration | Multi-modal (ECG, Echo, CT, and MRI) | Generative AI for counterfactual scenario simulation; digital twin population modeling | In silico clinical trial population modeling; risk stratification; treatment effect estimation | Framework for using digital twins to generate evidence across clinical trial populations; personalized simulation of cardiovascular scenarios; improved interpretability in HF prognostication |
| [47] | Federated Learning—Privacy-Preserving MRI Synthesis | MRI (knee—fat suppression) | GAN-based fat-suppressed MRI synthesis with federated training | Improves multi-site generalizability of synthesis model without sharing patient data | Federated GAN outperformed single-site trained GAN on external fastMRI data; demonstrated privacy-preserving multi-institutional synthesis collaboration |
| [48] | Federated Learning—Fairness via Synthetic Augmentation | Multi-modal (dermatology, radiology, and histopathology) | Conditional diffusion model with federated/privacy-preserving generation | Reduces model bias for under-represented populations without access to centralized patient data | Synthetic augmentation closed fairness gap across institutions; diffusion + FL enables scale-up without privacy violation |
| Metric | Full Name | Description | Primary Use Case | Known Limitations | Dimensionality/Computation |
|---|---|---|---|---|---|
| PSNR | Peak Signal-to-Noise Ratio | Log-ratio (dB) of the maximum possible pixel value to the mean squared error between generated and reference volumes, computed voxel-wise. | CT/MRI reconstruction quality; low-dose CT denoising; accelerated MRI benchmarking. | No perceptual or structural sensitivity; favors over-smoothed outputs; insensitive to clinically relevant local distortions. Sensitive to global intensity offsets. | Computed per 2D slice in most reviewed studies, then averaged across slices; native whole-volume (3D) PSNR is possible but rarely reported explicitly. |
| SSIM | Structural Similarity Index Measure | Local window-based index combining luminance, contrast, and structural comparison between two images; values in [−1, 1], higher is better. | Accelerated MRI reconstruction; anatomy-preserving synthesis; cross-modal image translation evaluation. | Window-size sensitive; poor at capturing global geometry; high SSIM images can still fail clinical assessments; poor correlation with expert radiologist scores in several studies. | Typically computed with a 2D sliding window per slice, then averaged across slices; a volumetric (3D) sliding-window SSIM exists but is used less often in the reviewed literature. |
| NRMSE | Normalized Root Mean Squared Error | Root mean squared voxel-wise intensity deviation normalized by the range or mean of the reference volume; lower is better. | Quantitative MRI reconstruction; k-space recovery evaluation; HU-accuracy in CT synthesis. | Sensitive to outlier voxels; dominated by high-intensity structures (bone in CT, fat in MRI); not invariant to global intensity scale; no spatial structural information. | Computed voxel-wise over the full 3D volume directly—inherently volume-native since it is a simple normalized error rather than a windowed statistic. |
| MAE/MSE | Mean Absolute/Mean Squared Error | Average absolute or squared voxel intensity difference between generated and reference volumes; both are lower-is-better scalar error summaries. | Denoising; SUV accuracy in PET synthesis; HU accuracy in synthetic CT for radiotherapy planning. | Entirely insensitive to spatial structure and anatomical coherence; MSE penalises large errors disproportionately; neither reflects perceptual or clinical image quality. | Voxel-wise and natively 3D by construction; any slice-wise reporting is a display convenience and does not change the underlying computation. |
| LPIPS | Learned Perceptual Image Patch Similarity | Feature-space distance between image patches using activations from a deep CNN (AlexNet or VGG) trained on natural images; lower is better. | Perceptual quality assessment of MRI/CT synthesis; diffusion model output evaluation; complements SSIM/PSNR in paired evaluation. | Pretrained on RGB natural images—limited validation for grayscale volumetric medical data; 3D extension not standardized; sensitive to domain shift between natural and medical image distributions. | Inherently 2D: computed slice by slice with a 2D CNN backbone (AlexNet/VGG) and averaged across slices; no native 3D LPIPS backbone is used in the reviewed literature (Section 7.2). |
| FID/FID-3D | Fréchet Inception Distance (3D adapted) | Fréchet distance between Gaussian distributions fitted to deep feature embeddings of real vs. generated volume sets; lower indicates closer distributional alignment. | Diversity and realism assessment of unconditional 3D generation; GAN/diffusion model comparison on brain MRI and chest CT. | Requires large sample sizes (≥2000 recommended); Inception-v3 trained on ImageNet—applicability to medical images debated; no standardized 3D feature extractor; FID-optimal models can still produce clinically incorrect images. | Standard FID uses 2D slices through a 2D Inception network; FID-3D uses a 3D feature extractor on the whole volume, but the extractor is not standardized across studies (Section 7.2). |
| Dice/IoU | Dice Similarity Coefficient/Intersection over Union | Overlap-based metrics applied to binary segmentation masks of anatomical structures in synthetic vs. real volumes; Dice = 2|A∩B|/(|A| + |B|). | Downstream segmentation task evaluation; assessing whether synthetic data augmentation improves segmentation model performance. | Requires segmentation ground truth; cannot directly assess synthesis quality; sensitive to class imbalance; insensitive to boundary sharpness. | Computed natively in 3D over the full segmentation volume in most reviewed studies, though 2D per-slice Dice is sometimes reported separately for slice-wise comparison. |
| HD95 | Hausdorff Distance (95th percentile) | 95th percentile of the maximum one-sided surface distance between predicted and reference segmentation contours; reported in mm; lower is better. Uses the 95th percentile to reduce sensitivity to extreme outliers present in the full HD. | Organ boundary fidelity in synthetic CT for radiotherapy; shape accuracy in surgical planning evaluation. | Still influenced by surface outliers beyond the 95th percentile; implementation-dependent (voxel spacing, mesh vs. voxel computation); does not encode volumetric overlap; requires segmentation masks. | Computed over the full 3D surface point set in most reviewed studies; some 2D contour-based variants exist for slice-wise comparison. |
| ASSD | Average Symmetric Surface Distance | Mean of all bidirectional nearest-surface distances between two segmentation contours; symmetric and lower is better; complements HD95 by averaging rather than taking an extreme value. | Organ shape fidelity assessment; complementary to HD95 in synthetic CT and MRI evaluation pipelines. | Can mask localized large errors by averaging; also requires segmentation masks; sensitive to the segmentation method used for evaluation. | Computed over the full 3D surface mesh/point set in most reviewed studies. |
| Ref | Dataset Name | Size | Modality | Anatomy/Task | Access/Homepage |
|---|---|---|---|---|---|
| [49] | BraTS 2021 | 1251 training/5 validation/219 testing volumes | MRI (T1, T1ce, T2, and FLAIR) | Brain tumor segmentation and multi-modal MRI synthesis | www.med.upenn.edu/cbica/brats2021/ (accessed on 3 August 2026) |
| [50] | BraSyn 2023 | Based on BraTS 2021—challenge-defined splits | MRI (T1, T1ce, T2, and FLAIR) | Brain MR image synthesis for missing modalities | www.med.upenn.edu/cbica/brats/ (BraTS challenge series, Synapse portal) (accessed on 3 August 2026) |
| [51] | LIDC-IDRI | 1018 cases; 4-radiologist annotations; 7 academic centers + 8 companies | CT (thoracic) | Lung nodule detection and segmentation and CT synthesis | www.cancerimagingarchive.net/collection/lidc-idri/ (accessed on 3 August 2026) |
| [52] | fastMRI | ~8344 raw k-space volumes (1594 knee + 6970 brain); DICOM set adds 20,000+ | MRI (knee and brain) | Accelerated MRI reconstruction and k-space synthesis | https://fastmri.med.nyu.edu/ (accessed on 3 August 2026) |
| [53] | CHAOS | 80 patients: 40 CT + 40 MRI; 120 MRI DICOM series | CT + MRI (T1-DUAL in/opp-phase, T2-SPIR) | Healthy abdominal organ segmentation; CT↔MRI translation | https://chaos.grand-challenge.org/ (accessed on 3 August 2026) |
| [54] | ACDC | 150 patients; 100 training/50 testing; 5 diagnostic classes | Cardiac cine MRI | Cardiac segmentation (LV, RV, and myocardium) and diagnosis | www.creatis.insa-lyon.fr/Challenge/acdc/ (accessed on 3 August 2026) |
| [55] | ADNI (1–4) | >2000 subjects cumulative; longitudinal, multi-timepoint | MRI (T1 and 3T) + PET (FDG, amyloid, and tau) | Alzheimer’s disease progression; PET synthesis from MRI | https://adni.loni.usc.edu/ (accessed on 3 August 2026) |
| [56] | AAPM Low-Dose CT | 299 scans: 49 head, 100 chest, 100 abdomen; full-dose + 25%/10% low-dose pairs | CT (head, chest, and abdomen) | Low-dose CT denoising and dose-reduction synthesis | www.aapm.org/GrandChallenge/LowDoseCT/ (accessed on 3 August 2026) |
| [57] | HECKTOR 2022 | 883 cases: 524 training (7 centers) + 359 test (3 centers) | FDG-PET + CT | Head and neck tumor segmentation; multi-modal PET/CT synthesis | https://hecktor.grand-challenge.org/ (accessed on 3 August 2026) |
| [58] | crossMoDA 2021–23 | 2021: 227 ceT1 + 295 hrT2; 2023: ~560 training subjects | MRI (ceT1 → hrT2) | Vestibular schwannoma and cochlea segmentation; unpaired cross-modal synthesis | https://crossmoda-challenge.ml/ (accessed on 3 August 2026) |
| Stage | What Happens | Minimum Evidence—Research Only | Minimum Evidence—Augmentation/Education | Minimum Evidence—Clinical Decision Support |
|---|---|---|---|---|
| 1. Data governance | Define data sources, consent scope, de-identification, and retention policy | Documented data-use agreement | Explicit consent for generative reuse; documented provenance | Explicit consent for generative reuse; IRB/ethics approval; documented chain of custody |
| 2. Model development | Architecture selection, training, and hyperparameter search | Version-controlled code | Version-controlled code; training-data manifest | Version-controlled code; training-data manifest; predetermined change-control plan |
| 3. Technical validation | Voxel/structural fidelity metrics on held-out internal data | PSNR/SSIM/FID on internal split | PSNR/SSIM/FID plus downstream task metric (e.g., Dice) | Full Table 5 metric panel plus uncertainty quantification |
| 4. External validation | Evaluation on data from sites/scanners not seen in training | Not required | At least one external dataset or site | Multi-site, multi-vendor external validation with stratified reporting (Section 8.5) |
| 5. Clinical evaluation | Expert reader study/clinical-utility assessment | Not required | Reader study on downstream task performance | Blinded multi-reader clinical-equivalence or non-inferiority study |
| 6. Risk assessment | Formal identification of failure modes and their clinical consequence | Informal | Documented failure-mode list | Formal risk file mapped to regulatory risk class (Section 8.5) |
| 7. Human review | Definition of the human-in-the-loop checkpoint before use | N/A | Recommended at point of downstream model deployment | Mandatory clinician sign-off before use in patient care |
| 8. Deployment | Integration into clinical/research workflow | N/A | Institutional approval | Regulatory clearance/CE marking or equivalent; integrity-checked model deployment (Section 8.5) |
| 9. Continuous monitoring | Post-deployment surveillance of performance and safety | N/A | Periodic re-evaluation against Stage 3 metrics | Continuous drift monitoring, override/rejection tracking, and adverse-event reporting (Section 8.5) |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Kumar, C.K.; Kumar, M.K.; Gurrala, V.; Grandhi, A.; Chikkala, R.B.; Praveen, S.P. From Pixels to Volumes: Generative AI in 3D Medical Imaging. Math. Comput. Appl. 2026, 31, 196. https://doi.org/10.3390/mca31050196
Kumar CK, Kumar MK, Gurrala V, Grandhi A, Chikkala RB, Praveen SP. From Pixels to Volumes: Generative AI in 3D Medical Imaging. Mathematical and Computational Applications. 2026; 31(5):196. https://doi.org/10.3390/mca31050196
Chicago/Turabian StyleKumar, Chanumolu Kiran, Maheswara Kishore Kumar, Venkataramana Gurrala, Appalaraju Grandhi, Rajendra Babu Chikkala, and Surapaneni Phani Praveen. 2026. "From Pixels to Volumes: Generative AI in 3D Medical Imaging" Mathematical and Computational Applications 31, no. 5: 196. https://doi.org/10.3390/mca31050196
APA StyleKumar, C. K., Kumar, M. K., Gurrala, V., Grandhi, A., Chikkala, R. B., & Praveen, S. P. (2026). From Pixels to Volumes: Generative AI in 3D Medical Imaging. Mathematical and Computational Applications, 31(5), 196. https://doi.org/10.3390/mca31050196

