Diff-GTISR: Guided Thermal Image Super-Resolution via Diffusion Model and Refinement
Abstract
1. Introduction
- (1)
- Diff-GTISR, a diffusion model for guided thermal image super-resolution, which simultaneously improves the image quality subjectively and objectively.
- (2)
- Design of a modality-specific dual encoder and cross-modal guidance attention module, effectively injecting visible light structural details into thermal image restoration.
- (3)
- Utilization and analysis of an optional refiner to suppress unrealistic patterns that may appear in the resulting images.
2. Related Works
2.1. Single Image Super-Resolution
2.2. Guided Thermal Image Super-Resolution
2.3. Diffusion Model for Image Super-Resolution
3. Proposed Method
3.1. Modality-Specific Dual Encoder
3.2. Cross-Modal Guidance Attention Module
3.3. Decoder and Loss Function
3.4. Refiner Network
4. Experiment Results
4.1. Dataset
4.2. Implementation Details
4.3. Evaluation Metrics
4.4. Compared Methods
4.5. Objective Comparison
4.6. Subjective Comparison
4.7. Ablation Study
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Bagavathiappan, S.; Lahiri, B.B.; Saravanan, T.; Philip, J.; Jayakumar, T. Infrared thermography for condition monitoring—A review. Infrared Phys. Technol. 2013, 60, 35–55. [Google Scholar] [CrossRef] [Scilit]
- Gade, R.; Moeslund, T.B. Thermal cameras and applications: A survey. Mach. Vis. Appl. 2014, 25, 245–262. [Google Scholar] [CrossRef] [Scilit]
- Ma, J.; Ma, Y.; Li, C. Infrared and visible image fusion methods and applications: A survey. Inf. Fusion 2019, 45, 153–178. [Google Scholar] [CrossRef] [Scilit]
- Altay, F.; Velipasalar, S. The use of thermal cameras for pedestrian detection. IEEE Sens. J. 2022, 22, 11489–11498. [Google Scholar] [CrossRef] [Scilit]
- Dong, C.; Loy, C.C.; He, K.; Tang, X. Image super-resolution using deep convolutional networks. IEEE Trans. Pattern Anal. Mach. Intell. 2016, 38, 295–307. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Dong, C.; Loy, C.C.; Tang, X. Accelerating the super-resolution convolutional neural network. In Computer Vision—ECCV 2016; European Conference on Computer Vision (ECCV); Springer: Cham, Switzerland, 2016; Volume 9906, pp. 391–407. [Google Scholar] [CrossRef] [Scilit]
- Liang, J.; Cao, J.; Sun, G.; Zhang, K.; Van Gool, L.; Timofte, R. SwinIR: Image restoration using Swin Transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), Montreal, BC, Canada, 11–17 October 2021; pp. 1833–1844. [Google Scholar] [CrossRef] [Scilit]
- Chen, X.; Wang, X.; Zhou, J.; Qiao, Y.; Dong, C. Activating more pixels in image super-resolution transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 17–24 June 2023; pp. 22367–22377. [Google Scholar] [CrossRef] [Scilit]
- Ledig, C.; Theis, L.; Huszár, F.; Caballero, J.; Cunningham, A.; Acosta, A.; Aitken, A.; Tejani, A.; Totz, J.; Wang, Z.; et al. Photo-realistic single image super-resolution using a generative adversarial network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 4681–4690. [Google Scholar] [CrossRef] [Scilit]
- Wang, X.; Yu, K.; Wu, S.; Gu, J.; Liu, Y.; Dong, C.; Qiao, Y.; Loy, C.C. ESRGAN: Enhanced super-resolution generative adversarial networks. In Computer Vision—ECCV 2018 Workshops; European Conference on Computer Vision Workshops (ECCVW); Springer: Cham, Switzerland, 2018; Volume 11133, pp. 63–79. [Google Scholar] [CrossRef] [Scilit]
- Zou, Y.; Zhang, L.; Liu, C.; Wang, B.; Hu, Y.; Chen, Q. Super-resolution reconstruction of infrared images based on a convolutional neural network with skip connections. Opt. Lasers Eng. 2021, 146, 106717. [Google Scholar] [CrossRef] [Scilit]
- Rivadeneira, R.E.; Sappa, A.D.; Vintimilla, B.X. Thermal image super-resolution: A novel architecture and dataset. In Proceedings of the International Conference on Computer Vision Theory and Applications (VISAPP), Valletta, Malta, 27–29 February 2020; pp. 111–119. [Google Scholar] [CrossRef] [Scilit]
- Huang, Y.; Jiang, Z.; Lan, R.; Zhang, S.; Pi, K. Infrared image super-resolution via transfer learning and PSRGAN. IEEE Signal Process. Lett. 2021, 28, 982–986. [Google Scholar] [CrossRef] [Scilit]
- Wu, X.; Zhou, B.; Wang, X.; Peng, J.; Lin, P.; Cao, R.; Huang, F. SwinIPISR: A super-resolution method for infrared polarization imaging sensors via Swin Transformer. IEEE Sens. J. 2024, 24, 468–477. [Google Scholar] [CrossRef] [Scilit]
- Rivadeneira, R.E.; Sappa, A.D.; Wang, C.; Jiang, J.; Zhong, Z.; Chen, P.; Wang, S. Thermal image super-resolution challenge results—PBVS 2024. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Seattle, WA, USA, 17–18 June 2024; pp. 3113–3122. [Google Scholar] [CrossRef] [Scilit]
- Almasri, F.; Debeir, O. Multimodal sensor fusion in single thermal image super-resolution. In Computer Vision—ECCV 2018 Workshops; Asian Conference on Computer Vision (ACCV); Springer: Cham, Switzerland, 2018; Volume 11367, pp. 418–433. [Google Scholar] [CrossRef] [Scilit]
- Gupta, H.; Mitra, K. Pyramidal edge-maps and attention based guided thermal super-resolution. In Computer Vision—ECCV 2020 Workshops; Bartoli, A., Fusiello, A., Eds.; Springer: Cham, Switzerland, 2021; Volume 12537, pp. 698–715. [Google Scholar] [CrossRef] [Scilit]
- Fang, Y.; Fan, L.; Cai, Y. Guided super-resolution for image fusion: A novel approach to enhancing crack segmentation in masonry structures. IEEE Sens. J. 2025, 25, 11491–11507. [Google Scholar] [CrossRef] [Scilit]
- Jiang, H.; Chen, Z. Flexible window-based self-attention transformer in thermal image super-resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Seattle, WA, USA, 17–18 June 2024; pp. 3076–3085. [Google Scholar] [CrossRef] [Scilit]
- Arnold, C.; Jouvet, P.; Seoud, L. SwinFuSR: An image fusion-inspired model for RGB-guided thermal image super-resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Seattle, WA, USA, 17–18 June 2024; pp. 3027–3036. [Google Scholar] [CrossRef] [Scilit]
- Shi, L.; Wu, G.; Wang, Y.; Liu, Y.; Chai, T. DuaDiff: Dual-conditional diffusion model for guided thermal image super-resolution. IEEE Trans. Neural Netw. Learn. Syst. 2025; in press. [CrossRef] [Scilit] [PubMed]
- Cortés-Mendez, C.; Hayet, J.-B. Exploring the usage of diffusion models for thermal image super-resolution: A generic, uncertainty-aware approach for guided and non-guided schemes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Seattle, WA, USA, 17–18 June 2024; pp. 3123–3130. [Google Scholar] [CrossRef] [Scilit]
- Saharia, C.; Ho, J.; Chan, W.; Salimans, T.; Fleet, D.J.; Norouzi, M. Image super-resolution via iterative refinement. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 45, 4713–4726. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Li, H.; Yang, Y.; Chang, M.; Chen, S.; Feng, H.; Xu, Z.; Li, Q.; Chen, Y. SRDiff: Single image super-resolution with diffusion probabilistic models. Neurocomputing 2022, 479, 47–59. [Google Scholar] [CrossRef] [Scilit]
- Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 10674–10685. [Google Scholar] [CrossRef] [Scilit]
- Yue, Z.; Wang, J.; Loy, C.C. ResShift: Efficient diffusion model for image super-resolution by residual shifting. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2023; Volume 36, pp. 13294–13307. [Google Scholar] [CrossRef] [Scilit]
- Li, X.; Ren, Y.; Jin, X.; Lan, C.; Wang, X.; Zeng, W.; Wang, X.; Chen, Z. Diffusion models for image restoration and enhancement: A comprehensive survey. Int. J. Comput. Vis. 2025, 133, 8078–8108. [Google Scholar] [CrossRef] [Scilit]
- Cohen, R.; Kligvasser, I.; Rivlin, E.; Freedman, D. Looks too good to be true: An information-theoretic analysis of hallucinations in generative restoration models. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2024; Volume 37, pp. 22596–22623. [Google Scholar] [CrossRef] [Scilit]
- Li, X.; Orchard, M.T. New edge-directed interpolation. IEEE Trans. Image Process. 2001, 10, 1521–1527. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Choi, B.D.; Yoo, H. Design of piecewise weighted linear interpolation based on even-odd decomposition and its application to image resizing. IEEE Trans. Consum. Electron. 2009, 55, 2280–2286. [Google Scholar] [CrossRef] [Scilit]
- Yang, J.; Wright, J.; Huang, T.S.; Ma, Y. Image super-resolution via sparse representation. IEEE Trans. Image Process. 2010, 19, 2861–2873. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Timofte, R.; De Smet, V.; Van Gool, L. Anchored neighborhood regression for fast example-based super-resolution. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Sydney, Australia, 1–8 December 2013; pp. 1920–1927. [Google Scholar] [CrossRef] [Scilit]
- Chen, X.; Zhai, G.; Wang, J.; Hu, C.; Chen, Y. Color guided thermal image super resolution. In Proceedings of the 2016 Visual Communications and Image Processing (VCIP), Chengdu, China, 27–30 November 2016; pp. 1–4. [Google Scholar] [CrossRef] [Scilit]
- Ho, J.; Jain, A.; Abbeel, P. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2020; Volume 33, pp. 6840–6851. [Google Scholar]
- Esser, P.; Rombach, R.; Ommer, B. Taming transformers for high-resolution image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021; pp. 12868–12878. [Google Scholar] [CrossRef] [Scilit]
- Li, X.; Wang, Z.; Zou, Y.; Chen, Z.; Ma, J.; Jiang, Z.; Ma, L.; Liu, J. DifIISR: A diffusion model with gradient guidance for infrared image super-resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 10–17 June 2025; pp. 7534–7544. [Google Scholar] [CrossRef] [Scilit]
- Weng, Z.; Liu, X.; Liu, C.; Guo, X.; Shi, Y.; Lin, L. DroneSR: Rethinking few-shot thermal image super-resolution from drone-based perspective. IEEE Sens. J. 2025, 25, 37722–37731. [Google Scholar] [CrossRef] [Scilit]
- Blau, Y.; Michaeli, T. The perception-distortion trade-off. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–23 June 2018; pp. 6228–6237. [Google Scholar] [CrossRef] [Scilit]
- Rivadeneira, R.E.; Velesaca, H.O.; Sappa, A. Cross-spectral image registration: A comparative study and a new benchmark dataset. In Innovations in Computational Intelligence and Computer Vision; International Conference on Innovative Computing Intelligence and Computer Vision; Springer: Singapore, 2024; pp. 1–12. [Google Scholar] [CrossRef] [Scilit]
- Wang, Z.; Bovik, A.C.; Sheikh, H.R.; Simoncelli, E.P. Image quality assessment: From error visibility to structural similarity. IEEE Trans. Image Process. 2004, 13, 600–612. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhang, R.; Isola, P.; Efros, A.A.; Shechtman, E.; Wang, O. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–23 June 2018; pp. 586–595. [Google Scholar] [CrossRef] [Scilit]








| Method | Task | Model | Parameters |
|---|---|---|---|
| ResShift [26] | SISR | Diffusion | 118.59 M |
| SwinIR [7] | SISR | Transformer | 12.04 M |
| DiffGUA [22] | GTISR | Diffusion | 123.7 M |
| SwinFuSR [20] | GTISR | Transformer | 0.94 M |
| Method | PSNR [dB]↑ | SSIM↑ | LPIPS↓ |
|---|---|---|---|
| Bicubic | 24.99 | 0.7650 | 0.4876 |
| ResShift [26] | 25.83 | 0.7999 | 0.2113 |
| SwinIR [7] | 26.79 | 0.8297 | 0.2648 |
| DiffGUA [22] | 26.97 | 0.8189 | 0.2405 |
| SwinFuSR [20] | 28.18 | 0.8633 | 0.2010 |
| Diff-GTISR (Ours) | 28.23 | 0.8613 | 0.1753 |
| Method | PSNR [dB]↑ | SSIM↑ | LPIPS↓ |
|---|---|---|---|
| Bicubic | 21.86 | 0.6894 | 0.6569 |
| ResShift [26] | 22.29 | 0.6962 | 0.3569 |
| SwinIR [7] | 22.92 | 0.7279 | 0.4953 |
| DiffGUA [22] | 24.16 | 0.7515 | 0.3273 |
| SwinFuSR [20] | 24.73 | 0.7855 | 0.3129 |
| Diff-GTISR (Ours) | 25.34 | 0.7992 | 0.2491 |
| Structure | Attention | Q/K/V | Connection | PSNR [dB]↑ | SSIM↑ | LPIPS↓ |
|---|---|---|---|---|---|---|
| A (baseline) | - | - | - | 25.83 | 0.7999 | 0.2113 |
| B | CA | thr/vis/vis | residual | 26.20 | 0.8073 | 0.1950 |
| C | CA | thr/vis/vis | concat | 26.53 | 0.8158 | 0.1956 |
| D | CA | vis/thr/thr | concat | 27.64 | 0.8448 | 0.1565 |
| E (Ours) | CMGA | vis/vis/thr | concat | 27.64 | 0.8484 | 0.1520 |
| Structure | PSNR [dB]↑ | SSIM↑ | LPIPS↓ |
|---|---|---|---|
| w/o Refiner | 27.64 | 0.8484 | 0.1520 |
| with Refiner | 28.23 | 0.8613 | 0.1753 |
| Structure | PSNR [dB]↑ | SSIM↑ | LPIPS↓ |
|---|---|---|---|
| w/o Refiner | 24.95 | 0.7869 | 0.2116 |
| with Refiner | 25.34 | 0.7992 | 0.2491 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Hong, C.; Yoo, H. Diff-GTISR: Guided Thermal Image Super-Resolution via Diffusion Model and Refinement. Appl. Sci. 2026, 16, 3435. https://doi.org/10.3390/app16073435
Hong C, Yoo H. Diff-GTISR: Guided Thermal Image Super-Resolution via Diffusion Model and Refinement. Applied Sciences. 2026; 16(7):3435. https://doi.org/10.3390/app16073435
Chicago/Turabian StyleHong, ChaeHui, and Hoon Yoo. 2026. "Diff-GTISR: Guided Thermal Image Super-Resolution via Diffusion Model and Refinement" Applied Sciences 16, no. 7: 3435. https://doi.org/10.3390/app16073435
APA StyleHong, C., & Yoo, H. (2026). Diff-GTISR: Guided Thermal Image Super-Resolution via Diffusion Model and Refinement. Applied Sciences, 16(7), 3435. https://doi.org/10.3390/app16073435
