CSCGAN: Cross-Space Contrastive Learning for Blind Image Inpainting
Abstract
1. Introduction
- A mask-prediction-free framework is proposed for blind inpainting. Our method does not need to predict the noise area, thus avoiding inferior inpainting results caused by inaccurate mask prediction.
- Two novel constraints named cross-space contrastive loss and mask-aware adversarial loss are proposed for blind inpainting. To the best of our knowledge, we are the first to introduce contrastive learning to blind inpainting task.
- We propose the first blind inpainting metric called GCM, which can reasonably measure the quality of blind inpainting.
- Extensive experiments on benchmark datasets prove that our method achieves better quality than the state-of-the-art methods.
2. Related Work
2.1. Image Inpainting
2.2. Blind Image Inpainting
2.3. Contrastive Learning
3. Proposed Method
3.1. Degraded Image Generation
3.2. Network Architecture
3.3. Loss Function
3.3.1. Cross-Space Contrastive Loss (CSC Loss)
3.3.2. Mask-Aware Adversarial Loss (MAA Loss)
3.3.3. Other Inpainting Losses
3.3.4. Refine Losses
4. Experimental Results
4.1. Experimental Settings
4.2. Qualitative Comparison
4.2.1. Compared with General Inpainting Methods
4.2.2. Compared with Blind Inpainting Methods
4.3. Quantitative Comparison
4.4. Ablation Studies
4.4.1. W/ and W/O CSC Loss and MAA Loss
4.4.2. CSC Loss vs. Latent L1 Loss
4.4.3. Preservation of Reasonable Areas
4.4.4. Complexity and Computational Cost
4.4.5. Other Applications
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Yu, T.; Guo, Z.; Jin, X.; Wu, S.; Chen, Z.; Li, W.; Zhang, Z.; Liu, S. Region normalization for image inpainting. Proc. AAAI Conf. Artif. Intell. 2020, 34, 12733–12740. [Google Scholar] [CrossRef] [Scilit]
- Yu, J.; Lin, Z.; Yang, J.; Shen, X.; Lu, X.; Huang, T.S. Generative image inpainting with contextual attention. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; IEEE: Piscataway, NJ, USA, 2018; pp. 5505–5514. [Google Scholar]
- Wang, Y.; Tao, X.; Qi, X.; Shen, X.; Jia, J. Image inpainting via generative multi-column convolutional neural networks. arXiv 2018, arXiv:1810.08771. [Google Scholar] [CrossRef] [Scilit]
- Nazeri, K.; Ng, E.; Joseph, T.; Qureshi, F.Z.; Ebrahimi, M. EdgeConnect: Generative Image Inpainting with Adversarial Edge Learning. arXiv 2019, arXiv:1901.00212. [Google Scholar] [CrossRef] [Scilit]
- Suvorov, R.; Logacheva, E.; Mashikhin, A.; Remizova, A.; Ashukha, A.; Silvestrov, A.; Kong, N.; Goka, H.; Park, K.; Lempitsky, V. Resolution-robust large mask inpainting with fourier convolutions. In Proceedings of the 2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA, 3–8 January 2022; IEEE: Piscataway, NJ, USA, 2022; pp. 2149–2159. [Google Scholar]
- Zhao, S.; Cui, J.; Sheng, Y.; Dong, Y.; Liang, X.; Chang, E.I.; Xu, Y. Large scale image completion via co-modulated generative adversarial networks. arXiv 2021, arXiv:2103.10428. [Google Scholar] [CrossRef] [Scilit]
- Liu, Y.; Pan, J.; Su, Z. Deep blind image inpainting. In Intelligence Science and Big Data Engineering, Proceedings of the 9th International Conference, IScIDE 2019, Nanjing, China, 17–20 October 2019; Springer: Cham, Switzerland, 2019; pp. 128–141. [Google Scholar]
- Chen, H.; Giuffrida, M.V.; Doerner, P.; Tsaftaris, S.A. Blind Inpainting of Large-scale Masks of Thin Structures with Adversarial and Reinforcement Learning. arXiv 2019, arXiv:1912.02470. [Google Scholar] [CrossRef] [Scilit]
- Wang, J.; Chen, S.; Wu, Z.; Jiang, Y.G. FT-TDR: Frequency-guided Transformer and Top-Down Refinement Network for Blind Face Inpainting. IEEE Trans. Multimed. 2022, 25, 2382–2392. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Chen, Y.C.; Tao, X.; Jia, J. VCNet: A Robust Approach to Blind Image Inpainting. In Computer Vision—ECCV 2020, Proceedings of the 16th European Conference, Glasgow, UK, 23–28 August 2020; Springer: Cham, Switzerland, 2020. [Google Scholar]
- Chen, T.; Kornblith, S.; Norouzi, M.; Hinton, G. A simple framework for contrastive learning of visual representations. In Proceedings of the 37th International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2020; pp. 1597–1607. [Google Scholar]
- Zbontar, J.; Jing, L.; Misra, I.; LeCun, Y.; Deny, S. Barlow twins: Self-supervised learning via redundancy reduction. arXiv 2021, arXiv:2103.03230. [Google Scholar] [CrossRef] [Scilit]
- Pathak, D.; Krahenbuhl, P.; Donahue, J.; Darrell, T.; Efros, A.A. Context encoders: Feature learning by inpainting. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; IEEE: Piscataway, NJ, USA, 2016; pp. 2536–2544. [Google Scholar]
- Yang, J.; Qi, Z.; Shi, Y. Learning to incorporate structure knowledge for image inpainting. Proc. AAAI Conf. Artif. Intell. 2020, 34, 12605–12612. [Google Scholar] [CrossRef] [Scilit]
- Wang, T.; Ouyang, H.; Chen, Q. Image Inpainting with External-internal Learning and Monochromic Bottleneck. arXiv 2021, arXiv:2104.09068. [Google Scholar] [CrossRef] [Scilit]
- Ko, K.; Kim, C.S. Continuously masked transformer for image inpainting. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 1–6 October 2023; IEEE: Piscataway, NJ, USA, 2023; pp. 13169–13178. [Google Scholar]
- Li, Z.; Zhang, Y.; Du, Y.; Wang, X.; Wen, C.; Zhang, Y.; Geng, G.; Jia, F. STNet: Structure and texture-guided network for image inpainting. Pattern Recognit. 2024, 156, 110786. [Google Scholar] [CrossRef] [Scilit]
- Li, Z.; Han, N.; Wang, Y.; Zhang, Y.; Yan, J.; Du, Y.; Geng, G. Image inpainting based on CNN-Transformer framework via structure and texture restoration. Appl. Soft Comput. 2025, 170, 112671. [Google Scholar] [CrossRef] [Scilit]
- Zhang, S.; Liu, Y.; Jin, L.; Huang, Y.; Lai, S. Ensnet: Ensconce text in the wild. Proc. AAAI Conf. Artif. Intell. 2019, 33, 801–808. [Google Scholar] [CrossRef] [Scilit]
- Wang, X.; Li, Y.; Zhang, H.; Shan, Y. Towards real-world blind face restoration with generative facial prior. In Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021; IEEE: Piscataway, NJ, USA, 2021; pp. 9168–9178. [Google Scholar]
- Zhao, H.; Wang, Y.; Gu, Z.; Zheng, B.; Zheng, H. Context-aware mutual learning for blind image inpainting and beyond. Expert Syst. Appl. 2025, 268, 126224. [Google Scholar] [CrossRef] [Scilit]
- Chihaoui, H.; Lemkhenter, A.; Favaro, P. Blind image restoration via fast diffusion inversion. Adv. Neural Inf. Process. Syst. 2024, 37, 34513–34532. [Google Scholar]
- Meng, J.; Liu, W.; Shi, C.; Li, Z.; Liu, J. Self-information and prediction mask enhanced blind inpainting network for dunhuang murals. Eng. Appl. Artif. Intell. 2025, 159, 111769. [Google Scholar] [CrossRef] [Scilit]
- He, K.; Fan, H.; Wu, Y.; Xie, S.; Girshick, R. Momentum contrast for unsupervised visual representation learning. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; IEEE: Piscataway, NJ, USA, 2020; pp. 9729–9738. [Google Scholar]
- Henaff, O. Data-efficient image recognition with contrastive predictive coding. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2020; pp. 4182–4192. [Google Scholar]
- Hjelm, R.D.; Fedorov, A.; Lavoie-Marchildon, S.; Grewal, K.; Bachman, P.; Trischler, A.; Bengio, Y. Learning deep representations by mutual information estimation and maximization. arXiv 2018, arXiv:1808.06670. [Google Scholar]
- Hadsell, R.; Chopra, S.; LeCun, Y. Dimensionality reduction by learning an invariant mapping. In Proceedings of the 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’06), New York, NY, USA, 17–22 June 2006; IEEE: Piscataway, NJ, USA, 2006; Volume 2, pp. 1735–1742. [Google Scholar]
- Dosovitskiy, A.; Springenberg, J.T.; Riedmiller, M.; Brox, T. Discriminative unsupervised feature learning with convolutional neural networks. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2014. [Google Scholar]
- Park, T.; Efros, A.A.; Zhang, R.; Zhu, J.Y. Contrastive learning for unpaired image-to-image translation. In Computer Vision—ECCV 2020, Proceedings of the 16th European Conference, Glasgow, UK, 23–28 August 2020; Springer: Springer: Cham, Switzerland, 2020; pp. 319–345. [Google Scholar]
- Chopra, S.; Hadsell, R.; LeCun, Y. Learning a similarity metric discriminatively, with application to face verification. In Proceedings of the 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05), San Diego, CA, USA, 20–25 June 2005; IEEE: Piscataway, NJ, USA, 2005; Volume 1, pp. 539–546. [Google Scholar]
- Wu, G.; Jiang, J.; Jiang, K.; Liu, X. Learning from history: Task-agnostic model contrastive learning for image restoration. Proc. AAAI Conf. Artif. Intell. 2024, 38, 5976–5984. [Google Scholar] [CrossRef] [Scilit]
- Wang, Z.; Peng, J.; Zhao, H.; Yao, L.; Zhao, K. Orthogonal Decoupling Contrastive Regularization: Towards Uncorrelated Feature Decoupling for Unpaired Image Restoration. IEEE Trans. Pattern Anal. Mach. Intell. 2025, 48, 1842–1859. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lee, C.H.; Liu, Z.; Wu, L.; Luo, P. MaskGAN: Towards Diverse and Interactive Facial Image Manipulation. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; IEEE: Piscataway, NJ, USA, 2020. [Google Scholar]
- Zhou, B.; Lapedriza, A.; Khosla, A.; Oliva, A.; Torralba, A. Places: A 10 million image database for scene recognition. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 40, 1452–1464. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Doersch, C.; Singh, S.; Gupta, A.; Sivic, J.; Efros, A. What makes Paris look like Paris? ACM Trans. Graph. 2012, 31, 101. [Google Scholar] [CrossRef]
- Liu, G.; Reda, F.A.; Shih, K.J.; Wang, T.C.; Tao, A.; Catanzaro, B. Image inpainting for irregular holes using partial convolutions. In Computer Vision—ECCV 2018, Proceedings of the 15th European Conference, Munich, Germany, 8–14 September 2018; Springer: Cham, Switzerland, 2018; pp. 85–100. [Google Scholar]
- Zhang, R.; Isola, P.; Efros, A.A.; Shechtman, E.; Wang, O. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; IEEE: Piscataway, NJ, USA, 2018; pp. 586–595. [Google Scholar]
- Heusel, M.; Ramsauer, H.; Unterthiner, T.; Nessler, B.; Hochreiter, S. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Proceedings of the 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA, 4–9 December 2017; Neural Information Processing Systems Foundation, Inc.: Long Beach, CA, USA, 2017; Volume 30. [Google Scholar]
- Isola, P.; Zhu, J.Y.; Zhou, T.; Efros, A.A. Image-to-image translation with conditional adversarial networks. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; IEEE: Piscataway, NJ, USA, 2017; pp. 1125–1134. [Google Scholar]










| CelebA | Places2 | Paris Streetview | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SSIM ↑ | LPIPS ↓ | FID ↓ | GCM ↓ | SSIM ↑ | LPIPS ↓ | FID ↓ | GCM ↓ | SSIM ↑ | LPIPS ↓ | FID ↓ | GCM ↓ | |
| LaMa [5] | 0.85 | 0.11 | 16.55 | 0.16 | 0.83 | 0.12 | 33.25 | 0.14 | 0.80 | 0.18 | 80.22 | 0.10 |
| GMCNN [3] | 0.84 | 0.10 | 18.50 | 0.20 | 0.81 | 0.13 | 34.72 | 0.26 | 0.76 | 0.21 | 82.17 | 0.22 |
| RN [1] | 0.86 | 0.08 | 16.69 | 0.15 | 0.81 | 0.12 | 33.52 | 0.14 | 0.79 | 0.18 | 79.90 | 0.11 |
| VCNet [10] | 0.87 | 0.10 | 15.63 | 0.13 | 0.80 | 0.14 | 34.05 | 0.15 | 0.80 | 0.18 | 81.38 | 0.09 |
| CSCGAN w/o CSC loss | 0.86 | 0.09 | 15.77 | 0.12 | 0.79 | 0.13 | 33.46 | 0.15 | 0.79 | 0.19 | 80.05 | 0.10 |
| CSCGAN w/o MAA loss | 0.87 | 0.07 | 14.45 | 0.08 | 0.82 | 0.12 | 33.01 | 0.12 | 0.81 | 0.18 | 77.64 | 0.08 |
| CSCGAN | 0.88 | 0.06 | 13.23 | 0.05 | 0.82 | 0.11 | 32.89 | 0.11 | 0.82 | 0.17 | 74.28 | 0.07 |
| Item | Stage I | Stage II | Overall |
|---|---|---|---|
| Trainable parameters (M) | 42.3 | 6.4 | 48.7 |
| FLOPs per image (G) | 55.8 | 7.6 | 63.4 |
| Peak GPU memory (training, GB) | 8.7 | 2.1 | 9.8 |
| Training time per epoch (min) | 27 | 4 | 31 |
| Inference time per image (ms) | 101 | 25 | 126 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Jin, S.; Zhang, W.; Chu, T.; Zhang, Z.; Zhao, L.; Xing, W.; Lin, H.; Chen, L. CSCGAN: Cross-Space Contrastive Learning for Blind Image Inpainting. Appl. Sci. 2026, 16, 2969. https://doi.org/10.3390/app16062969
Jin S, Zhang W, Chu T, Zhang Z, Zhao L, Xing W, Lin H, Chen L. CSCGAN: Cross-Space Contrastive Learning for Blind Image Inpainting. Applied Sciences. 2026; 16(6):2969. https://doi.org/10.3390/app16062969
Chicago/Turabian StyleJin, Sheng, Weijing Zhang, Tianyi Chu, Zhanjie Zhang, Lei Zhao, Wei Xing, Huaizhong Lin, and Lixia Chen. 2026. "CSCGAN: Cross-Space Contrastive Learning for Blind Image Inpainting" Applied Sciences 16, no. 6: 2969. https://doi.org/10.3390/app16062969
APA StyleJin, S., Zhang, W., Chu, T., Zhang, Z., Zhao, L., Xing, W., Lin, H., & Chen, L. (2026). CSCGAN: Cross-Space Contrastive Learning for Blind Image Inpainting. Applied Sciences, 16(6), 2969. https://doi.org/10.3390/app16062969

