WaveletMask: Wavelet-Domain Mask-Guided Degradation Detection for Old-Film Restoration
Abstract
1. Introduction
- WaveletMask: frequency-aware gating for recurrent OFR. We formulate degradation sensing in the Haar wavelet domain and replace a single pixel-domain difference gate with a two-branch design that separately models high-frequency local damage and low-frequency flicker cues.
- LHFM for transient structural defects. LHFM applies geometric-mean differencing to the LH, HL, and HH sub-bands across three consecutive frames, emphasizing frame-specific scratches, dust, and other local defects while suppressing responses from persistent edges and one-sided temporal changes.
- GFM for frame-level brightness instability. GFM applies a minimum-difference criterion in the LL sub-band to detect current-frame brightness deviations relative to both temporal neighbors, yielding a dedicated cue for low-frequency flicker.
- A diagnostic evaluation framework for frequency-domain degradation masks. Under the SRWOV protocol, we establish a diagnostic evaluation framework that combines a unified re-training benchmark over ten methods with a strict paired statistical analysis against the RTN, and analyzes mask behavior through multi-seed reproducibility tests, paired clip-level significance analysis, artifact-stratified and failure-mode probes, fusion and wavelet-family ablations, and runtime measurement.
2. Related Work
2.1. Old-Film Restoration Methods
2.2. Wavelet Transforms in Image and Video Restoration
2.3. Degradation Detection and Adaptive Mechanisms
3. Materials and Methods
3.1. Problem Setup
3.2. Pipeline Overview

3.3. Local High-Frequency Mask Module (LHFM)

3.4. Global Flicker Mask Module (GFM)

3.5. Mask Fusion and Integration
3.6. Response Analysis
3.7. Training Objective
4. Experiments
4.1. Datasets
4.2. Implementation Details
4.3. Evaluation Metrics
4.4. Comparison with State-of-the-Art Methods
4.4.1. Quantitative Comparison on Synthetic Data
4.4.2. No-Reference Comparison on Real-World Data
4.4.3. Qualitative Comparison
4.5. Ablation Study
- Q1 (branch effectiveness). Do LHFM and GFM each improve over the RTN pixel-domain gate, and are the two branches complementary?
- Q2 (fusion rule). Is parameter-free branch-level maximum fusion preferable to learned pixel-wise fusion?
- Q3 (wavelet family). How sensitive is the gate to the wavelet basis and the number of decomposition levels?
- Q4 (robustness). Do the gains persist across artifact types, motion and occlusion levels, and unseen content (Appendix B)?

4.6. Computational Complexity
5. Discussion
5.1. Why the Wavelet-Domain Cues Help
5.2. Transferability, Boundary Conditions, and Evaluation Scope
5.3. Limitations and Future Work
6. Conclusions
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| OFR | Old-film restoration |
| DWT | Discrete wavelet transform |
| LHFM | Local High-Frequency Mask Module |
| GFM | Global Flicker Mask Module |
| HF | High frequency |
| LF | Low frequency |
| RTN | Recurrent Transformer Network |
| RRTN | Recursive Recurrent Transformer Network |
| SRWOV | Synthetic and Real-World Old Video (benchmark) |
| REDS | Realistic and Dynamic Scenes |
| RAFT | Recurrent All-Pairs Field Transform |
| ROI | Region of interest |
| GT | Ground truth |
| PSNR | Peak signal-to-noise ratio |
| SSIM | Structural Similarity Index Measure |
| LPIPS | Learned Perceptual Image Patch Similarity |
| DISTS | Deep Image Structure and Texture Similarity |
| NIQE | Natural Image Quality Evaluator |
| BRISQUE | Blind/Referenceless Image Spatial Quality Evaluator |
| CLIPIQA+ | CLIP-based Image Quality Assessment (prompt-tuned variant) |
| CLIP | Contrastive Language–Image Pre-training |
| CNN | Convolutional Neural Network |
| RNN | Recurrent Neural Network |
| GAN | Generative Adversarial Network |
| VGG | Visual Geometry Group |
| GPU | Graphics Processing Unit |
| FLOPs | Floating Point Operations |
| MAE | Mean Absolute Error |
| RGB | Red–Green–Blue |
| BT.601 | ITU-R BT.601 luminance transform |
| LL/LH/HL/HH | Low–Low/Low–High/High–Low/High–High |
Appendix A. Statistical Reliability of the Paired Comparison
| Metric | RTN | WaveletMask | Paired Improvement | 95% CI/Two-Sided p |
|---|---|---|---|---|
| PSNR ↑ | 26.1648 | 26.6125 | , | |
| SSIM ↑ | 0.8923 | 0.8954 | , | |
| LPIPS ↓ | 0.1023 | 0.0886 | , | |
| DISTS ↓ | 0.0946 | 0.0827 | , |
| Method | PSNR ↑ | SSIM ↑ | LPIPS ↓ | DISTS ↓ |
|---|---|---|---|---|
| RTN [3] | 24.5570 ± 1.6667 | 0.8808 ± 0.0129 | 0.1083 ± 0.0118 | 0.0922 ± 0.0040 |
| WaveletMask (ours) | 26.2183 ± 0.3725 | 0.8983 ± 0.0045 | 0.0879 ± 0.0026 | 0.0831 ± 0.0015 |
| Mean difference | +1.6613 | +0.0175 | −0.0205 | −0.0091 |
Appendix B. Robustness and Cross-Content Generalization
Appendix B.1. Artifact-Dominant Synthetic Breakdown
| Improvement over RTN | ||||||
|---|---|---|---|---|---|---|
| Stratum | Clips | RTN PSNR | WaveletMask PSNR | PSNR (dB) | SSIM | LPIPS/DISTS |
| Scratch-dominant | 5 | 27.47 | 28.11 | +0.64 | +0.0020 | +0.0110/+0.0116 |
| Dust-dominant | 8 | 26.10 | 26.55 | +0.45 | +0.0049 | +0.0196/+0.0120 |
| Flicker-dominant | 6 | 23.90 | 24.37 | +0.47 | +0.0055 | +0.0176/+0.0153 |
| Mixed residual | 11 | 26.83 | 27.17 | +0.34 | +0.0017 | +0.0102/+0.0109 |
Appendix B.2. Motion- and Occlusion-Binned Analysis
| Improvement over RTN | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| Group | Bin | Clips | Motion | Occl. | PSNR (dB) | SSIM | LPIPS | DISTS | LHFM/GFM Clean-FP |
| Motion | Low | 10 | 0.77 | 0.009 | +0.085 | +0.0009 | +0.0115 | +0.0127 | 0.646/0.440 |
| Motion | Medium | 10 | 1.87 | 0.027 | +0.477 | +0.0036 | +0.0172 | +0.0135 | 0.782/0.627 |
| Motion | High | 10 | 3.90 | 0.053 | +0.769 | +0.0056 | +0.0143 | +0.0103 | 0.681/0.584 |
| Occlusion | Low | 10 | 0.85 | 0.005 | +0.186 | +0.0018 | +0.0119 | +0.0133 | 0.716/0.515 |
| Occlusion | Medium | 10 | 2.02 | 0.022 | +0.498 | +0.0018 | +0.0139 | +0.0113 | 0.703/0.541 |
| Occlusion | High | 10 | 3.67 | 0.062 | +0.647 | +0.0065 | +0.0172 | +0.0119 | 0.690/0.594 |
Appendix B.3. Cross-Content Generalization on DAVIS
| Model | PSNR ↑ | SSIM ↑ | LPIPS ↓ | DISTS ↓ |
|---|---|---|---|---|
| RTN | 25.8417 | 0.8827 | 0.1223 | 0.1089 |
| WaveletMask | 25.9883 | 0.8795 | 0.1206 | 0.1048 |
| Improvement | +0.15 | −0.003 | +0.002 | +0.004 |
Appendix C. Detector Diagnostic Probes
Appendix C.1. GFM Failure-Mode Probe
| Probe | Clips | Clean Activation | Probe Activation | Ratio | FPR | Guarded Activation | Suppression |
|---|---|---|---|---|---|---|---|
| Scene cut | 12 | 0.0041 | 0.0066 | 2.35× | 0.0505 | 0.0024 | 0.6141 |
| Exposure flash | 12 | 0.0041 | 0.0257 | 10.41× | 0.1574 | 0.0015 | 0.9470 |
| Illumination ramp | 12 | 0.0041 | 0.0051 | 1.52× | 0.0352 | 0.0049 | 0.0247 |
Appendix C.2. Persistent-Scratch Response Probe
| Mode | L | Clips | Raw LHFM | Raw Relative Response | Raw Suppression | Relative Response |
|---|---|---|---|---|---|---|
| Fixed | 1 | 12 | 0.0862 | 1.0000 | 0.0000 | 1.0000 |
| Fixed | 2 | 12 | 0.0462 | 0.5355 | 0.4645 | 0.0167 |
| Fixed | 3 | 12 | 0.0302 | 0.3498 | 0.6502 | 0.0108 |
| Fixed | 5 | 12 | 0.0302 | 0.3498 | 0.6502 | 0.0108 |
| Drift | 1 | 12 | 0.0862 | 1.0000 | 0.0000 | 1.0000 |
| Drift | 2 | 12 | 0.0894 | 1.0364 | −0.0364 | 0.1363 |
| Drift | 3 | 12 | 0.0957 | 1.1097 | −0.1097 | 0.0419 |
| Drift | 5 | 12 | 0.0957 | 1.1097 | −0.1097 | 0.0419 |
Appendix C.3. Mask Interpretability and Residual-Proxy Overlap
| Mask | Proxy Area | AUPRC | Best F1 | Best IoU | Prec. | Rec. | In/Out | Top-5 Hit |
|---|---|---|---|---|---|---|---|---|
| 0.0820 | 0.0552 | 0.1515 | 0.0820 | 0.0820 | 1.0000 | 0.6103 | 0.0412 | |
| 0.2858 | 0.4888 | 0.4949 | 0.3288 | 0.4087 | 0.6271 | 3.5714 | 0.6854 | |
| 0.2907 | 0.4943 | 0.4996 | 0.3330 | 0.4128 | 0.6328 | 3.5638 | 0.6908 |
Appendix C.4. Cue-Complementarity Analysis
| Cue Pair | Frames | Pearson | Spearman | Top-5 Overlap/Jaccard |
|---|---|---|---|---|
| Spatial/ | 60 | 0.0717 | −0.0418 | 0.1641/0.0938 |
| Spatial/ | 60 | 0.3763 | 0.4850 | 0.2300/0.1408 |
| / | 60 | −0.1635 | −0.3364 | 0.0024/0.0012 |
Appendix C.5. Cross-Domain Frequency-Operator Diagnostic
| Operator | Family | AUPRC ↑ | Top-5 Hit ↑ | Top-5 Non-Proxy ↓ |
|---|---|---|---|---|
| Haar fused | Haar DWT | 0.3166 | 0.3685 | 0.6315 |
| Sobel | Gradient edge | 0.3062 | 0.3470 | 0.6530 |
| Laplacian | Gradient edge | 0.2993 | 0.3352 | 0.6648 |
| Gaussian fused | Gaussian pyramid | 0.3270 | 0.3815 | 0.6185 |
| Fourier temporal fused | Fourier temporal | 0.3885 | 0.4384 | 0.5616 |
| Learned shallow frequency | Proxy-supervised | 0.4265 | 0.5278 | 0.4722 |
Appendix D. Complexity Details and Full Ablation Sweeps
Appendix D.1. Full Fusion Strategy Comparison
| Variant | LHFM | GFM | Fusion | PSNR ↑ | SSIM ↑ | LPIPS ↓ | DISTS ↓ |
|---|---|---|---|---|---|---|---|
| RTN | — | — | — | 26.1580 | 0.8917 | 0.1034 | 0.0953 |
| +LHFM | ✓ | — | direct | 26.3778 | 0.8966 | 0.0914 | 0.0790 |
| +LHFM + learned | ✓ | — | CNN | 25.9858 | 0.8978 | 0.0889 | 0.0830 |
| +GFM | — | ✓ | direct | 26.1024 | 0.8980 | 0.0969 | 0.0916 |
| +GFM + learned | — | ✓ | CNN | 25.9488 | 0.8956 | 0.0946 | 0.0814 |
| WaveletMask | ✓ | ✓ | max | 26.6125 | 0.8954 | 0.0886 | 0.0827 |
| WaveletMask + gated | ✓ | ✓ | gated CNN | 26.3515 | 0.8966 | 0.0869 | 0.0842 |
| WaveletMask + learned | ✓ | ✓ | max + CNN | 24.9827 | 0.8896 | 0.1068 | 0.0867 |
Appendix D.2. Full Wavelet-Family and Stronger-Fusion Sweeps
| Variant | PSNR ↑ | SSIM ↑ | LPIPS ↓ | DISTS ↓ |
|---|---|---|---|---|
| Haar, single level (seed 2021) | 26.6125 | 0.8954 | 0.0886 | 0.0827 |
| Haar, two levels | 26.5175 | 0.9011 | 0.0847 | 0.0760 |
| DTCWT (dual-tree complex) | 26.4506 | 0.8992 | 0.0862 | 0.0780 |
| Coiflet-1 | 25.8533 | 0.8874 | 0.1026 | 0.0893 |
| Symlet-2 | 24.2087 | 0.8509 | 0.1507 | 0.1226 |
| Daubechies-2 | 17.9208 | 0.6473 | 0.4136 | 0.3154 |
| Attention fusion | 23.2479 | 0.8281 | 0.1768 | 0.1359 |
| Residual fusion | 25.9777 | 0.8929 | 0.0966 | 0.0811 |
Appendix D.3. Runtime Measurement Protocol
Appendix D.4. Per-Module Parameter Breakdown
| Module | Parameters | Description |
|---|---|---|
| Haar DWT | 0 | Fixed filter kernels, no learnable weights |
| WaveletNoiseMask CNNs | 898 | LHFM and GFM lightweight branches (Section 3.3 and Section 3.4) |
| PixelWiseMaskFusion | 14,497 | Learned spatial confidence weighting, ablated only (Appendix D.1) |
| WaveletMask detector | 898 | LHFM + GFM + maximum fusion (proposed) |
| Learned-fusion variant | 15,395 | Detector with PixelWiseMaskFusion (ablation) |
Appendix D.5. No-Reference Ablation on Real-World Data
| Variant | LHFM | GFM | NIQE ↓ | BRISQUE ↓ | CLIPIQA+ ↑ |
|---|---|---|---|---|---|
| RTN | — | — | 5.3340 | 29.3119 | 0.4359 |
| +LHFM only | ✓ | — | 5.3392 | 19.5562 | 0.4115 |
| +GFM only | — | ✓ | 5.0400 | 28.7878 | 0.4388 |
| WaveletMask | ✓ | ✓ | 5.2288 | 24.2023 | 0.4356 |
References
- Kokaram, A.C. Motion Picture Restoration: Digital Algorithms for Artefact Suppression in Degraded Motion Picture Film and Video; Springer: London, UK, 1998. [Google Scholar] [CrossRef]
- Iizuka, S.; Simo-Serra, E. DeepRemaster: Temporal Source-Reference Attention Networks for Comprehensive Video Enhancement. ACM Trans. Graph. 2019, 38, 176:1–176:13. [Google Scholar] [CrossRef]
- Wan, Z.; Zhang, B.; Chen, D.; Liao, J. Bringing Old Films Back to Life. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 19–24 June 2022; pp. 17694–17703. [Google Scholar] [CrossRef]
- Wan, Z.; Zhang, B.; Chen, D.; Zhang, P.; Chen, D.; Liao, J.; Wen, F. Bringing Old Photos Back to Life. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Virtual, 14–19 June 2020; pp. 2747–2757. [Google Scholar] [CrossRef]
- Su, J.; Xu, B.; Yin, H. A Survey of Deep Learning Approaches to Image Restoration. Neurocomputing 2022, 487, 46–65. [Google Scholar] [CrossRef]
- Tian, C.; Fei, L.; Zheng, W.; Xu, Y.; Zuo, W.; Lin, C.W. Deep Learning on Image Denoising: An Overview. Neural Netw. 2020, 131, 251–275. [Google Scholar] [CrossRef] [PubMed]
- Lin, S.; Simo-Serra, E. Restoring Degraded Old Films with Recursive Recurrent Transformer Networks. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA, 3–8 January 2024; pp. 6718–6728. [Google Scholar] [CrossRef]
- Mao, Y.; Luo, H.; Zhong, Z.; Chen, P.; Zhang, Z.; Wang, S. Making Old Film Great Again: Degradation-aware State Space Model for Old Film Restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 11–15 June 2025; pp. 28039–28049. [Google Scholar] [CrossRef]
- Guo, H.; Li, J.; Dai, T.; Ouyang, Z.; Ren, X.; Xia, S.T. MambaIR: A Simple Baseline for Image Restoration with State-Space Model. In Proceedings of the Computer Vision—ECCV 2024, Milan, Italy, 29 September–4 October 2024; pp. 222–241. [Google Scholar] [CrossRef]
- Mallat, S.G. A Theory for Multiresolution Signal Decomposition: The Wavelet Representation. IEEE Trans. Pattern Anal. Mach. Intell. 1989, 11, 674–693. [Google Scholar] [CrossRef]
- Bernardini, L.; Bono, F.M.; Collina, A. Drive-by Damage Detection and Localization Exploiting Continuous Wavelet Transform and Multiple Sparse Autoencoders. Railw. Eng. Sci. 2025, 33, 721–745. [Google Scholar] [CrossRef]
- Liu, P.; Zhang, H.; Zhang, K.; Lin, L.; Zuo, W. Multi-Level Wavelet-CNN for Image Restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Salt Lake City, UT, USA, 18–22 June 2018; pp. 773–782. [Google Scholar] [CrossRef]
- Kokaram, A.C.; Morris, R.D.; Fitzgerald, W.J.; Rayner, P.J.W. Detection of Missing Data in Image Sequences. IEEE Trans. Image Process. 1995, 4, 1496–1508. [Google Scholar] [CrossRef] [PubMed]
- Joyeux, L.; Boukir, S.; Besserer, B.; Buisson, O. Reconstruction of Degraded Image Sequences. Application to Film Restoration. Image Vis. Comput. 2001, 19, 503–516. [Google Scholar] [CrossRef]
- van Roosmalen, P.M.B.; Lagendijk, R.L.; Biemond, J. Correction of Intensity Flicker in Old Film Sequences. IEEE Trans. Circuits Syst. Video Technol. 1999, 9, 1013–1019. [Google Scholar] [CrossRef][Green Version]
- Rota, C.; Buzzelli, M.; Bianco, S.; Schettini, R. Video Restoration Based on Deep Learning: A Comprehensive Survey. Artif. Intell. Rev. 2023, 56, 5317–5364. [Google Scholar] [CrossRef]
- Chan, K.C.; Wang, X.; Yu, K.; Dong, C.; Loy, C.C. BasicVSR: The Search for Essential Components in Video Super-Resolution and Beyond. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Virtual, 19–25 June 2021; pp. 4947–4956. [Google Scholar] [CrossRef]
- Chan, K.C.; Zhou, S.; Xu, X.; Loy, C.C. BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and Alignment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 19–24 June 2022; pp. 5972–5981. [Google Scholar] [CrossRef]
- Ranjan, A.; Black, M.J. Optical Flow Estimation Using a Spatial Pyramid Network. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 4161–4170. [Google Scholar] [CrossRef]
- Teed, Z.; Deng, J. RAFT: Recurrent All-Pairs Field Transforms for Optical Flow. In Proceedings of the Computer Vision—ECCV 2020, Glasgow, UK, 23–28 August 2020; pp. 402–419. [Google Scholar] [CrossRef]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, L.; Polosukhin, I. Attention Is All You Need. In Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA, 4–9 December 2017; Volume 30, pp. 5998–6008. [Google Scholar]
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Virtual, 11–17 October 2021; pp. 10012–10022. [Google Scholar] [CrossRef]
- Liang, J.; Fan, Y.; Xiang, X.; Ranjan, R.; Ilg, E.; Green, S.; Cao, J.; Zhang, K.; Timofte, R.; Van Gool, L. Recurrent Video Restoration Transformer with Guided Deformable Attention. In Proceedings of the Advances in Neural Information Processing Systems, New Orleans, LA, USA, 28 November–3 December 2022; Volume 35, pp. 378–393. [Google Scholar] [CrossRef]
- Liang, J.; Cao, J.; Fan, Y.; Zhang, K.; Ranjan, R.; Li, Y.; Timofte, R.; Van Gool, L. VRT: A Video Restoration Transformer. IEEE Trans. Image Process. 2024, 33, 2171–2182. [Google Scholar] [CrossRef] [PubMed]
- Li, D.; Shi, X.; Zhang, Y.; Cheung, K.C.; See, S.; Wang, X.; Qin, H.; Li, H. A Simple Baseline for Video Restoration with Grouped Spatial-Temporal Shift. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 17–24 June 2023; pp. 9822–9832. [Google Scholar] [CrossRef]
- Liang, J.; Cao, J.; Sun, G.; Zhang, K.; Van Gool, L.; Timofte, R. SwinIR: Image Restoration Using Swin Transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), Virtual, 11–17 October 2021; pp. 1833–1844. [Google Scholar] [CrossRef]
- Zamir, S.W.; Arora, A.; Khan, S.; Hayat, M.; Khan, F.S.; Yang, M.H. Restormer: Efficient Transformer for High-Resolution Image Restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 19–24 June 2022; pp. 5728–5739. [Google Scholar] [CrossRef]
- Wang, X.; Xie, L.; Dong, C.; Shan, Y. Real-ESRGAN: Training Real-World Blind Super-Resolution with Pure Synthetic Data. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), Virtual, 11–17 October 2021; pp. 1905–1914. [Google Scholar] [CrossRef]
- Zeng, Y.; Fu, J.; Chao, H. Learning Joint Spatial-Temporal Transformations for Video Inpainting. In Proceedings of the Computer Vision—ECCV 2020, Glasgow, UK, 23–28 August 2020; pp. 528–543. [Google Scholar] [CrossRef]
- Zhou, S.; Li, C.; Chan, K.C.; Loy, C.C. ProPainter: Improving Propagation and Transformer for Video Inpainting. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 2–6 October 2023; pp. 10477–10486. [Google Scholar] [CrossRef]
- Huang, H.; He, R.; Sun, Z.; Tan, T. Wavelet-SRNet: A Wavelet-Based CNN for Multi-Scale Face Super Resolution. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017; pp. 1689–1697. [Google Scholar] [CrossRef]
- Deng, X.; Yang, R.; Xu, M.; Dragotti, P.L. Wavelet Domain Style Transfer for an Effective Perception-Distortion Tradeoff in Single Image Super-Resolution. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 3076–3085. [Google Scholar] [CrossRef]
- Chen, L.F.; Liao, H.Y.M.; Lin, J.C. Wavelet-Based Optical Flow Estimation. IEEE Trans. Circuits Syst. Video Technol. 2002, 12, 1–12. [Google Scholar] [CrossRef]
- Selesnick, I.W.; Baraniuk, R.G.; Kingsbury, N.G. The Dual-Tree Complex Wavelet Transform. IEEE Signal Process. Mag. 2005, 22, 123–151. [Google Scholar] [CrossRef]
- Li, B.; Liu, X.; Hu, P.; Wu, Z.; Lv, J.; Peng, X. All-in-One Image Restoration for Unknown Corruption. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 19–24 June 2022; pp. 17452–17462. [Google Scholar] [CrossRef]
- Potlapalli, V.; Zamir, S.W.; Khan, S.H.; Khan, F.S. PromptIR: Prompting for All-in-One Image Restoration. In Proceedings of the Advances in Neural Information Processing Systems, New Orleans, LA, USA, 10–16 December 2023; Volume 36, pp. 71275–71293. [Google Scholar] [CrossRef]
- Johnson, J.; Alahi, A.; Li, F.F. Perceptual Losses for Real-Time Style Transfer and Super-Resolution. In Proceedings of the Computer Vision—ECCV 2016, Amsterdam, The Netherlands, 8–16 October 2016; pp. 694–711. [Google Scholar] [CrossRef]
- Simonyan, K.; Zisserman, A. Very Deep Convolutional Networks for Large-Scale Image Recognition. In Proceedings of the International Conference on Learning Representations (ICLR), San Diego, CA, USA, 7–9 May 2015. [Google Scholar]
- Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; Bengio, Y. Generative Adversarial Nets. In Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada, 8–13 December 2014; Volume 27, pp. 2672–2680. [Google Scholar]
- Nah, S.; Baik, S.; Hong, S.; Moon, G.; Son, S.; Timofte, R.; Lee, K.M. NTIRE 2019 Challenge on Video Deblurring and Super-Resolution: Dataset and Study. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Long Beach, CA, USA, 16–20 June 2019; pp. 1996–2005. [Google Scholar] [CrossRef]
- Antic, J. DeOldify. 2019. Available online: https://github.com/jantic/DeOldify (accessed on 1 April 2026).
- Kingma, D.P.; Ba, J. Adam: A Method for Stochastic Optimization. In Proceedings of the International Conference on Learning Representations (ICLR), San Diego, CA, USA, 7–9 May 2015. [Google Scholar]
- Wang, X.; Xie, L.; Yu, K.; Chan, K.C.; Loy, C.C.; Dong, C. BasicSR: Open Source Image and Video Restoration Toolbox. GitHub Repository. 2022. Available online: https://github.com/XPixelGroup/BasicSR (accessed on 1 April 2026).
- Wang, Z.; Bovik, A.C.; Sheikh, H.R.; Simoncelli, E.P. Image Quality Assessment: From Error Visibility to Structural Similarity. IEEE Trans. Image Process. 2004, 13, 600–612. [Google Scholar] [CrossRef] [PubMed]
- Zhang, R.; Isola, P.; Efros, A.A.; Shechtman, E.; Wang, O. The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018; pp. 586–595. [Google Scholar] [CrossRef]
- Ding, K.; Ma, K.; Wang, S.; Simoncelli, E.P. Image Quality Assessment: Unifying Structure and Texture Similarity. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 44, 2567–2581. [Google Scholar] [CrossRef] [PubMed]
- Mittal, A.; Soundararajan, R.; Bovik, A.C. Making a “Completely Blind” Image Quality Analyzer. IEEE Signal Process. Lett. 2013, 20, 209–212. [Google Scholar] [CrossRef]
- Mittal, A.; Moorthy, A.K.; Bovik, A.C. No-Reference Image Quality Assessment in the Spatial Domain. IEEE Trans. Image Process. 2012, 21, 4695–4708. [Google Scholar] [CrossRef] [PubMed]
- Wang, J.; Chan, K.C.; Loy, C.C. Exploring CLIP for Assessing the Look and Feel of Images. Proc. AAAI Conf. Artif. Intell. 2023, 37, 2555–2563. [Google Scholar] [CrossRef]
- Kim, K.; Kim, Y.; Kim, Y.J. Hybrid Frequency–Spatial Domain Learning for Image Restoration in Under-Display Camera Systems Using Augmented Virtual Big Data Generated by the Angular Spectrum Method. Appl. Sci. 2025, 15, 30. [Google Scholar] [CrossRef]
- Park, C.H.; Choi, H.D.; Lim, M.T. Harnessing Spatial-Frequency Information for Enhanced Image Restoration. Appl. Sci. 2025, 15, 1856. [Google Scholar] [CrossRef]
- Rubel, A.; Ieremeiev, O.; Lukin, V.; Fastowicz, J.; Okarma, K. Combined No-Reference Image Quality Metrics for Visual Quality Assessment Optimized for Remote Sensing Images. Appl. Sci. 2022, 12, 1986. [Google Scholar] [CrossRef]
- Yeo, W.H.; Ryu, H.C. An Integration Framework for the Inpainting and Colorization of Arbitrary Masked Grayscale Images. Appl. Sci. 2025, 15, 1978. [Google Scholar] [CrossRef]



| Method | Year | Backbone | Gating Cue Domain | Temporal Criterion |
|---|---|---|---|---|
| DeepRemaster [2] | 2019 | 3D CNN | – | – |
| RTN [3] | 2022 | Bi-RNN + Swin | Pixel difference | Single-frame |
| RRTN [7] | 2024 | Recursive RNN + Swin | Pixel difference | Bilateral |
| MambaOFR [8] | 2025 | Mamba | Latent prompt | – |
| WaveletMask (ours) | 2026 | Bi-RNN + Swin | Haar wavelet sub-bands | Bilateral geometric mean/minimum |
| Pixel-Domain Detection | Wavelet-Domain Detection (Ours) | |
|---|---|---|
| Signal domain | Spatial intensity values | Haar wavelet sub-bands |
| High-frequency damage | Response can be diluted when the defect footprint is spatially sparse | Detail bands often preserve sharper local responses |
| Flicker discrimination | Global intensity changes can be confounded with legitimate scene variation | The approximation band provides a more direct brightness cue |
| Temporal criterion | Single-frame absolute difference | Bilateral geometric mean (LHFM) and minimum strategy (GFM) |
| Method | Year | PSNR ↑ | SSIM ↑ | LPIPS ↓ | DISTS ↓ |
|---|---|---|---|---|---|
| DeepRemaster [2] | 2019 | 25.0590 | 0.9185 | 0.1108 | 0.0895 |
| DeOldify [41] | 2019 | 23.2197 | 0.8923 | 0.2775 | 0.1260 |
| OldPhoto [4] | 2020 | 24.6174 | 0.8524 | 0.2252 | 0.1144 |
| RTN [3] | 2022 | 25.8514 | 0.8869 | 0.1214 | 0.0862 |
| RVRT [23] | 2022 | 22.4317 | 0.8710 | 0.1655 | 0.1078 |
| ShiftNet [25] | 2023 | 21.4817 | 0.8234 | 0.2030 | 0.1303 |
| RRTN [7] | 2024 | 24.1450 | 0.9006 | 0.1082 | 0.0668 |
| VRT [24] | 2024 | 22.3866 | 0.8631 | 0.1750 | 0.1087 |
| MambaOFR [8] | 2025 | 25.9938 | 0.9059 | 0.0815 | 0.0822 |
| WaveletMask (ours) | 2026 | 26.6017 | 0.8951 | 0.0891 | 0.0831 |
| Method | NIQE ↓ | BRISQUE ↓ | CLIPIQA+ ↑ |
|---|---|---|---|
| Degraded input | 5.7745 | 43.1050 | 0.3873 |
| DeOldify [41] | 5.7218 | 41.6118 | 0.3768 |
| DeepRemaster [2] | 5.9917 | 40.6158 | 0.3991 |
| OldPhoto [4] | 6.5601 | 30.4532 | 0.4503 |
| RTN [3] | 5.1723 | 19.0031 | 0.4163 |
| VRT [24] | 5.6551 | 42.7296 | 0.3985 |
| RVRT [23] | 5.7217 | 42.6402 | 0.4013 |
| ShiftNet [25] | 5.7126 | 42.8195 | 0.3975 |
| RRTN [7] | 5.6664 | 32.8572 | 0.3969 |
| MambaOFR [8] | 5.3381 | 29.7853 | 0.4351 |
| WaveletMask (ours) | 5.2288 | 24.2023 | 0.4356 |
| Variant | PSNR ↑ | SSIM ↑ | LPIPS ↓ | DISTS ↓ |
|---|---|---|---|---|
| (a) Detection-branch components | ||||
| RTN | 26.1648 | 0.8923 | 0.1023 | 0.0946 |
| +LHFM only | 26.3778 | 0.8966 | 0.0914 | 0.0790 |
| +GFM only | 26.1024 | 0.8980 | 0.0969 | 0.0916 |
| WaveletMask (max fusion) | 26.6125 | 0.8954 | 0.0886 | 0.0827 |
| (b) Fusion strategy (key contrasts) | ||||
| Branch-level maximum (proposed) | 26.6125 | 0.8954 | 0.0886 | 0.0827 |
| Gated CNN fusion | 26.3515 | 0.8966 | 0.0869 | 0.0842 |
| Max + pixel-wise CNN fusion | 24.9827 | 0.8896 | 0.1068 | 0.0867 |
| (c) Wavelet family (key contrasts) | ||||
| Haar, single level (proposed) | 26.6125 | 0.8954 | 0.0886 | 0.0827 |
| Haar, two levels | 26.5175 | 0.9011 | 0.0847 | 0.0760 |
| DTCWT (dual-tree complex) | 26.4506 | 0.8992 | 0.0862 | 0.0780 |
| Daubechies-2 | 17.9208 | 0.6473 | 0.4136 | 0.3154 |
| Method | Params (M) | FLOPs (G) | Latency (ms/clip) ↓ | Throughput (fps) ↑ | Peak Mem. (GB) ↓ |
|---|---|---|---|---|---|
| DeepRemaster [2] | 9.8781 | 914.0 | 17.9 | 391.6 | 0.35 |
| RTN [3] | 5.7351 | 2184.7 | 169.5 | 41.3 | 0.24 |
| RRTN [7] | 7.0503 | 2782.2 | 298.4 | 23.5 | 0.75 |
| MambaOFR [8] | 8.2984 | 2832.8 | 404.8 | 17.3 | 0.60 |
| WaveletMask | 5.7360 | 2184.9 | 171.4 | 40.8 | 0.24 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Cai, F.; Zhang, Q.; Xu, C.; Ding, Y. WaveletMask: Wavelet-Domain Mask-Guided Degradation Detection for Old-Film Restoration. Appl. Sci. 2026, 16, 6415. https://doi.org/10.3390/app16136415
Cai F, Zhang Q, Xu C, Ding Y. WaveletMask: Wavelet-Domain Mask-Guided Degradation Detection for Old-Film Restoration. Applied Sciences. 2026; 16(13):6415. https://doi.org/10.3390/app16136415
Chicago/Turabian StyleCai, Feifan, Qi Zhang, Chang’an Xu, and Youdong Ding. 2026. "WaveletMask: Wavelet-Domain Mask-Guided Degradation Detection for Old-Film Restoration" Applied Sciences 16, no. 13: 6415. https://doi.org/10.3390/app16136415
APA StyleCai, F., Zhang, Q., Xu, C., & Ding, Y. (2026). WaveletMask: Wavelet-Domain Mask-Guided Degradation Detection for Old-Film Restoration. Applied Sciences, 16(13), 6415. https://doi.org/10.3390/app16136415

