Enhanced Lightweight Image Super-Resolution via Residual Aggregation and Wavelet Loss
Abstract
1. Introduction
- We develop a lightweight SISR architecture that organizes feature aggregation across local, mesoscale, and non-local spatial ranges. LAB preserves neighborhood textures, MRGAB integrates grouped-residual attention for efficient regional interaction, and NLSAB captures long-range dependencies through sparse aggregation.
- We integrate grouped residual linear projections into the mesoscale RGSA stage to reduce redundancy in conventional query, key, and value projections. This design coordinates efficient mesoscale interaction with the local features extracted by LAB and the non-local information aggregated by NLSAB.
- We introduce a composite training objective that combines RGB-domain L1 loss with SWT-domain supervision. The SWT loss explicitly constrains low-frequency structures and directional high-frequency details without introducing an additional inference branch.
- Experiments on five standard benchmark datasets demonstrate that RAW achieves a competitive trade-off between reconstruction quality and model complexity. For SR, the complete model reduces the parameter count and MACs by 14.3% and 15.4%, respectively, relative to the baseline, while improving the PSNR and SSIM on Manga109 by 0.22 dB and 0.0015, respectively.
2. Related Work
2.1. CNN- and Transformer-Based SR Methods
2.2. Attention Mechanism
3. Proposed Method
3.1. Framework
3.2. Residual Aggregation Block (RAB)
3.3. Loss Function
4. Experimental Results and Analyses
4.1. Experimental Configuration
4.2. Comparison with State-of-the-Art Methods
4.3. Ablation Study
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| CNN | Convolutional neural network |
| GRL | Grouped residual linear |
| LAB | Local aggregation block |
| MACs | Multiply-accumulate operations |
| MRGAB | Mesoscale residual grouping aggregation block |
| NLSAB | Non-local sparse aggregation block |
| OSA | Omni-scale aggregation |
| PSNR | Peak signal-to-noise ratio |
| RAB | Residual aggregation block |
| RAW | Residual aggregation and wavelet loss |
| RGB | Red–green–blue |
| SISR | Single-image super-resolution |
| SR | Super-resolution |
| SSIM | Structural similarity index measure |
| SWT | Stationary wavelet transform |
| YCbCr | YCbCr color space |
References
- Dong, C.; Loy, C.C.; He, K.; Tang, X. Image super-resolution using deep convolutional networks. IEEE Trans. Pattern Anal. Mach. Intell. 2015, 38, 295–307. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lai, W.S.; Huang, J.B.; Ahuja, N.; Yang, M.H. Fast and accurate image super-resolution with deep laplacian pyramid networks. IEEE Trans. Pattern Anal. Mach. Intell. 2018, 41, 2599–2613. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhang, Y.; Li, K.; Li, K.; Zhong, B.; Fu, Y. Residual non-local attention networks for image restoration. arXiv 2019, arXiv:1903.10082. [Google Scholar]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. Adv. Neural Inf. Process. Syst. 2017, 30, 5998–6008. [Google Scholar]
- Dosovitskiy, A. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv 2020, arXiv:2010.11929. [Google Scholar]
- Talreja, J.; Aramvith, S.; Onoye, T. Dhtcun: Deep hybrid transformer cnn u network for single-image super-resolution. IEEE Access 2024, 12, 122624–122641. [Google Scholar] [CrossRef] [Scilit]
- Bian, Y.; Huang, J.; Cai, X.; Yuan, J.; Church, K. On attention redundancy: A comprehensive study. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Online, 6–11 June 2021; pp. 930–945. [Google Scholar]
- Zhang, X.; Zeng, H.; Guo, S.; Zhang, L. Efficient long-range attention network for image super-resolution. In Proceedings of the European Conference on Computer Vision; Springer: Cham, Switzerland, 2022; pp. 649–667. [Google Scholar]
- Li, Y.; Deng, Z.; Cao, Y.; Liu, L. GRFormer: Grouped residual self-attention for lightweight single image super-resolution. In Proceedings of the 32nd ACM International Conference on Multimedia, Melbourne, VIC, Australia, 28 October–1 November 2024; pp. 9378–9386. [Google Scholar]
- Mehta, S.; Rastegari, M. Mobilevit: Light-weight, general-purpose, and mobile-friendly vision transformer. arXiv 2021, arXiv:2110.02178. [Google Scholar]
- Wang, Y.; Li, Y.; Wang, G.; Liu, X. Multi-scale attention network for single image super-resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 5950–5960. [Google Scholar]
- Korkmaz, C.; Tekalp, A.M. Training transformer models by wavelet losses improves quantitative and visual performance in single image super-resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 6661–6670. [Google Scholar]
- Kim, J.; Lee, J.K.; Lee, K.M. Accurate image super-resolution using very deep convolutional networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; pp. 1646–1654. [Google Scholar]
- Lim, B.; Son, S.; Kim, H.; Nah, S.; Mu Lee, K. Enhanced deep residual networks for single image super-resolution. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Honolulu, HI, USA, 21–26 July 2017; pp. 136–144. [Google Scholar]
- Zhang, Y.; Li, K.; Li, K.; Wang, L.; Zhong, B.; Fu, Y. Image super-resolution using very deep residual channel attention networks. In Proceedings of the European conference on computer vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 286–301. [Google Scholar]
- Liang, J.; Cao, J.; Sun, G.; Zhang, K.; Van Gool, L.; Timofte, R. Swinir: Image restoration using swin transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 11–17 October 2021; pp. 1833–1844. [Google Scholar]
- Chen, H.; Wang, Y.; Guo, T.; Xu, C.; Deng, Y.; Liu, Z.; Ma, S.; Xu, C.; Xu, C.; Gao, W. Pre-trained image processing transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 19–25 June 2021; pp. 12299–12310. [Google Scholar]
- Lu, Z.; Li, J.; Liu, H.; Huang, C.; Zhang, L.; Zeng, T. Transformer for single image super-resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 457–466. [Google Scholar]
- Zhou, Y.; Zhang, Y.; Xie, X.; Kung, S.Y. Image super-resolution based on dense convolutional auto-encoder blocks. Neurocomputing 2021, 423, 98–109. [Google Scholar] [CrossRef] [Scilit]
- Dai, T.; Cai, J.; Zhang, Y.; Xia, S.T.; Zhang, L. Second-order attention network for single image super-resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019; pp. 11065–11074. [Google Scholar]
- Mei, Y.; Fan, Y.; Zhang, Y.; Yu, J.; Zhou, Y.; Liu, D.; Fu, Y.; Huang, T.S.; Shi, H. Pyramid attention network for image restoration. Int. J. Comput. Vis. 2023, 131, 3207–3225. [Google Scholar] [CrossRef] [Scilit]
- Niu, B.; Wen, W.; Ren, W.; Zhang, X.; Yang, L.; Wang, S.; Zhang, K.; Cao, X.; Shen, H. Single image super-resolution via a holistic attention network. In Proceedings of the European Conference on Computer Vision; Springer: Cham, Switzerland, 2020; pp. 191–207. [Google Scholar]
- Hui, Z.; Gao, X.; Yang, Y.; Wang, X. Lightweight image super-resolution with information multi-distillation network. In Proceedings of the 27th ACM International Conference on Multimedia, Nice, France, 21–25 October 2019; pp. 2024–2032. [Google Scholar]
- Wang, H.; Chen, X.; Ni, B.; Liu, Y.; Liu, J. Omni aggregation networks for lightweight image super-resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 22378–22387. [Google Scholar]
- Kim, J.; Lee, J.K.; Lee, K.M. Deeply-recursive convolutional network for image super-resolution. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; pp. 1637–1645. [Google Scholar]
- Agustsson, E.; Timofte, R. Ntire 2017 challenge on single image super-resolution: Dataset and study. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Honolulu, HI, USA, 21–26 July 2017; pp. 126–135. [Google Scholar]
- Bevilacqua, M.; Roumy, A.; Guillemot, C.; Alberi-Morel, M.L. Low-complexity single-image super-resolution based on nonnegative neighbor embedding. In Proceedings of the British Machine Vision Conference (BMVC), Guildford, UK, 3–7 September 2012. [Google Scholar]
- Zeyde, R.; Elad, M.; Protter, M. On single image scale-up using sparse-representations. In Proceedings of the International Conference on Curves and Surfaces; Springer: Berlin/Heidelberg, Germany, 2010; pp. 711–730. [Google Scholar]
- Martin, D.; Fowlkes, C.; Tal, D.; Malik, J. A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics. In Proceedings of the Eighth IEEE International Conference on Computer Vision; ICCV 2001; IEEE: Piscataway, NJ, USA, 2001; Volume 2, pp. 416–423. [Google Scholar]
- Huang, J.B.; Singh, A.; Ahuja, N. Single image super-resolution from transformed self-exemplars. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA, 7–12 June 2015; pp. 5197–5206. [Google Scholar]
- Matsui, Y.; Ito, K.; Aramaki, Y.; Fujimoto, A.; Ogawa, T.; Yamasaki, T.; Aizawa, K. Sketch-based manga retrieval using manga109 dataset. Multimed. Tools Appl. 2017, 76, 21811–21838. [Google Scholar] [CrossRef] [Scilit]
- Zhang, K.; Zuo, W.; Gu, S.; Zhang, L. Learning deep CNN denoiser prior for image restoration. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 3929–3938. [Google Scholar]
- Luo, X.; Xie, Y.; Zhang, Y.; Qu, Y.; Li, C.; Fu, Y. LatticeNet: Towards Lightweight Image Super-Resolution with Lattice Block. In Proceedings of the Computer Vision—ECCV 2020; Springer International Publishing: Cham, Switzerland, 2020; pp. 272–289. [Google Scholar]
- Sun, L.; Pan, J.; Tang, J. Shufflemixer: An efficient convnet for image super-resolution. Adv. Neural Inf. Process. Syst. 2022, 35, 17314–17326. [Google Scholar] [CrossRef] [Scilit]
- Ruangsang, W.; Aramvith, S.; Onoye, T. Multi-FusNet of cross channel network for image super-resolution. IEEE Access 2023, 11, 56287–56299. [Google Scholar] [CrossRef] [Scilit]
- Wang, L.; Li, K.; Tang, J.; Liang, Y. Image super-resolution via lightweight attention-directed feature aggregation network. ACM Trans. Multimed. Comput. Commun. Appl. 2023, 19, 1–23. [Google Scholar] [CrossRef] [Scilit]
- Wang, Z.; Gao, G.; Li, J.; Yan, H.; Zheng, H.; Lu, H. Lightweight feature de-redundancy and self-calibration network for efficient image super-resolution. ACM Trans. Multimed. Comput. Commun. Appl. 2023, 19, 1–15. [Google Scholar] [CrossRef] [Scilit]
- Wan, C.; Yu, H.; Li, Z.; Chen, Y.; Zou, Y.; Liu, Y.; Yin, X.; Zuo, K. Swift parameter-free attention network for efficient super-resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 6246–6256. [Google Scholar]
- Zhang, L.; Wan, Y. Partial convolutional reparameterization network for lightweight image super-resolution. J. Real-Time Image Process. 2024, 21, 187. [Google Scholar] [CrossRef] [Scilit]
- Liu, C.; Gao, G.; Wu, F.; Guo, Z.; Yu, Y. An efficient feature reuse distillation network for lightweight image super-resolution. Comput. Vis. Image Underst. 2024, 249, 104178. [Google Scholar] [CrossRef] [Scilit]
- Wang, S.; Liu, G.; Liu, X.; Liao, X.; Ren, C. Efficient image super resolution via Mixed Window and Dimension Interaction. Neurocomputing 2025, 620, 129211. [Google Scholar] [CrossRef] [Scilit]
- Chen, J.; Mao, M.; Guan, A.; Ayush, A. Residual trio feature network for efficient super-resolution. Complex Intell. Syst. 2025, 11, 9. [Google Scholar] [CrossRef] [Scilit]





| Scale | Method | Params (K) | MACs (G) | Set5 | Set14 | B100 | Urban100 | Manga109 |
|---|---|---|---|---|---|---|---|---|
| PSNR/SSIM | PSNR/SSIM | PSNR/SSIM | PSNR/SSIM | PSNR/SSIM | ||||
| EDSR-baseline | 1370 | 316.30 | 37.99/0.9604 | 33.57/0.9175 | 32.16/0.8994 | 31.98/0.9272 | 38.54/0.9769 | |
| LatticeNet | 756 | 169.50 | 38.15/0.9610 | 33.78/0.9193 | 32.25/0.9005 | 32.43/0.9302 | -/- | |
| ShuffleMixer | 394 | 45.05 | 38.01/0.9606 | 33.63/0.9180 | 32.17/0.8995 | 31.89/0.9257 | 39.83/0.9774 | |
| MFCC | 1861 | 364.50 | 38.16/0.9611 | 33.85/0.9195 | 32.28/0.9010 | 32.65/0.9331 | 39.11/0.9780 | |
| AFAN | 1208 | 289.50 | 38.06/0.9608 | 33.74/0.9198 | 32.20/0.9001 | 32.38/0.9305 | 38.79/0.9773 | |
| FDSCSR | 823 | 215.80 | 38.12/0.9609 | 33.69/0.9191 | 32.24/0.9004 | 32.50/0.9315 | 38.89/0.9775 | |
| SPAN | 481 | - | 38.08/0.9608 | 33.71/0.9183 | 32.22/0.9002 | 32.24/0.9294 | 38.94/0.9777 | |
| PCRN | 368 | - | 38.05/0.9611 | 33.62/0.9184 | 32.23/0.9000 | 32.23/0.9286 | 39.05/0.9779 | |
| EFRDN | 768 | 111.60 | 38.03/0.9609 | 33.65/0.9185 | 32.16/0.8997 | 32.34/0.9298 | 38.86/0.9776 | |
| MWDIN | 536 | 112.80 | 38.16/0.9614 | 33.97/0.9206 | 32.33/0.9021 | 32.58/0.9313 | 39.31/0.9785 | |
| DCAE-Multi | - | - | 37.74/0.9594 | 33.29/0.9144 | 32.05/0.8977 | 31.44/0.9207 | -/- | |
| Ours | 659 | 145.17 | 38.18/0.9619 | 33.88/0.9220 | 32.38/0.9024 | 32.95/0.9360 | 39.41/0.9786 | |
| EDSR-baseline | 1555 | 160.10 | 34.37/0.9270 | 30.28/0.8417 | 29.09/0.8052 | 28.15/0.8527 | 33.45/0.9439 | |
| LatticeNet | 765 | 76.30 | 34.53/0.9281 | 30.39/0.8424 | 29.15/0.8059 | 28.33/0.8538 | -/- | |
| ESRT | 770 | - | 34.42/0.9268 | 30.43/0.8433 | 29.15/0.8063 | 28.46/0.8574 | 33.95/0.9455 | |
| ShuffleMixer | 415 | 21.50 | 34.40/0.9272 | 30.37/0.8423 | 29.12/0.8051 | 28.08/0.8498 | 33.69/0.9448 | |
| MFCC | 2230 | 374.10 | 34.67/0.9294 | 30.51/0.8456 | 29.22/0.8080 | 28.64/0.8616 | 34.15/0.9478 | |
| AFAN | 1208 | 143.10 | 34.46/0.9271 | 30.38/0.8436 | 29.11/0.8064 | 28.31/0.8556 | 33.61/0.9451 | |
| FDSCSR | 830 | 96.40 | 34.50/0.9281 | 30.43/0.8442 | 29.15/0.8068 | 28.40/0.8576 | 33.78/0.9460 | |
| EFRDN | 768 | 49.50 | 34.44/0.9275 | 30.38/0.8414 | 29.10/0.8059 | 28.29/0.8542 | 33.73/0.9453 | |
| MWDIN | 545 | 50.90 | 34.58/0.9290 | 30.48/0.8450 | 29.21/0.8093 | 28.43/0.8567 | 34.14/0.9471 | |
| DCAE-Multi | - | - | 34.12/0.9251 | 30.02/0.8353 | 29.00/0.8018 | 27.63/0.8394 | -/- | |
| Ours | 668 | 66.22 | 34.46/0.9304 | 30.25/0.8484 | 29.21/0.8112 | 28.89/0.8670 | 34.44/0.9501 | |
| EDSR-baseline | 1518 | 114.20 | 32.09/0.8938 | 28.58/0.7813 | 27.57/0.7357 | 26.04/0.7849 | 30.35/0.9067 | |
| LatticeNet | 777 | 43.60 | 32.30/0.8962 | 28.68/0.7830 | 27.62/0.7367 | 26.25/0.7873 | -/- | |
| ESRT | 751 | - | 32.19/0.8947 | 28.69/0.7833 | 27.69/0.7379 | 26.39/0.7962 | 30.75/0.9100 | |
| ShuffleMixer | 411 | 14.00 | 32.21/0.8953 | 28.66/0.7827 | 27.61/0.7366 | 26.08/0.7835 | 30.65/0.9093 | |
| MFCC | 2157 | 395.20 | 32.42/0.8973 | 28.73/0.7849 | 27.67/0.7399 | 26.48/0.7977 | 30.98/0.9131 | |
| AFAN | 1226 | 90.10 | 32.30/0.8951 | 28.66/0.7838 | 27.61/0.7383 | 26.27/0.7913 | 30.63/0.9109 | |
| FDSCSR | 839 | 54.80 | 32.36/0.8970 | 28.67/0.7840 | 27.63/0.7384 | 26.33/0.7935 | 30.69/0.9113 | |
| SPAN | 498 | - | 32.20/0.8953 | 28.66/0.7834 | 27.62/0.7374 | 26.18/0.7879 | 30.66/0.9103 | |
| PCRN | 389 | - | 32.28/0.8964 | 28.68/0.7832 | 27.64/0.7369 | 26.21/0.7883 | 30.68/0.9105 | |
| EFRDN | 767 | 27.90 | 32.33/0.8964 | 28.67/0.7833 | 27.63/0.7384 | 26.37/0.7939 | 30.76/0.9113 | |
| MWDIN | 557 | 29.40 | 32.42/0.8987 | 28.77/0.7857 | 27.70/0.7415 | 26.38/0.7920 | 31.06/0.9134 | |
| RTFN | - | - | 32.26/0.8953 | 28.63/0.7818 | 27.61/0.7364 | 26.27/0.7887 | 30.55/0.9075 | |
| DCAE-Multi | - | - | 31.79/0.8895 | 28.31/0.7741 | 27.42/0.7291 | 25.66/0.7695 | -/- | |
| Ours | 680 | 38.32 | 32.41/0.8996 | 28.55/0.7887 | 27.80/0.7443 | 26.64/0.8033 | 31.32/0.9187 |
| Model | Baseline | GRL | NLSAB | SWT | Params/MACs |
|---|---|---|---|---|---|
| 1 | ✓ | 792 K/45.29 G | |||
| 2 | ✓ | ✓ | 732 K/41.67 G | ||
| 3 | ✓ | ✓ | ✓ | 732 K/41.67 G | |
| 4 | ✓ | ✓ | ✓ | ✓ | 680 K/38.32 G |
| Model | Set5 | Set14 | B100 | Urban100 | Manga109 |
|---|---|---|---|---|---|
| 1 | 32.34/0.8992 | 28.56/0.7884 | 27.77/0.7433 | 26.57/0.8017 | 31.10/0.9172 |
| 2 | 32.39/0.8987 | 28.50/0.7875 | 27.75/0.7428 | 26.52/0.8007 | 31.15/0.9169 |
| 3 | 32.40/0.8991 | 28.51/0.7878 | 27.77/0.7434 | 26.60/0.8023 | 31.25/0.9175 |
| 4 | 32.41/0.8996 | 28.55/0.7887 | 27.80/0.7443 | 26.64/0.8033 | 31.32/0.9187 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Nan, J.; Wang, W.; Zhang, F.; Liu, Y.; Sun, L.; Chen, L.; Zheng, R.; Liao, Y.; Wei, L.; Fu, X.; et al. Enhanced Lightweight Image Super-Resolution via Residual Aggregation and Wavelet Loss. Electronics 2026, 15, 4271. https://doi.org/10.3390/electronics15184271
Nan J, Wang W, Zhang F, Liu Y, Sun L, Chen L, Zheng R, Liao Y, Wei L, Fu X, et al. Enhanced Lightweight Image Super-Resolution via Residual Aggregation and Wavelet Loss. Electronics. 2026; 15(18):4271. https://doi.org/10.3390/electronics15184271
Chicago/Turabian StyleNan, Jiahui, Wenkai Wang, Feng Zhang, Ying Liu, Li Sun, Longjia Chen, Renkui Zheng, Yiwen Liao, Longyu Wei, Xinyu Fu, and et al. 2026. "Enhanced Lightweight Image Super-Resolution via Residual Aggregation and Wavelet Loss" Electronics 15, no. 18: 4271. https://doi.org/10.3390/electronics15184271
APA StyleNan, J., Wang, W., Zhang, F., Liu, Y., Sun, L., Chen, L., Zheng, R., Liao, Y., Wei, L., Fu, X., & Song, J. (2026). Enhanced Lightweight Image Super-Resolution via Residual Aggregation and Wavelet Loss. Electronics, 15(18), 4271. https://doi.org/10.3390/electronics15184271

