Multi-View Industrial Image Super-Resolution via Hierarchical Multi-Scale Data Fusion
Abstract
1. Introduction
- (1)
- The proposed hybrid-resolution multi-view data fusion model architecture helps significantly enhance performance in large-scale SR.
- (2)
- This study uses reference-based SR to address the neglect of transferring high-frequency workpiece texture in HR images in previous studies, enabling realistic texture generation.
- (3)
- An SR model estimates complementary information between LR multi-view workpiece images to extract low-frequency contents to capture the overall geometry and silhouette of the object.
2. Literature Review
2.1. Preliminaries and Challenges
2.2. Multiple-View Image Super-Resolution
2.3. Reference-Based Super-Resolution
3. Methodology
3.1. Overview of Model Structure
3.2. Search and Transfer at Feature Levels
3.3. Super-Resolution at Image Levels
3.4. Loss Function
4. Experiment
4.1. Dataset and Data Processing
4.2. Experiment Results
4.2.1. Ablation Study
4.2.2. Results with Synthetic Images
- LFhybridSR [6]: The method captured the multi-view information of light field images by a hybrid-lens system, which takes an HR image in the center view while taking LR images in other views.
- ERVSR [11]: VSR reconstruction is performed using an HR intermediate frame in the video as a reference image, where optical flow is used to compute the similarity of different frames.
4.2.3. Results with Real Images
4.2.4. Depth Estimation Results with Real Images
5. Discussion
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Tian, J.; Ma, K.-K. A survey on super-resolution imaging. Signal Image Video Process. 2011, 5, 329–342. [Google Scholar] [CrossRef] [Scilit]
- Wang, Z.H.; Chen, J.; Hoi, S.C.H. Deep Learning for Image Super-Resolution: A Survey. IEEE Trans. Pattern Anal. Mach. Intell. 2021, 43, 3365–3387. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Cheng, R.; Sun, Y.; Yan, B.; Tan, W.; Ma, C. Geometry-Aware Reference Synthesis for Multi-View Image Super-Resolution. In Proceedings of the 30th ACM International Conference on Multimedia, Lisboa, Portugal, 10–14 October 2022; pp. 6083–6093. [Google Scholar]
- Zhao, M.D.; Wu, G.C.; Li, Y.P.; Hao, X.Y.; Fang, L.; Liu, Y.B. Cross-Scale Reference-Based Light Field Super-Resolution. IEEE Trans. Comput. Imaging 2018, 4, 406–418. [Google Scholar] [CrossRef] [Scilit]
- Al-Mekhlafi, H.; Liu, S. Single image super-resolution: A comprehensive review and recent insight. Front. Comput. Sci. 2023, 18, 181702. [Google Scholar] [CrossRef] [Scilit]
- Jin, J.; Hou, J.; Chen, J.; Kwong, S.; Yu, J. Light Field Super-resolution via Attention-Guided Fusion of Hybrid Lenses. In Proceedings of the 28th ACM International Conference on Multimedia, Seattle, WA, USA, 12–16 October 2020; pp. 193–201. [Google Scholar]
- Zheng, H.T.; Guo, M.H.; Wang, H.Q.; Liu, Y.B.; Fang, L. Combining Exemplar-based Approach and learning-based Approach for Light Field Super-resolution Using a Hybrid Imaging System. In Proceedings of the 2017 IEEE International Conference on Computer Vision Workshops (ICCVW), Venice, Italy, 22–29 October 2017; pp. 2481–2486. [Google Scholar]
- Boominathan, V.; Mitra, K.; Veeraraghavan, A. Improving Resolution and Depth-of-Field of Light Field Cameras Using a Hybrid Imaging System. In Proceedings of the 2014 IEEE International Conference on Computational Photography (ICCP), Santa Clara, CA, USA, 2–4 May 2014; pp. 1–10. [Google Scholar]
- Wang, Y.W.; Liu, Y.B.; Heidrich, W.; Dai, Q.H. The Light Field Attachment: Turning a DSLR into a Light Field Camera Using a Low Budget Camera Ring. IEEE Trans. Vis. Comput. Graph. 2017, 23, 2357–2364. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yang, F.Z.; Yang, H.; Fu, J.L.; Lu, H.T.; Guo, B.N. Learning Texture Transformer Network for Image Super-Resolution. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 5790–5799. [Google Scholar]
- Kim, Y.; Lim, J.; Cho, H.; Lee, M.; Lee, D.; Yoon, K.-J.; Choi, H.-J. Efficient reference-based video super-resolution (ERVSR): Single reference image is all you need. In Proceedings of the 2023 IEEE/CVF Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA, 2–7 January 2023; pp. 1828–1837. [Google Scholar]
- Chang, S.; Lin, Y.; Zhang, S. Flexible Hybrid Lenses Light Field Super-Resolution using Layered Refinement. In Proceedings of the 30th ACM International Conference on Multimedia, Lisboa, Portugal, 10–14 October 2022; pp. 5584–5592. [Google Scholar]
- Jin, J.; Guo, M.; Hou, J.; Liu, H.; Xiong, H. Light Field Reconstruction via Deep Adaptive Fusion of Hybrid Lenses. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 12050–12067. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lu, L.Y.; Li, W.B.; Tao, X.; Lu, J.B.; Jia, J.Y. MASA-SR: Matching Acceleration and Spatial Adaptation for Reference-Based Image Super-Resolution. In Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021; pp. 6364–6373. [Google Scholar]
- Cao, J.Z.; Liang, J.Y.; Zhang, K.; Li, Y.W.; Zhang, Y.L.; Wang, W.G.; Van Gool, L. Reference-Based Image Super-Resolution with Deformable Attention Transformer. In Computer Vision—ECCV 2022, Proceedings of the 2022 European Conference on Computer Vision (ECCV), Tel Aviv, Israel, October 23–27 2022; Springer: Cham, Switzerland, 2022; pp. 325–342. [Google Scholar]
- De Roovere, P.; Moonen, S.; Michiels, N.; Wyffels, F. Sim-to-Real Dataset of Industrial Metal Objects. Machines 2024, 12, 99. [Google Scholar] [CrossRef] [Scilit]
- Yeung, H.W.F.; Hou, J.H.; Chen, X.M.; Chen, J.; Chen, Z.B.; Chung, Y.Y. Light Field Spatial Super-Resolution Using Deep Efficient Spatial-Angular Separable Convolution. IEEE Trans. Image Process. 2019, 28, 2319–2330. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Richter, T.; Seiler, J.; Schnurrer, W.; Kaup, A. Robust Super-Resolution for Mixed-Resolution Multiview Image Plus Depth Data. IEEE Trans. Circuits Syst. Video Technol. 2016, 26, 814–828. [Google Scholar] [CrossRef] [Scilit]
- Su, H.; Li, Y.; Xu, Y.F.; Fu, X.; Liu, S. A review of deep-learning-based super-resolution: From methods to applications. Pattern Recognit. 2025, 157, 110935. [Google Scholar] [CrossRef] [Scilit]
- Dong, C.; Loy, C.C.; He, K.; Tang, X. Image super-resolution using deep convolutional networks. IEEE Trans. Pattern Anal. Mach. Intell. 2015, 38, 295–307. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- An, T.; Zhang, X.; Huo, C.L.; Xue, B.; Wang, L.F.; Pan, C.H. TR-MISR: Multiimage Super-Resolution Based on Feature Fusion With Transformers. IEEE J. Sel. Top. Appl. Earth Observ. Remote Sens. 2022, 15, 1373–1388. [Google Scholar] [CrossRef] [Scilit]
- Richard, A.; Cherabier, I.; Oswald, M.R.; Tsiminaki, V.; Pollefeys, M.; Schindler, K. Learned Multi-View Texture Super-Resolution. In Proceedings of the 2019 International Conference on 3D Vision (3DV), Quebec City, QC, Canada, 16–19 September 2019; pp. 533–543. [Google Scholar]
- Wang, L.G.; Wang, Y.Q.; Liang, Z.F.; Lin, Z.P.; Yang, J.G.; An, W.; Guo, Y.L. Learning Parallax Attention for Stereo Image Super-Resolution. In Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019; pp. 12242–12251. [Google Scholar]
- Chu, X.J.; Chen, L.Y.; Yu, W.Q. NAFSSR: Stereo Image Super-Resolution Using NAFNet. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), New Orleans, LA, USA, 19–20 June 2022; pp. 1238–1247. [Google Scholar]
- Chen, L.Y.; Chu, X.J.; Zhang, X.Y.; Sun, J. Simple Baselines for Image Restoration. In Computer Vision—ECCV 2022, Proceedings of the 2022 European Conference on Computer Vision (ECCV), Tel Aviv, Israel, October 23–27 2022; Springer: Cham, Switzerland, 2022; pp. 17–33. [Google Scholar]
- Yoon, Y.; Jeon, H.G.; Yoo, D.; Lee, J.Y.; Kweon, I.S. Learning a Deep Convolutional Network for Light-Field Image Super-Resolution. In Proceedings of the 2015 IEEE International Conference on Computer Vision Workshop (ICCVW), Santiago, Chile, 7–13 December 2015; pp. 57–65. [Google Scholar]
- Zhang, S.; Lin, Y.F.; Sheng, H. Residual Networks for Light Field Image Super-Resolution. In Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019; pp. 11038–11047. [Google Scholar]
- Mo, Y.; Wang, Y.Q.; Xiao, C.; Yang, J.G.; An, W. Dense Dual-Attention Network for Light Field Image Super-Resolution. IEEE Trans. Circuits Syst. Video Technol. 2022, 32, 4431–4443. [Google Scholar] [CrossRef] [Scilit]
- Chao, W.; Zhao, J.; Duan, F.; Wang, G. Lfsrdiff: Light field image super-resolution via diffusion models. In Proceedings of the ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Hyderabad, India, 6–11 April 2025; IEEE: New York, NY, USA, 2025; pp. 1–5. [Google Scholar]
- Shim, G.; Park, J.; Kweon, I.S. Robust Reference-based Super-Resolution with Similarity-Aware Deformable Convolution. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 8422–8431. [Google Scholar]
- Zheng, H.T.; Ji, M.Q.; Wang, H.Q.; Liu, Y.B.; Fang, L. CrossNet: An End-to-End Reference-Based Super Resolution Network Using Cross-Scale Warping. In Computer Vision—ECCV 2018, Proceedings of the 2018 European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; Springer: Cham, Switzerland, 2018; pp. 87–104. [Google Scholar]
- Zhang, Z.F.; Wang, Z.W.; Lin, Z.; Qi, H.R. Image Super-Resolution by Neural Texture Transfer. In Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019; pp. 7974–7983. [Google Scholar]
- Zhao, W.; Lee, C.K.; Cheung, B.C.F. Reference-based Super-resolution for Light Field Image. In Proceedings of the Third Australian International Conference on Industrial Engineering and Operations Management, Sydney, Australia, 24–26 September 2024; IEOM Society International: Southfield, MI, USA, 2024. [Google Scholar]
- Jin, J.; Hou, J.H.; Chen, J.; Yeung, H.; Kwong, S. Light Field Spatial Super-resolution via CNN Guided by A Single High-resolution RGB Image. In Proceedings of the 2018 IEEE 23rd International Conference on Digital Signal Processing (DSP), Shanghai, China, 19–21 November 2018. [Google Scholar]
- Lai, W.S.; Huang, J.B.; Ahuja, N.; Yang, M.H. Fast and Accurate Image Super-Resolution with Deep Laplacian Pyramid Networks. IEEE Trans. Pattern Anal. Mach. Intell. 2019, 41, 2599–2613. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ranftl, R.; Lasinger, K.; Hafner, D.; Schindler, K.; Koltun, V. Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-Shot Cross-Dataset Transfer. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 44, 1623–1637. [Google Scholar] [CrossRef] [Scilit] [PubMed]








| Method | Category | Multiple LR Input | Use HR Reference | Strengths | Limitations |
|---|---|---|---|---|---|
| Dong et al. [20] | SISR | ✕ | ✕ | Simple and fast | Over-smooths low-texture surfaces |
| MVSRnet [3] | MVISR | ✕ | ✕ | Geometry-aware warping; fuses complementary LR views | Requires depth priors; fails under large parallax and reflections |
| PASSRnet [23] | Stereo SR | ✓ (stereo pair) | ✕ | Parallax attention for cross-view fusion | Limited to two views; no HR texture guidance |
| LFhybridSR [6] | Light field SR | ✓ | ✓ | Hybrid input; attention-guided fusion | Not large parallax |
| ERVSR [11] | Reference-based video SR | ✓ (temporal) | ✓ | Optical flow alignment; efficient | Temporal assumption mismatched with angular views |
| TTSR [10] | Reference-based SR | ✕ | ✓ | Transformer-based texture transfer | Single view only; Not large-scale |
| Images | Information | Value |
|---|---|---|
| Real-world image | Spatial resolution (HR) | 1920 × 1080 |
| Spatial resolution (LR) | 240 × 135 | |
| Synthetic image | Spatial resolution (HR) | 2560 × 2048 |
| Spatial resolution (LR) | 320 × 256 | |
| Angular resolution | 3 × 3 |
| Variant | Description | PSNR | SSIM | LPIPS |
|---|---|---|---|---|
| A | a weight-shared encoder for feature extraction | 35.30 | 0.847 | 0.226 |
| B | remove the SAS block | 35.16 | 0.842 | 0.241 |
| C | remove the fusion module | 35.17 | 0.840 | 0.244 |
| D | L1 loss function | 35.33 | 0.847 | 0.227 |
| Ours | proposed full model | 35.37 | 0.848 | 0.221 |
| Scenes | Bicubic (PSNR/SSIM) | LFhybridSR (PSNR/SSIM) | ERVSR (PSNR/SSIM) | Ours (PSNR/SSIM) |
|---|---|---|---|---|
| 1 | 33.06/0.904 | 33.56/0.882 | 37.71/0.935 | 42.25/0.937 |
| 2 | 32.52/0.894 | 33.05/0.874 | 36.96/0.926 | 41.32/0.930 |
| 3 | 31.29/0.875 | 31.81/0.862 | 35.36/0.910 | 39.68/0.924 |
| 4 | 32.93/0.896 | 33.45/0.874 | 37.39/0.927 | 41.84/0.930 |
| 5 | 33.06/0.899 | 33.53/0.878 | 37.56/0.930 | 42.06/0.933 |
| 6 | 32.33/0.885 | 32.66/0.871 | 36.71/0.918 | 41.27/0.929 |
| Mean | 32.53/0.892 | 33.01/0.874 | 36.95/0.924 | 41.40/0.930 |
| Scenes | Bicubic (PSNR/SSIM) | LFhybridSR (PSNR/SSIM) | ERVSR (PSNR/SSIM) | Ours (PSNR/SSIM) |
|---|---|---|---|---|
| 1 | 33.39/0.755 | 33.60/0.747 | 28.56/0.785 | 34.73/0.814 |
| 2 | 34.36/0.823 | 34.50/0.808 | 30.67/0.855 | 35.88/0.878 |
| 3 | 34.19/0.816 | 34.21/0.797 | 30.05/0.847 | 35.53/0.869 |
| 4 | 35.86/0.841 | 35.73/0.837 | 34.59/0.856 | 36.61/0.870 |
| 5 | 33.86/0.796 | 33.98/0.781 | 29.22/0.830 | 35.24/0.855 |
| 6 | 32.94/0.734 | 33.14/0.724 | 27.97/0.769 | 34.23/0.799 |
| Mean | 34.1/0.794 | 34.19/0.782 | 30.18/0.823 | 35.37/0.848 |
| Input | Abs Rel | δ1 | δ2 | δ3 |
|---|---|---|---|---|
| SR | 0.1024 | 0.8989 | 0.9383 | 0.9521 |
| LR | 0.5349 | 0.4007 | 0.5951 | 0.6569 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Zhao, W.; Lee, C.K.M.; Li, D.; Cheung, B.C.F. Multi-View Industrial Image Super-Resolution via Hierarchical Multi-Scale Data Fusion. AI 2026, 7, 172. https://doi.org/10.3390/ai7050172
Zhao W, Lee CKM, Li D, Cheung BCF. Multi-View Industrial Image Super-Resolution via Hierarchical Multi-Scale Data Fusion. AI. 2026; 7(5):172. https://doi.org/10.3390/ai7050172
Chicago/Turabian StyleZhao, Wenqin, Carman Ka Man Lee, Da Li, and Benny Chi Fai Cheung. 2026. "Multi-View Industrial Image Super-Resolution via Hierarchical Multi-Scale Data Fusion" AI 7, no. 5: 172. https://doi.org/10.3390/ai7050172
APA StyleZhao, W., Lee, C. K. M., Li, D., & Cheung, B. C. F. (2026). Multi-View Industrial Image Super-Resolution via Hierarchical Multi-Scale Data Fusion. AI, 7(5), 172. https://doi.org/10.3390/ai7050172

