Next Article in Journal
A Wearable-Sensor System with AI Technology for Real-Time Biomechanical Feedback Training in Hammer Throw
Next Article in Special Issue
Optical Panel Inspection Using Explicit Band Gaussian Filtering Methods in Discrete Cosine Domain
Previous Article in Journal
Current Technologies for Detection of COVID-19: Biosensors, Artificial Intelligence and Internet of Medical Things (IoMT): Review
Previous Article in Special Issue
Inline Quality Monitoring of Reverse Extruded Aluminum Parts with Cathodic Dip-Paint Coating (KTL)
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Conv-Former: A Novel Network Combining Convolution and Self-Attention for Image Quality Assessment

1
Changchun Institute of Optics, Fine Mechanics and Physics, Chinese Academy of Sciences, Changchun 130033, China
2
College of Materials Science and Opto-Electronic Technology, University of Chinese Academy of Sciences, Beijing 100049, China
3
Department of Electrical and Optical Engineering, Space Engineering University, Beijing 101416, China
*
Author to whom correspondence should be addressed.
Sensors 2023, 23(1), 427; https://doi.org/10.3390/s23010427
Submission received: 7 November 2022 / Revised: 21 December 2022 / Accepted: 26 December 2022 / Published: 30 December 2022

Abstract

To address the challenge of no-reference image quality assessment (NR-IQA) for authentically and synthetically distorted images, we propose a novel network called the Combining Convolution and Self-Attention for Image Quality Assessment network (Conv-Former). Our model uses a multi-stage transformer architecture similar to that of ResNet-50 to represent appropriate perceptual mechanisms in image quality assessment (IQA) to build an accurate IQA model. We employ adaptive learnable position embedding to handle images with arbitrary resolution. We propose a new transformer block (TB) by taking advantage of transformers to capture long-range dependencies, and of local information perception (LIP) to model local features for enhanced representation learning. The module increases the model’s understanding of the image content. Dual path pooling (DPP) is used to keep more contextual image quality information in feature downsampling. Experimental results verify that Conv-Former not only outperforms the state-of-the-art methods on authentic image databases, but also achieves competing performances on synthetic image databases which demonstrate the strong fitting performance and generalization capability of our proposed model.
Keywords: image quality assessment; vision transformer; self-attention; neural network; deep model interation image quality assessment; vision transformer; self-attention; neural network; deep model interation

Share and Cite

MDPI and ACS Style

Han, L.; Lv, H.; Zhao, Y.; Liu, H.; Bi, G.; Yin, Z.; Fang, Y. Conv-Former: A Novel Network Combining Convolution and Self-Attention for Image Quality Assessment. Sensors 2023, 23, 427. https://doi.org/10.3390/s23010427

AMA Style

Han L, Lv H, Zhao Y, Liu H, Bi G, Yin Z, Fang Y. Conv-Former: A Novel Network Combining Convolution and Self-Attention for Image Quality Assessment. Sensors. 2023; 23(1):427. https://doi.org/10.3390/s23010427

Chicago/Turabian Style

Han, Lintao, Hengyi Lv, Yuchen Zhao, Hailong Liu, Guoling Bi, Zhiyong Yin, and Yuqiang Fang. 2023. "Conv-Former: A Novel Network Combining Convolution and Self-Attention for Image Quality Assessment" Sensors 23, no. 1: 427. https://doi.org/10.3390/s23010427

APA Style

Han, L., Lv, H., Zhao, Y., Liu, H., Bi, G., Yin, Z., & Fang, Y. (2023). Conv-Former: A Novel Network Combining Convolution and Self-Attention for Image Quality Assessment. Sensors, 23(1), 427. https://doi.org/10.3390/s23010427

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop