Next Article in Journal
Optimization of Ca–Al–Mn–Si Substitution Level for Enhanced Magnetic Properties of M-Type Sr-Hexaferrites for Permanent Magnet Application
Previous Article in Journal
Studying Corrosion Failure Prediction Models and Methods for Submarine Oil and Gas Transport Pipelines
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Transformer-Based Visual Object Tracking with Global Feature Enhancement

1
School of Computer Science and Technology, Anhui Engineering Research Center for Intelligent Computing and Application on Cognitive Behavior (ICACB), Huaibei Normal University, Huaibei 235000, China
2
College of Electronic and Information Engineering, Hebei University, Baoding 071000, China
3
Department of Geography, Universidade Estadual do Maranhão, São Luís 65055-000, Brazil
*
Author to whom correspondence should be addressed.
Appl. Sci. 2023, 13(23), 12712; https://doi.org/10.3390/app132312712
Submission received: 27 October 2023 / Revised: 19 November 2023 / Accepted: 21 November 2023 / Published: 27 November 2023

Abstract

With the rise of general models, transformers have been adopted in visual object tracking algorithms as feature fusion networks. In these trackers, self-attention is used for global feature enhancement. Cross-attention is applied to fuse the features of the template and the search regions to capture the global information of the object. However, studies have found that the feature information fused by cross-attention does not pay enough attention to the object region. In order to enhance cross-attention for the object region, an enhanced cross-attention (ECA) module is proposed for global feature enhancement. By calculating the average attention score for each position in the fused feature sequence and assigning higher weights to the positions with higher attention scores, the proposed ECA module can improve the feature information in the object region and further enhance the matching accuracy. In addition, to reduce the computational complexity of self-attention, orthogonal random features are introduced to implement a fast attention operation. This decomposes the attention matrix into the product of a random non-linear function between the original query and key. This module can reduce the spatial complexity and improve the inference speed by avoiding the explicit construction of a quadratic attention matrix. Finally, a tracking method named GFETrack is proposed, which comprises a Siamese backbone network and an enhanced attention mechanism. Experimental results show that the proposed GFETrack achieves competitive results on four challenging datasets.
Keywords: visual object tracking; transformer; global feature enhancement; fast self-attention; enhanced cross-attention visual object tracking; transformer; global feature enhancement; fast self-attention; enhanced cross-attention

Share and Cite

MDPI and ACS Style

Wang, S.; Fang, G.; Liu, L.; Wang, J.; Zhu, K.; Melo, S.N. Transformer-Based Visual Object Tracking with Global Feature Enhancement. Appl. Sci. 2023, 13, 12712. https://doi.org/10.3390/app132312712

AMA Style

Wang S, Fang G, Liu L, Wang J, Zhu K, Melo SN. Transformer-Based Visual Object Tracking with Global Feature Enhancement. Applied Sciences. 2023; 13(23):12712. https://doi.org/10.3390/app132312712

Chicago/Turabian Style

Wang, Shuai, Genwen Fang, Lei Liu, Jun Wang, Kongfen Zhu, and Silas N. Melo. 2023. "Transformer-Based Visual Object Tracking with Global Feature Enhancement" Applied Sciences 13, no. 23: 12712. https://doi.org/10.3390/app132312712

APA Style

Wang, S., Fang, G., Liu, L., Wang, J., Zhu, K., & Melo, S. N. (2023). Transformer-Based Visual Object Tracking with Global Feature Enhancement. Applied Sciences, 13(23), 12712. https://doi.org/10.3390/app132312712

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop